Stereo Vision
Stereo vision gives a computer a form of depth perception by comparing two images of the same scene taken from slightly different viewpoints. It works much like human eyes: nearby objects appear to shift more between the left and right views than distant objects do.
From image shift to depthThe key measurement is disparity: the horizontal position difference between the same physical point in the two images. Before measuring it, a system calibrates the cameras to learn their focal length and relative placement, then rectifies both images so matching points fall on the same image row. For a calibrated stereo pair, depth follows the relationship Z = fB / d, where Z is depth, f is focal length, B is the distance between cameras (the baseline), and d is disparity. A large disparity means the point is close; a small one means it is far away.
Finding reliable matchesThe difficult part is stereo matching: deciding which pixel in the left image corresponds to each pixel in the right image. Algorithms compare local image patterns along each row and build a disparity map, an image whose pixel values represent distance-related shifts. Semi-Global Block Matching (SGBM), available as OpenCV’s cv::StereoSGBM, balances accuracy and speed. Smooth, textureless walls, repeating patterns, reflections, and areas visible to only one camera create ambiguous or missing matches, so practical systems filter unreliable depth estimates.
Why depth changes what vision can doA depth map lets a system reason about physical layout rather than only appearance. For example:
- An autonomous vehicle can estimate the distance to a pedestrian or parked car.
- A robot can locate a box in front of a shelf and plan a safe grasp.
- A factory inspection system can detect whether a component sits too high, low, or far forward.
- A phone camera can separate a face from its background for portrait effects.
Without stereo depth, two objects of different sizes can look equally large in an image, making distance hard to infer. Stereo vision supplies that missing geometric evidence directly from paired cameras.
Stereo vision estimates scene depth and 3D structure by comparing images of the same scene captured from two or more cameras at different viewpoints. Corresponding pixel locations produce disparity, which converts to distance using camera geometry. It enables robots, vehicles, and augmented-reality systems to perceive spatial layout, avoid obstacles, and measure object distances without relying solely on dedicated depth sensors.
Imagine holding one finger in front of your face and closing each eye in turn. Your finger seems to shift against the background. Your brain uses that tiny difference between your two views to judge how far away things are.
Stereo vision gives a machine a similar ability. It uses two cameras placed a short distance apart, like a pair of eyes, to view the same scene from slightly different angles. By comparing those views, the system can estimate depth: what is near, what is far, and where objects sit in three-dimensional space. This helps robots avoid obstacles, cars detect road hazards, and machines navigate real-world spaces more safely.