Notes

Disparity Map

A disparity map is an image in which each pixel records how far that point appears to shift between two views of the same scene. It is the key intermediate result that lets a stereo camera pair turn ordinary photographs into an estimate of 3D shape and distance.

How the shift reveals depth
Stereo vision imitates human eyesight: a left and right camera are separated by a known baseline. A nearby object appears in noticeably different horizontal positions in the two images, while a distant object barely shifts. For corresponding pixels, disparity is commonly written as d = xleft − xright, measured in pixels. After the cameras are rectified, matching points lie on the same image row, making this a one-dimensional search. Depth follows the relationship Z = fB / d, where f is focal length and B is the camera separation. Large disparity means close; small disparity means far.

What the map looks like

  • In a grayscale disparity map, brighter regions commonly represent larger disparity and therefore nearer surfaces.
  • A dense map assigns a value to nearly every image pixel, preserving object boundaries and surface shape.
  • Invalid or unreliable areas occur where one camera cannot see a surface, such as behind an object edge, or where texture is too repetitive to match confidently.

Why it matters in practice
A disparity map gives autonomous vehicles an immediate cue for separating a nearby pedestrian from a distant background, helps robots judge where to grasp an object, and supports 3D measurement in factory inspection. Stereo matching algorithms compare small image patches or learned features to find the best correspondence; a widely used classical option is OpenCV’s StereoSGBM. Poor calibration, reflective surfaces, shadows, and blank walls can produce incorrect disparities, which then turn into incorrect depth. Reliable disparity estimation is therefore central to safe distance-aware vision rather than merely an attractive visualization.

A disparity map is an image in which each pixel stores the horizontal offset between corresponding points in a calibrated stereo image pair. This offset, or disparity, is inversely related to scene depth: larger disparities indicate closer surfaces. Disparity maps enable dense 3D reconstruction, depth estimation, obstacle detection, and robot navigation from two-camera imagery.

Hold a finger in front of your face and close one eye, then the other. Your finger seems to jump more than the distant background. A disparity map is a picture that records this kind of shift for every part of a scene, using two images taken from slightly different viewpoints, like human eyes.

Areas that shift a lot are usually close; areas that barely shift are farther away. This gives an AI a useful sense of depth: where a road lies, how far away a pedestrian is, or whether a robot can safely grasp an object. It helps machines turn flat images into an understanding of 3D space.