Notes

Image Pyramid

An object can look completely different to a vision system when it appears near the camera versus far away: a nearby car may occupy hundreds of pixels, while a distant one occupies only a few. An image pyramid gives the system several versions of the same image at different sizes, so it can examine visual patterns at the scale where they are easiest to recognize.

How the pyramid is built
A typical image pyramid starts with the original, full-resolution image as its bottom level. Each higher level is created by first applying a low-pass blur, which removes fine detail that would cause artifacts, then downsampling—for example, keeping every second pixel in each direction. This produces progressively smaller images: 1024×1024, 512×512, 256×256, and so on. The stack resembles a pyramid because image area shrinks rapidly at higher levels. A Gaussian pyramid uses Gaussian blur before shrinking; it is the most common form.

Why multiple resolutions help
Instead of redesigning a detector for every possible object size, a system can apply the same detector at each pyramid level. A face detector trained to find a 40-pixel-wide face can find:

  • a large face in a reduced version of the image,
  • a small face in the original image, and
  • medium-sized faces at intermediate levels.
This was central to classical sliding-window detectors using features such as HOG. It also supports image alignment and optical flow: matching a small, blurry version first makes large movements easier to estimate, then finer levels refine the result.

Related pyramid forms
A Laplacian pyramid stores the detail lost between adjacent Gaussian-pyramid levels, making it useful for image blending, compression, and reconstruction. Modern deep-learning detectors use the same core idea in Feature Pyramid Networks (FPNs): rather than resizing only raw images, they combine feature maps from different depths so a detector can recognize both tiny distant objects and large nearby ones. OpenCV provides practical tools such as cv2.pyrDown and cv2.pyrUp for constructing these levels.

An image pyramid is a sequence of progressively lower-resolution versions of an image, created by repeated smoothing and downsampling; related forms store differences between levels. It represents visual information across spatial scales, so algorithms can analyze both fine detail and large structures. Image pyramids are essential for scale-robust feature detection, image blending, registration, and efficient coarse-to-fine matching.

Imagine looking at the same photograph from different distances: up close, you notice tiny details; farther away, you see the big shapes and overall scene. An image pyramid gives a computer those same views. It creates several versions of one image, each smaller and less detailed than the last, stacked like a pyramid.

This matters because objects can appear at very different sizes. A face close to a camera and a face across the street should both be recognized. By examining images at multiple sizes, computer vision systems can spot small details as well as large patterns, making tasks like finding objects, tracking movement, and matching photos more reliable.