Notes

Feature Pyramid Network (FPN)

Objects in a photograph do not arrive at a convenient single size: a distant traffic sign might occupy a few pixels, while a nearby bus fills much of the frame. A Feature Pyramid Network (FPN) helps an object detector stay alert to both by giving it useful image features at several spatial scales.

How the pyramid is built
A convolutional backbone such as ResNet naturally creates a hierarchy of feature maps. Early layers preserve fine detail and have high resolution, but understand little about meaning. Deeper layers recognize richer concepts—such as “car” or “face”—but are low resolution, so small objects can disappear. FPN combines the strengths of both through two paths:

  • A bottom-up pathway: the backbone produces progressively smaller, more semantic feature maps.
  • A top-down pathway: semantic information from deep layers is upsampled and passed back toward higher-resolution layers.
  • Lateral connections: features from matching levels are merged, usually by addition after a small convolution aligns their channel dimensions.

How detectors use it
The result is a set of maps, commonly called P2, P3, P4, P5, where each level has strong semantic information but represents a different object scale. A detector can assign small objects to a high-resolution level and large objects to a lower-resolution level. In Faster R-CNN with FPN, both region proposals and final object classification use these multi-scale features. Mask R-CNN also relies on FPN to produce detailed instance masks.

Why it matters
Without an FPN, a model forced to detect every object from one feature-map resolution performs poorly when sizes vary widely—especially for small pedestrians, defects on a production line, tiny tumors in scans, or distant vehicles. FPN made multi-scale detection practical without repeatedly running the whole image through separate resized networks. Modern detection frameworks, including Detectron2 and many RetinaNet-style models, use this idea because it improves scale handling while keeping computation manageable.

A Feature Pyramid Network (FPN) is a neural-network architecture that combines feature maps from multiple depths into a pyramid of semantically strong, multi-resolution representations. It gives detectors dedicated features for objects at different scales, from small to large. FPNs are important because scale variation is a central challenge in object detection and segmentation; they improve accuracy while reusing a backbone’s existing feature hierarchy efficiently.

Think of looking for animals in a landscape using both binoculars and a wide-angle view. The wide view helps you spot large things, like a bus or elephant, while the close view helps you notice small things, like a bird or stop sign.

A Feature Pyramid Network (FPN) gives an AI vision system these different “views” of an image at once. It helps the system find objects of many sizes, even when tiny and large objects appear together. This matters for tasks like spotting pedestrians, cars, and distant bicycles in a street photo. An FPN is not a type of unsupervised learning; it is a visual helper that can be used with models trained in different ways.