Anchor Box
When a detector scans a street photo, it needs a practical starting guess for where cars, people, and signs could be. An anchor box is that starting guess: a set of pre-placed reference rectangles spread across an image or its feature map.
How anchors guide detectionEach anchor has a position, width, height, and aspect ratio. Rather than inventing every possible bounding box from scratch, an anchor-based model asks two focused questions for each reference rectangle: “Does this contain an object?” and “How should I shift and resize it to fit the object?” The model predicts small numeric adjustments—such as moving the center left or making the box taller—then produces a refined bounding box and a class label. It is like beginning with a sheet of differently shaped picture frames, then adjusting the best frame around each visible object.
Training and matchingDuring training, anchors are compared with the human-labeled boxes using intersection over union (IoU). Anchors with high overlap become positive examples; anchors with little overlap become background examples. A detector commonly uses multiple scales and shapes at every location so it can cover varied objects:
- Wide, short anchors help fit vehicles or text lines.
- Tall, narrow anchors help fit standing people.
- Larger anchors help detect nearby objects; smaller ones help find distant objects.
Anchors made fast detectors such as Faster R-CNN, SSD, and early YOLO versions effective because they turn open-ended localization into predictable regression from useful starting points. Poorly chosen anchor sizes can miss tiny defects in production-line inspection, crowd tightly packed faces, or struggle with unusually long medical structures. Detectors also generate many overlapping refined boxes, so non-maximum suppression removes duplicates. Newer anchor-free detectors avoid these predefined rectangles, but anchors remain an important way to understand how many widely used object detectors locate objects.
An anchor box is a predefined bounding box with a fixed size and aspect ratio placed at locations across an image or feature map. An object detector compares anchors with ground-truth objects, then predicts class labels and coordinate offsets that refine matching anchors into final detections. Anchor boxes let detectors efficiently represent objects of varied scales and shapes; poor anchor design reduces detection accuracy, especially for unusual object sizes or aspect ratios.
Imagine looking for houses in a city using a set of differently sized picture frames. You place small, medium, large, wide, and tall frames across the map, then check which frames best surround a house. An anchor box is like one of those ready-made frames.
In AI systems that find objects in images, anchor boxes give the system starting guesses for where an object might be and how big it might be. A small frame may suit a face in the distance; a wide frame may suit a car or bus. The AI then adjusts the best-fitting frame and labels what it found. This helps it spot many objects of different shapes and sizes in one picture.