Average Precision (AP)
Imagine reviewing a detector’s predictions from most confident to least confident and asking: “How reliably does it find the real objects before it starts making too many mistakes?” Average Precision (AP) captures the answer in one score, rewarding detectors that find objects accurately and rank their best predictions first.
How AP is calculatedFor one object class, such as cars, every predicted bounding box has a confidence score. Predictions are sorted from highest to lowest score, then matched against the labeled objects in the image. A prediction counts as a true positive when it has the correct class and overlaps an unused ground-truth box enough, measured by Intersection over Union (IoU). Otherwise, it is a false positive. Missed labeled objects contribute to false negatives.
- Precision asks: among predictions accepted so far, what fraction is correct?
- Recall asks: what fraction of all real objects has been found?
- Plotting precision against recall at progressively lower confidence thresholds produces a precision–recall curve.
AP is the area under this curve. A high AP means the detector maintains strong precision while reaching high recall; it finds most objects without flooding the result with incorrect boxes.
Why the matching rule mattersAP always depends on an IoU threshold. AP@0.5 treats a detection as correct when IoU is at least 0.50, while AP@0.75 demands tighter box placement. The COCO benchmark’s headline AP averages results across IoU thresholds from 0.50 to 0.95, making it sensitive to both recognition and localization quality. For example, a production-line detector can identify every defective bottle but still earn a lower strict AP if its boxes are poorly aligned. Mean Average Precision (mAP) then averages AP across classes, providing one headline score for a multi-class detector. This makes AP central when comparing models such as YOLO, Faster R-CNN, and DETR: accuracy alone cannot reveal whether their confidence-ranked detections are genuinely dependable.
Average Precision (AP) is an object-detection metric that summarizes a model’s precision–recall curve for one class, typically by integrating precision across recall levels. A prediction counts as correct when its class is correct and its bounding box meets a specified intersection over union (IoU) threshold. AP measures both localization and classification quality; it is the basis for mean Average Precision (mAP), the standard detector benchmark.
Imagine judging a lost-and-found helper. You want it to find as many missing items as possible, but you also do not want it loudly claiming that every ordinary object is a match. Average Precision (AP) is a score for how well an AI object detector strikes that balance.
For example, when spotting bicycles in photos, a strong detector finds most real bicycles and rarely draws a box around something that is not a bicycle. AP combines these two qualities—finding the right things and avoiding false alarms—into one easy-to-compare score. Higher AP means the detector is more reliable across different levels of confidence.