Detection Head
A detection head is the part of an object-detection model that turns visual features into usable answers: what is in this image, where is it, and how confident is the model? Think of the earlier layers as gathering clues from pixels, while the head makes the final labeled bounding-box predictions.
What the head produces
A vision model’s backbone, such as ResNet or a Vision Transformer, converts an image into feature maps: compact representations of edges, textures, shapes, and object parts. A detection head reads those features and predicts, for many candidate locations:
- Class scores, such as “car,” “person,” or “scratch.”
- Bounding-box coordinates, usually expressed as a box center, width, and height, or corner positions.
- Objectness/confidence, indicating whether a candidate region actually contains an object.
For example, a YOLO head makes dense predictions across a grid of image locations. A Faster R-CNN head instead classifies and refines a smaller set of proposed regions. Afterward, non-maximum suppression removes duplicate boxes that refer to the same object.
Why it matters in transfer learning
When adapting a pretrained detector to a new dataset, the detection head is commonly replaced or retrained. A model trained to detect 80 everyday categories cannot directly output “defective seal” or “tumor region” if its final head has no outputs for those classes. The backbone’s reusable visual knowledge can be retained, while a new head learns the new class count and the target dataset’s box patterns. In PyTorch’s torchvision, this is reflected in replacing a Faster R-CNN model’s box predictor.
Practical consequences
The head strongly affects speed, accuracy, and what errors a detector makes. A lightweight head supports real-time pedestrian detection in video; a richer multi-scale head helps find tiny defects on a production line or small lesions in scans. Poorly matched heads can miss small objects, produce imprecise boxes, or confuse visually similar classes—even when the backbone itself is strong.
A detection head is the task-specific output module attached to a vision model’s pretrained backbone. It converts extracted visual features into predictions such as object classes, bounding-box coordinates, and confidence scores. In transfer learning, replacing or fine-tuning the detection head adapts a general pretrained model to a new labeled object-detection dataset while retaining useful visual representations from the backbone.
Think of an AI vision system like a person looking at a photo. Most of its “brain” notices broad visual clues: edges, textures, shapes, and patterns. The detection head is the final specialist that turns those clues into a useful answer.
For example, after examining a street image, it may say: “There is a car here,” “a pedestrian there,” and draw a box around each one. In transfer learning, a model can keep its experienced visual “eyes” and swap or retrain this final part for a new job—such as spotting defects on products instead of animals in photos.