Notes

Bounding Box

A bounding box is the rectangle a vision system draws around an object it has found: a car in a street scene, a face in a photo, or a defect on a manufactured part. It gives the model a compact way to say both what it sees and where it is.

How a box represents location
A standard, axis-aligned bounding box has sides parallel to the image edges. It is stored as four numbers, using one of two common conventions:

  • (xmin, ymin, xmax, ymax): the coordinates of the top-left and bottom-right corners.
  • (x, y, width, height): the top-left corner plus the rectangle’s size.

Coordinates are measured in pixels, or normalized to a 0–1 range so the representation works across image sizes. A detector outputs a box alongside a class label, such as “person,” and a confidence score. Training images contain human-provided boxes, called ground-truth annotations, that teach the model what correct localization looks like.

How detectors use them
A detector such as YOLO predicts many candidate boxes per image. Several can surround the same object, so non-maximum suppression (NMS) keeps the highest-confidence box and removes close duplicates. Accuracy is judged with Intersection over Union (IoU): the overlap area between a predicted and ground-truth box divided by their combined area. A box can identify the correct object but still count as inaccurate when it is shifted or too large.

Why boxes matter—and their limits
Bounding boxes make object detection practical for traffic monitoring, shelf inventory, face detection, and visual quality inspection. They are quick for people to label and efficient for models to predict. But a rectangle includes background around irregular shapes: a box around a tumor, bicycle, or spilled liquid does not mark its exact boundary. When precise outlines matter—such as measuring a lesion or separating touching cells—semantic or instance segmentation is the better tool. Rotated boxes are also used for angled objects such as aerial-view vehicles and text.

A bounding box is a rectangular annotation that encloses an object in an image, typically represented by its corner coordinates or center position, width, and height. In object detection, models predict bounding boxes alongside class labels and confidence scores. They provide the standard representation for training, localizing objects, and evaluating detections using overlap measures such as Intersection over Union.

Imagine drawing a neat rectangle around every item you want someone to notice in a photo: one around a dog, another around a bicycle, and another around a stop sign. A bounding box is that rectangle.

In computer vision, it tells an AI not just “there is a dog here,” but also roughly where the dog is in the image. This matters for things like self-driving cars spotting pedestrians, security cameras finding people, or photo apps identifying objects.

The box does not trace an object’s exact outline. It is simply the smallest useful rectangle that contains it, making the object quick and practical for an AI to locate.