Bounding Box Annotation Format
A bounding box is the simplest way to tell a vision system, “the object is here.” A bounding box annotation format is the convention used to record that rectangle’s location in an image, along with the object’s class such as car, person, or defect.
What the numbers describe
Every image has a coordinate grid: the origin is usually the top-left corner, x increases to the right, and y increases downward. A box can be expressed in several equivalent ways:
- Corner coordinates:
(x_min, y_min, x_max, y_max), the top-left and bottom-right corners. - Top-left plus size:
(x, y, width, height). - Centre plus size:
(x_center, y_center, width, height).
Common formats in practice
The Pascal VOC format commonly stores corner coordinates in XML files. COCO annotations use [x, y, width, height] in JSON, measured in pixels. YOLO label files use class_id x_center y_center width height, but the last four values are normalized: each is divided by the image width or height, placing it between 0 and 1. For a 1000-pixel-wide image, an object center at x = 500 becomes 0.5. This makes labels portable across resized versions of the same image.
Why consistency matters
An object detector learns from these numbers and is judged by how closely its predicted box overlaps the annotated one, commonly using intersection over union (IoU). Mixing formats without conversion can place boxes in completely wrong locations: a YOLO-normalized value of 0.5 interpreted as pixels is nearly at the image edge rather than the center. This directly harms training, evaluation, and tasks such as locating vehicles for autonomous driving, lesions in medical scans, or faulty parts on a production line. Libraries such as Albumentations require the box format to be declared so image transformations can update annotations correctly.
Bounding box annotation format is a standardized way to label an object’s location in an image using a rectangular box, typically encoded as corner coordinates (xmin, ymin, xmax, ymax) or as center position plus width and height. It defines how training labels are stored and interpreted by object-detection systems. Consistent formats are essential for correctly training, evaluating, and exchanging detection datasets and model predictions.
Imagine giving someone directions to draw a rectangle around a cat in a photo: “Start here, end there.” A bounding box annotation format is the agreed-upon way of writing down those directions so a computer can understand them.
It records a box’s position and size, usually along with a label such as “cat,” “car,” or “person.” Different formats may describe the same box in slightly different ways, like using the top-left corner and width/height, or two opposite corners. This consistency matters because AI systems learn from many labeled examples. Clear box annotations help them learn not just what an object looks like, but where it appears in an image.