Notes

Spatial Coordinates

Every visual system needs a way to say where something is. Spatial coordinates provide that shared map: they identify the position of a pixel, object, landmark, or region so a computer can connect visual content to a precise location.

Coordinates inside an image
In a digital image, locations are usually described with two values: x for horizontal position and y for vertical position. The origin, written as (0, 0), is normally the top-left pixel. Moving right increases x; moving down increases y. This differs from the Cartesian graph familiar from mathematics, where y increases upward. Image-processing libraries also commonly describe a pixel as (row, column), which corresponds to (y, x)—a small convention difference that can cause misplaced results when ignored.

How vision systems use them
Spatial coordinates turn model outputs into usable visual decisions. For example:

  • An object detector represents a car with a bounding box, such as (x, y, width, height) or two corner points.
  • Face-recognition systems locate facial landmarks—the eyes, nose, and mouth corners—to align faces before comparison.
  • Medical segmentation assigns a label to each coordinate, marking exactly which pixels belong to a tumor or organ.
  • OCR uses text-region coordinates to identify where words and lines appear on a scanned page.

Coordinates must survive transformations
When an image is resized, cropped, rotated, or flipped, its annotations must be transformed too. Crop a photo without shifting its box coordinates, and the detector is trained to look in the wrong place. Coordinates can be stored in raw pixels or as normalized coordinates between 0 and 1, which makes labels easier to transfer across image sizes. Libraries such as OpenCV use coordinates throughout functions like cv2.rectangle, cv2.warpAffine, and cv2.resize. Spatial coordinates are therefore the bridge between “what is visible” and “where it is,” enabling vision systems to act precisely rather than merely recognize patterns.

Spatial coordinates specify a location in an image, video frame, or 3D scene using numerical axes—for example, pixel positions (x, y) in an image or (x, y, z) in space. They provide the common reference system for representing object positions, bounding boxes, keypoints, segmentation masks, and geometric relationships. Accurate coordinates are essential for detection, tracking, pose estimation, and mapping visual predictions back to the original scene.

Think of a city map: every place has an address that tells you exactly where it is. Spatial coordinates are the “addresses” of points in an image, video, or 3D scene.

In a photo, coordinates can say that a cat’s ear is near a particular spot on the screen, such as “120 pixels from the left and 80 pixels from the top.” In a 3D scan, they can describe where an object sits in real space: left or right, up or down, and near or far.

They matter because vision systems need more than “there is a car.” They need to know where the car is to draw a box around it, help a robot avoid it, or guide a self-driving vehicle safely.