Notes

Pixel

A digital image is built from tiny picture elements called pixels. Think of a pixel as one small measurement of light and colour at a particular position in an image; millions of these measurements together form the photos, video frames, scans, and camera feeds that a computer-vision system examines.

What a pixel stores
A pixel has a location, commonly written as coordinates (x, y), and one or more numeric values. In a grayscale image, one value represents brightness: lower numbers are darker and higher numbers are lighter. In a standard colour image, a pixel usually has three channels: red, green, and blue (RGB). With 8 bits per channel, each channel ranges from 0 to 255, so a pixel such as (255, 0, 0) is bright red. Libraries such as OpenCV store colour images in BGR order by default, a small but important detail when displaying or processing images.

Pixels as model input
Computer-vision pipelines turn an image into an array of pixel values. A 1920 × 1080 image contains more than two million pixel locations. Models use patterns across neighbouring pixels—not isolated values—to recognize edges, textures, shapes, and objects. For example:

  • An OCR system examines pixel patterns to distinguish “8” from “B”.
  • A medical segmentation model assigns each pixel a label such as tumour, organ, or background.
  • A production-line inspection system finds unusually dark, bright, or irregular pixel regions that signal a defect.

Why pixels matter
Pixel values determine what visual evidence a model can use. Image resolution controls how much spatial detail is available: a tiny distant pedestrian may occupy too few pixels for an autonomous-vehicle detector to recognize reliably. Pixel-level preprocessing—resizing, normalizing values, correcting colour, or reducing noise—also directly affects model predictions. In tasks such as semantic segmentation, the desired output is itself a pixel-by-pixel map, making the pixel both the raw input unit and the unit in which the final visual decision is expressed.

A pixel (picture element) is the smallest addressable unit of a digital image, representing a sampled value at a specific spatial location. Its value encodes brightness in grayscale images or channel intensities such as red, green, and blue in color images. Pixels form the numerical input that vision algorithms analyze; their arrangement and resolution determine the visual detail available for tasks such as detection, segmentation, and recognition.

A pixel is like a tiny colored tile in a giant mosaic. From far away, the tiles blend into a photo of a face, a dog, or a sunset. Zoom in far enough, and you can see that the image is really made of many little squares.

For a computer, an image is a grid of pixels. Each pixel records a color or brightness at one small spot. A phone photo may contain millions of them, which is why it can show fine detail.

Computer vision uses these tiny tiles as its raw visual information. By looking at patterns across many pixels, an AI can learn to recognize things such as roads, people, tumors in scans, or handwritten numbers.