Notes

Semantic Segmentation

Semantic segmentation gives a computer vision system a detailed map of an image: rather than saying “there is a car,” it marks which individual pixels belong to the car, road, sky, person, building, or another category. Think of it as coloring every pixel with its meaning.

How it works
A semantic-segmentation model treats an image as a dense classification problem. For every pixel location, it predicts a class label from a fixed set, such as road, sidewalk, vehicle, and background. The result is a mask with the same width and height as the input image. During training, the model compares its predicted mask against a hand-annotated ground-truth mask and adjusts its parameters to reduce pixel-level errors.

Models and practical examples
Modern systems commonly use an encoder-decoder design: the encoder extracts useful visual patterns, while the decoder restores fine spatial detail so boundaries align with objects. Architectures such as U-Net, DeepLab, and Fully Convolutional Networks (FCNs) are widely used.

  • An autonomous vehicle separates drivable road from curbs, pedestrians, lane markings, and other cars.
  • In medical imaging, a U-Net can outline a tumor or organ in an MRI scan, helping measure its size and location.
  • On a production line, a model can mark scratches, missing coating, or contaminated areas across a product surface.

Why pixel-level meaning matters
Image classification answers what is present; object detection adds bounding boxes; semantic segmentation provides the precise shape and extent of each class. This precision matters when a decision depends on boundaries, such as calculating a tumor’s area or determining where a vehicle can safely drive. Unlike instance segmentation, semantic segmentation does not distinguish two separate objects of the same class: every car pixel receives the label “car,” regardless of which car it belongs to. Performance is commonly measured with Intersection over Union (IoU), which checks how much a predicted region overlaps the correct one.

Semantic segmentation assigns a class label to every pixel in an image, grouping pixels into categories such as road, sky, vehicle, or person without distinguishing separate objects of the same class. It produces a dense, pixel-level scene map rather than a single image label or bounding box. Semantic segmentation enables precise scene understanding for applications such as autonomous driving, medical image analysis, and robotic navigation.

Imagine coloring a street photo with a different crayon for every kind of thing: blue for sky, gray for road, green for trees, and red for cars. Semantic segmentation does this for an AI system. Instead of merely saying “there is a car in this image,” it labels every tiny part of the picture according to what it represents.

This matters when knowing the exact shape and location of areas is important. A self-driving car needs to tell road from sidewalk and pedestrian from background. Doctors can use it to highlight organs or possible tumors in scans. It helps machines see images less like a single scene and more like a detailed, labeled map.