Ground Truth Mask
A ground truth mask is the pixel-by-pixel answer key used to teach and evaluate a vision system that needs to understand image regions, not just whole images. Instead of saying “there is a cat somewhere,” it marks exactly which pixels belong to the cat and which belong to the background.
What the mask containsA mask has the same height and width as its source image. Each pixel stores a label that represents the correct annotation:
- In binary segmentation, a pixel is commonly 1 for the target object and 0 for everything else.
- In semantic segmentation, pixel values identify classes such as road, car, pedestrian, sky, or tumor.
- In instance segmentation, separate masks distinguish individual objects of the same class—for example, each person in a crowd receives its own mask.
The word “ground truth” means the label is treated as the trusted reference, usually created or checked by human annotators. It is not necessarily perfect: fuzzy boundaries, hidden objects, and disagreements between annotators can introduce uncertainty.
How models use itDuring training, a segmentation model produces a predicted mask and compares it with the ground truth mask. The difference becomes a loss, such as cross-entropy loss or Dice loss, which guides the model toward more accurate pixel labels. At evaluation time, measures such as Intersection over Union (IoU) compare the overlap between predicted and ground-truth regions.
Why precise masks matterGround truth masks unlock tasks where object boundaries matter. A medical model can outline a tumor for treatment planning; an autonomous vehicle can separate drivable road from sidewalks and pedestrians; a factory inspection system can isolate a tiny scratch from an otherwise normal product. Poor masks teach poor boundaries: a model trained on rough labels can miss thin structures, merge nearby objects, or learn background artifacts instead of the object itself. Tools such as CVAT and Label Studio help annotators draw, refine, and export these masks for training pipelines.
A ground truth mask is a pixel-level annotation that assigns each image pixel to a target class, object, or background, representing the reference-correct segmentation. It is typically created by human annotators and used as the labeled target during training and evaluation of semantic or instance segmentation models. Accurate masks are essential for learning precise object boundaries and measuring segmentation quality.
Imagine tracing around every cat in a photo with a colored marker, carefully filling in the exact pixels that belong to the cat and leaving everything else blank. That colored outline-and-fill is a ground truth mask.
In computer vision, a mask is like a precise cutout map for an object, person, road, tumor, or any other region in an image. “Ground truth” means it is the trusted answer, usually created or checked by people. AI systems compare their own guesses with these masks to learn what objects look like and to measure whether their predictions are accurate. It matters whenever simply saying “there is a cat” is not enough—you need to know exactly where the cat is.