Instance Segmentation
Imagine asking a vision system not just to find “cars” in a street scene, but to trace the exact outline of each individual car, even when several overlap. Instance segmentation gives a separate pixel-level mask to every detectable object, so the system knows which pixels belong to car 1, car 2, and car 3.
How it differs from related tasks
A bounding box from object detection says roughly where an object is. A segmentation mask says precisely which pixels form its visible shape. The key word, instance, means individual object identity within an image.
- Semantic segmentation labels all car pixels simply as “car,” merging neighboring cars into one class region.
- Instance segmentation separates each car into its own mask and identifier.
- Panoptic segmentation extends this idea by labeling individual countable objects while also covering background regions such as road, sky, and grass.
What the output looks like
For each object, a result usually includes a class label, a confidence score, a bounding box, and a binary or probabilistic mask. When two people stand close together, their masks must remain separate. During training, the model compares its predicted masks with hand-annotated ground-truth masks; evaluation commonly uses mask intersection over union (IoU) and average precision metrics.
Why precise object boundaries matter
Instance segmentation unlocks decisions that boxes alone cannot support:
- An autonomous vehicle can distinguish each pedestrian and estimate its visible extent.
- A medical system can measure each separate cell or lesion in a scan.
- A factory inspection system can count, isolate, and flag individual defective parts.
- A photo-editing tool can select one person’s hair or clothing without selecting another person beside them.
Instance segmentation assigns a class label and a separate pixel-accurate mask to every individual object in an image, distinguishing objects even when they share the same category. Unlike semantic segmentation, it separates each car, person, or cell into its own instance. It is essential for applications requiring object-level localization and shape understanding, including robotics, autonomous driving, medical imaging, and image editing.
Imagine putting a different colored outline around every apple in a fruit bowl—even when some apples touch or overlap. Instance segmentation gives AI that ability for images: it identifies each individual object and marks the exact pixels that belong to it.
For example, in a street photo, it can separate one pedestrian from another, label each car individually, and trace each dog rather than simply saying “there are cars, people, and dogs here.” This matters when the identity and shape of every separate object count, such as for self-driving vehicles, medical image analysis, or robots picking items from a shelf.