Notes

Mask R-CNN

Imagine a photo containing several people, cars, and bicycles. A detector can draw a box around each object; Mask R-CNN goes further by tracing the exact pixels belonging to every individual object, even when objects of the same type overlap.

How it produces object masks
Mask R-CNN is an instance segmentation model: it identifies each object instance, assigns it a class, locates it with a bounding box, and creates a separate pixel mask for it. It extends Faster R-CNN, a two-stage object detector. First, a backbone network such as ResNet extracts visual features. A Region Proposal Network (RPN) then suggests image regions likely to contain objects. For each proposed region, the model predicts:

  • the object class, such as “person” or “car”;
  • a refined bounding box; and
  • a small binary mask marking the object’s foreground pixels.

The key detail: RoIAlign
A crucial part of Mask R-CNN is RoIAlign. Earlier detectors rounded region coordinates while extracting features, which caused slight misalignments. That is tolerable for a bounding box but damaging when predicting an edge pixel by pixel. RoIAlign samples feature values at precise fractional locations instead, keeping the predicted mask aligned with the original object. The mask branch works in parallel with classification and box prediction, so the model learns both “what is this?” and “which pixels belong to this particular one?”

Why it matters in practice
Mask R-CNN is useful wherever object boundaries carry meaning: separating overlapping cells in a microscope image, identifying each vehicle and pedestrian for autonomous-driving perception, isolating products for visual inspection, or selecting a person precisely for photo editing. A plain detector cannot distinguish object pixels from nearby background within its box; semantic segmentation can label “car” pixels but does not inherently separate two touching cars. Mask R-CNN supplies both identity and boundaries. It is implemented in widely used frameworks such as Detectron2 and torchvision.models.detection.maskrcnn_resnet50_fpn.

Mask R-CNN is a two-stage neural network for instance segmentation that extends Faster R-CNN by predicting a pixel-level mask for every detected object, alongside its class label and bounding box. Its RoIAlign operation preserves spatial alignment for accurate masks. Mask R-CNN enables systems to distinguish and precisely outline separate objects of the same class, supporting applications such as medical-image analysis, robotics, and autonomous driving.

Imagine cutting out every person, car, and dog from a photo with scissors, giving each one its own separate paper shape. Mask R-CNN helps AI do something similar: it finds individual objects in an image, names them, and carefully marks the exact pixels belonging to each one.

This matters when objects overlap. Rather than saying “there are three people here,” it can separate each person’s outline, even in a crowded scene. That makes it useful for tasks such as counting items in a shop, spotting cells in medical images, or helping robots understand what they can safely pick up. It is usually trained with human-labelled examples, not by unsupervised learning, where patterns are found without labels.