Segment Anything Model (SAM)
Imagine selecting an object in a photograph with a single click and having its exact outline appear—rather than drawing it pixel by pixel. The Segment Anything Model (SAM) is designed for that kind of flexible image understanding: it creates masks, meaning pixel-level regions, for objects or parts of objects based on simple user prompts.
How SAM produces a mask
SAM separates the work into three cooperating parts. A large Vision Transformer image encoder first turns the entire image into a reusable feature representation. A prompt encoder translates a hint into a form the model can use, and a lightweight mask decoder combines both to predict the selected region.
- Point prompts: click an object to include it, or click a distracting area to exclude it.
- Box prompts: draw a rough rectangle around an object.
- Mask prompts: refine a previous or partially supplied segmentation.
Handling visual ambiguity
A click on a person’s shirt could mean the shirt, the whole person, or a nearby group. SAM can return several plausible masks and assigns each a predicted quality score, helping software choose the most reliable result. It was trained on the SA-1B dataset, containing more than a billion masks, which gives it broad zero-shot ability: it can segment many unfamiliar object categories without being retrained for each one.
Why it matters in practice
SAM makes segmentation a reusable building block rather than a model trained from scratch for every label set. In medical-image workflows, a clinician can provide a few clicks to outline a structure; in production-line inspection, a box can isolate a part before checking for defects; and in photo or video tools, masks support background removal, object editing, and tracking. SAM does not inherently name objects or guarantee medically validated boundaries—it identifies regions matching the prompt—so domain-specific review remains essential when mistakes carry consequences.
Segment Anything Model (SAM) is a promptable foundation model for image segmentation that generates object masks from prompts such as points, bounding boxes, text, or existing masks. Trained on a large-scale dataset of images and annotations, it generalizes across diverse objects and scenes without task-specific retraining. SAM enables rapid interactive annotation, zero-shot segmentation, and mask generation for downstream vision systems.
Imagine outlining an object in a photo with a marker: you point at a dog, and someone instantly traces its exact shape, including its ears and tail. The Segment Anything Model (SAM) does something similar for images.
It can separate a picture into meaningful pieces, such as people, cars, plants, or individual objects. A person can guide it with a click, a rough box, or a short text prompt, and SAM marks the object’s boundaries pixel by pixel. This matters because many AI tasks need to know not just what is in an image, but exactly where it is—like helping photo editors, medical image tools, or robots understand their surroundings.