Notes

Vision Transformer (ViT)

A Vision Transformer (ViT) looks at an image less like a grid of nearby pixels and more like a sequence of meaningful pieces. It divides the image into small patches, then lets every patch compare itself with every other patch to build a picture-wide understanding.

How it turns an image into tokens

A ViT begins by splitting an image into fixed-size patches, such as 16×16 pixels. Each patch is flattened and converted into a numeric vector called a token embedding. Because transformers do not inherently know where a token came from, ViT adds positional embeddings that tell it each patch’s location. A special learnable class token is placed at the start of the sequence; after processing, this token gathers information from the whole image and is used for classification.

How it learns visual relationships

The token sequence passes through stacked transformer encoder layers. Their key mechanism, self-attention, lets a patch weigh information from every other patch. For example, a patch containing a car wheel can connect with distant patches showing windows and headlights, helping the model recognize a car even when parts are separated or partly hidden.

  • For photo classification, the class token can decide whether an image contains a dog, aircraft, or product defect.
  • For medical scans, patch relationships help identify the full extent of a tumor rather than treating each local area independently.
  • For object detection and segmentation, ViT-derived models provide rich image features that later components turn into boxes or pixel-level masks.
Why ViT matters

Traditional convolutional networks build understanding from local neighborhoods outward. ViT can capture long-range relationships directly, which is valuable for large objects, scene context, and fine visual structure. Its trade-off is that self-attention becomes expensive as image resolution increases, and the original ViT benefits greatly from large-scale pretraining. In practice, implementations such as torchvision.models.VisionTransformer and Hugging Face’s ViT models use pretrained weights, then adapt them to a smaller task-specific dataset.

Vision Transformer (ViT) is an image-recognition architecture that divides an image into fixed-size patches, converts them into token embeddings, and processes them with a transformer encoder using self-attention and positional information. A classification token aggregates image-level features. ViTs capture long-range relationships across an image and underpin many modern models for classification, detection, and segmentation.

Imagine looking at a large mural by first noticing many small sections, then connecting them to understand the whole scene. A Vision Transformer (ViT) is an AI model that does something similar with images. Rather than focusing only on nearby details, it can relate a bird’s wing in one area to the tree branches far away in another.

This matters because meaning in an image often depends on distant parts working together: a person holding an umbrella, a car on a road, or a face partly covered by sunglasses. ViTs can also learn from huge collections of unlabeled pictures, helping AI gain useful visual understanding even when humans have not named every image.