Notes

Image Patches

Rather than looking at every pixel as a separate piece of information, a Vision Transformer breaks an image into small, regular tiles called image patches. Think of it as cutting a photograph into a grid of tiny squares, then giving the model one compact description for each square.

How patches become transformer input
In the original Vision Transformer (ViT), an image is divided into non-overlapping patches, such as 16 × 16 pixels. A 224 × 224 image therefore becomes a 14 × 14 grid: 196 patches. Each patch contains its red, green, and blue pixel values, which are flattened into one long vector and passed through a learned linear layer. This creates a patch embedding, also called a visual token.

  • Patch tokens carry the visual content of each small image region.
  • Positional embeddings tell the transformer where each patch came from, since transformers do not inherently know left from right or top from bottom.
  • A special class token gathers information from all patches for image-level decisions, such as “this image contains a bicycle.”

Why patch size matters
Patch size controls a useful trade-off. Smaller patches preserve finer detail—helpful for tiny defects on a production line, small road signs in autonomous driving, or subtle boundaries in medical scans—but they create more tokens and make attention computation more expensive. Larger patches reduce cost, but can hide small objects or blur precise edges. In ViT, patch creation is commonly implemented as a convolution with kernel size and stride equal to the patch size; PyTorch’s Conv2d can perform this efficiently. By turning an image into a sequence of meaningful regions, patches let transformer attention connect distant parts of a scene: a face, its eyes, and a hat can influence one another even when they are far apart in the image.

Image patches are fixed-size, non-overlapping regions into which an image is divided before processing. In a Vision Transformer, each patch is flattened and projected into a vector token, allowing the model to apply transformer attention to image content. Patch size determines the visual detail available to the model: smaller patches preserve finer structure but increase computational cost. They enable transformers to process images as token sequences.

Imagine looking at a large jigsaw puzzle by picking up one small piece at a time. Image patches are like those small pieces: an image is divided into many tiny square regions so an AI can examine each region separately.

In a Vision Transformer, each patch becomes a small “note” describing what appears there—perhaps part of a face, a wheel, or a patch of sky. The AI then considers how these notes relate across the whole picture. This helps it understand both nearby details and far-apart connections, such as a person holding an umbrella. Patches are useful in both labeled training and unsupervised learning, where AI learns visual patterns without being told the answers.