Notes

Patch Embedding

Images arrive as grids of pixels, while transformers are designed to process sequences of tokens. Patch embedding is the bridge between those two formats: it turns small square regions of an image into the numeric tokens a Vision Transformer can read.

How an image becomes tokens

A Vision Transformer divides an image into fixed-size, non-overlapping patches—for example, a 224×224 image can be split into 16×16 pixel patches. This produces 14×14, or 196, patches. Each patch contains color values for its pixels, so it is first flattened into one long vector. A learned linear projection then converts that vector into an embedding of a chosen size, such as 768 numbers. Each resulting vector is a patch token: a compact learned description of one local image region.

Position is essential

Unlike a convolutional network, a transformer has no built-in sense of where a token came from. Therefore, positional embeddings are added to patch embeddings so the model can distinguish “an eye near the top” from “an eye near the bottom.” The token sequence, commonly with an extra learned class token for image-level classification, then enters transformer encoder layers. Their attention mechanism compares patches across the whole image: a wheel token can relate to another wheel token far away, helping identify a bicycle.

Why it matters in practice

Patch size controls a useful trade-off:

  • Smaller patches preserve finer details, useful for tiny defects in factory inspection or small tumors in medical scans, but create more tokens and higher computation costs.
  • Larger patches are cheaper to process but can lose thin text strokes for OCR or small distant objects in driving scenes.

In PyTorch, the original ViT-style operation is commonly implemented as a Conv2d whose kernel size and stride both equal the patch size. This efficiently extracts and projects every patch in one step, rather than manually flattening image regions.

Patch embedding converts an image into a sequence of fixed-size patches, then maps each flattened patch to a learned vector representation, or token. In a Vision Transformer, these tokens—combined with positional information—are the input processed by transformer layers. Patch embedding makes images compatible with sequence-based attention models, while its patch size directly controls spatial detail, computation cost, and the model’s ability to represent small visual structures.

Imagine trying to describe a large painting to someone by cutting it into small, equally sized tiles. You could talk about each tile one at a time: its colors, edges, textures, and tiny details. Patch embedding is the step that gives an AI this kind of manageable view of an image.

Instead of treating a whole photo as one enormous object, it divides the image into small patches, like a grid of tiles. Each patch is turned into a compact numerical description the model can work with. This lets the AI compare different parts of the picture and notice relationships—such as a wheel belonging to a car, even when the wheel and roof are far apart.