Notes

BEiT

Imagine covering several squares of a photograph and asking a model to infer what belongs behind them from the surrounding scene. BEiT teaches a vision model this kind of visual “fill in the blanks” skill before asking it to recognize objects, segment medical scans, or detect defects.

How BEiT learns from images
BEiT, short for Bidirectional Encoder representation from Image Transformers, is a self-supervised pretraining method for a Vision Transformer (ViT). An image is divided into fixed-size patches, such as 16×16-pixel squares, and each patch becomes a token processed by Transformer attention layers. During pretraining, BEiT hides a portion of those patch tokens. The model must predict a discrete visual token for every hidden patch using context from the visible patches.

Why predict visual tokens instead of pixels?
Rather than reconstructing exact red, green, and blue pixel values, original BEiT uses a separately trained image tokenizer—based on a discrete variational autoencoder—to assign each image patch a code representing its visual content. This resembles BERT in language processing: BERT predicts missing words or word pieces; BEiT predicts missing image codes. Predicting meaningful codes pushes the model to understand structures such as object parts, texture, boundaries, and scene layout, rather than merely copying nearby colors.

Why it matters in vision
After this broad pretraining on unlabeled images, the BEiT encoder is fine-tuned for a specific task with far less task-specific labeling than training from scratch. Its learned representations help with:

  • Image classification, such as identifying diseases in retinal images.
  • Object detection, such as locating pedestrians and vehicles in driving footage.
  • Semantic segmentation, such as outlining tumors or separating road, sky, and buildings.

BEiT helped establish masked image modeling as a powerful route to training visual Transformers, particularly when abundant unlabeled images are available but carefully annotated data is expensive.

BEiT (Bidirectional Encoder representation from Image Transformers) is a self-supervised vision-transformer pretraining method that masks image patches and trains the model to predict their discrete visual-token representations. By learning contextual relationships among patches without manual labels, BEiT produces transferable visual features that improve performance on image classification, object detection, and semantic segmentation after fine-tuning.

Imagine learning a language by reading a book with some words covered up, then guessing what belongs in the blanks from the surrounding sentence. BEiT applies a similar idea to pictures. It hides small parts of an image and trains an AI to understand what those missing parts are likely to represent.

Rather than needing a person to label every photo “cat,” “car,” or “tree,” BEiT can learn from huge collections of unlabeled images. This is called self-supervised learning: the image provides its own practice question. That early visual understanding helps the AI later recognize objects, scenes, and details more accurately.