Notes

BYOL

Imagine teaching a model that two differently edited versions of the same photo still depict the same underlying thing. BYOL, short for Bootstrap Your Own Latent, does this without needing human labels—or a collection of “negative” images to contrast against.

How it learns from views
BYOL starts with one image and creates two augmented views: for example, different crops, color changes, blur, or flips. One neural network, the online network, processes the first view. A second, slightly delayed copy called the target network, processes the other. Their outputs should represent the same image content, despite the visual edits.

  • The online network contains an encoder, a projection head, and a predictor.
  • The target network contains an encoder and projection head, but no predictor.
  • Training adjusts the online network so its prediction matches the target network’s representation.
  • The target network is not trained directly by gradient descent. Instead, its weights are updated as an exponential moving average of the online network’s weights.

Why this avoids a trivial answer
A naïve system could map every image to the same vector and claim all views match; this is called representation collapse. BYOL’s asymmetric design—the predictor on only one side, the stop-gradient target branch, and the slowly updated target weights—makes that shortcut unstable in practice. The model instead learns features that preserve meaningful image structure: object shape, parts, texture, and scene layout.

Why it matters in vision
After pretraining on large unlabeled image collections, BYOL’s encoder can be fine-tuned for image classification, object detection, semantic segmentation, medical-image analysis, or visual inspection on a factory line. This is valuable when labels are expensive: a model can first learn broad visual knowledge from raw images, then need far fewer annotated examples for the final task. BYOL also differs from methods such as SimCLR because it does not rely on explicitly pushing representations of different images apart.

BYOL (Bootstrap Your Own Latent) is a self-supervised learning method that trains an online image encoder to predict representations produced by a slowly updated target encoder for different augmented views of the same image. Unlike contrastive methods, it does not require negative image pairs. BYOL learns transferable visual features from unlabeled data, improving downstream classification, detection, and segmentation when annotations are limited.

Imagine learning to recognize a friend whether they are wearing sunglasses, standing in shadow, or seen from a different angle. BYOL, short for Bootstrap Your Own Latent, teaches an AI a similar kind of visual common sense.

It learns from ordinary unlabeled images, so nobody needs to tag each photo “cat,” “bicycle,” or “tree.” Instead, it encourages the AI to treat slightly changed versions of the same image as meaning the same thing. A cropped, blurred, or color-shifted photo of a dog should still be understood as that dog. This helps AI build useful visual understanding before it is trained for tasks such as sorting photos or spotting objects.