Notes

CLIP

Imagine teaching a vision system by showing it an image of a dog alongside the phrase “a photo of a dog,” rather than assigning it a fixed class number. CLIP, short for Contrastive Language–Image Pre-training, learns to connect pictures with the natural-language descriptions that fit them. This gives it a broad visual vocabulary without training a separate model for every possible category.

How CLIP learns image–text meaning
CLIP is trained on a very large collection of image-and-caption pairs gathered from the web. It has two encoders: one converts an image into a vector of numbers, and the other converts text into a vector in the same shared space. During training, the model is shown a batch of matching and non-matching pairs. Its contrastive learning objective pulls the vector for each real image-caption pair closer together and pushes incorrect pairings apart.

Classification through prompts
Rather than requiring a classifier trained specifically for “cat,” “car,” or “defective part,” CLIP can compare an image against text prompts such as:

  • “a photo of a pedestrian”
  • “a chest X-ray showing pneumonia”
  • “a scratched product on a production line”

The prompt whose text vector is most similar to the image vector becomes the prediction. This is called zero-shot classification: recognizing a category without seeing task-specific labeled examples during final training. Prompt wording matters; “a photo of a golden retriever” and “a golden retriever” can produce different results.

Why it matters in vision
CLIP makes visual systems more flexible: it supports image search using ordinary language, helps label or organize large photo collections, and supplies useful features for detection and segmentation systems. It also underpins tools such as OpenCLIP and text-guided image generation. Its broad web training brings limitations: captions can be incomplete or biased, and CLIP’s similarity score is not reliable proof of what is truly present in an image—especially in high-stakes settings such as medicine or driving.

CLIP (Contrastive Language–Image Pre-training) is a multimodal model trained to align images with their natural-language descriptions by contrasting matched and unmatched image–text pairs. It learns transferable visual representations from large-scale web data and can classify images using text prompts without task-specific training. CLIP enables zero-shot recognition, text-based image retrieval, and language-guided visual applications.

Imagine teaching someone what things look like by showing them millions of photos with their captions: a dog beside a bicycle, a sunset over the ocean, a red shoe. CLIP is an AI system that learns to connect images with the words that describe them.

Rather than needing people to carefully label every photo with fixed categories, CLIP learns from naturally paired images and text found online. This is close in spirit to self-supervised learning: it draws useful lessons from the data already available. That makes CLIP flexible—it can often recognize a new idea simply from a written description, such as “a watercolor painting of a city,” even without being specially trained for that exact label.