Notes

Knowledge Distillation

Knowledge distillation is a way to pass the useful judgment of a large, capable vision model into a smaller model that is cheaper and faster to run. Rather than training the smaller model only on correct answers, it learns from the larger model’s richer view of what an image could contain.

How the transfer works
A trained teacher model produces a score for every possible class. For a photo of a golden retriever, the teacher might assign high confidence to “golden retriever,” but also some probability to “Labrador” and “dog.” These relative scores contain information about visual similarity that a single ground-truth label hides. The smaller student model is trained to match both:

  • the true label, through ordinary classification loss; and
  • the teacher’s output distribution, through a distillation loss.

A temperature parameter softens the teacher’s probabilities so the student can see meaningful differences among non-winning classes. Training combines the label loss and distillation loss, usually with a weighting factor that controls how strongly the teacher influences the student.

More than class predictions
In computer vision, distillation can transfer more than final labels. A student object detector can imitate a teacher’s bounding-box predictions; a segmentation model can copy pixel-level probability maps; and a face-recognition model can learn the teacher’s feature embeddings. Some methods also align internal feature maps or attention patterns, helping the student learn where the teacher looks in an image.

Why it matters in practice
A large detector might identify defects accurately on a production line but be too slow for an edge camera. Distillation can produce a compact student that preserves much of that accuracy while meeting latency, memory, and power limits. It is also useful when labeled medical scans or specialized inspection images are scarce: a strong teacher supplies informative training signals beyond the limited annotations. Distillation does not guarantee improvement—the student still needs enough capacity, and a biased or inaccurate teacher can pass along its mistakes—but it is a central route from high-performing research models to deployable vision systems.

Knowledge distillation trains a smaller student model to reproduce the outputs or intermediate representations of a larger, high-performing teacher model. Instead of learning only from hard labels, the student uses the teacher’s soft probability predictions, which encode class similarities and confidence. It enables efficient vision models for deployment on resource-constrained devices while retaining much of the teacher’s accuracy.

Imagine a master chef teaching an apprentice. The master may have years of experience and a huge recipe book, but they can pass along the most important habits and shortcuts so the apprentice can cook well too. Knowledge distillation does something similar for AI.

A large, powerful AI model acts as the “teacher,” and a smaller, faster model is the “student.” The student learns to copy the teacher’s useful judgments—for example, recognizing whether a photo shows a dog, a bicycle, or a pedestrian. This matters because the smaller model can often run on phones, cameras, or cars where the large model would be too slow or costly.