Notes

Discriminative Learning Rates

A pretrained vision model has already learned a great deal: early layers detect edges and textures, while later layers recognize parts and objects. Discriminative learning rates fine-tune these parts at different speeds instead of forcing every layer to change equally.

How it works

During fine-tuning, a model is divided into parameter groups, such as the classification head, late backbone blocks, and early backbone blocks. Each group receives its own learning rate. The newly added task-specific head gets the largest rate because it starts with random weights and must learn quickly. Deeper pretrained layers use a medium rate, while the earliest layers use the smallest rate because their general visual features are already useful and easy to damage.

early layers:  0.00001
late layers:   0.0001
new head:      0.001
Why unequal rates help

Think of it as renovating a house: the new room needs substantial work, but the foundations should be adjusted carefully. A single large learning rate can cause catastrophic forgetting, where fine-tuning overwrites valuable knowledge gained from large-scale pretraining. A single tiny rate, on the other hand, leaves the new prediction head unable to adapt. Discriminative rates balance adaptation with preservation.

Where this matters

This approach is especially useful when labeled data is limited:

  • Adapting an ImageNet-pretrained model to identify product defects on a production line.
  • Training a medical-image classifier where subtle tissue patterns differ from ordinary photographs.
  • Fine-tuning an object detector for a new camera, lighting condition, or set of vehicle classes.

Libraries such as fastai expose this directly through layer-group learning rates, while PyTorch users create optimizer parameter groups manually. Combined with gradual unfreezing—training the head first, then progressively opening backbone layers—discriminative learning rates make transfer learning more stable, data-efficient, and less likely to erase useful visual knowledge.

Discriminative learning rates assign different update speeds to different parts of a pretrained model during fine-tuning, typically using a smaller rate for early backbone layers and a larger rate for newly added task-specific layers. This preserves broadly useful visual features while allowing higher-level representations to adapt rapidly to the target dataset. They improve transfer-learning stability and reduce catastrophic forgetting of pretrained knowledge.

Imagine renovating a house: you might carefully preserve the sturdy foundation while making bigger changes to the kitchen and paintwork. Discriminative learning rates do something similar when adapting an AI model to a new job.

An existing vision model already knows useful basics, such as edges, shapes, and textures. Those early skills should change slowly. But the later parts, which make task-specific decisions, may need to adapt quickly—for example, learning to spot defects on factory products instead of recognizing everyday objects.

Using different learning speeds for different parts helps the model keep valuable old knowledge while becoming good at its new task.