Pretrained Model
A pretrained model is a model that has already learned useful visual patterns before you give it your own problem. Instead of starting with a network whose weights are random, you begin with one that has spent substantial training time learning from a large image collection—much like hiring someone who already recognizes edges, textures, shapes, and common objects.
What it has learned
In computer vision, a pretrained network is commonly trained on datasets such as ImageNet, which contains millions of labeled images. Its early layers learn broadly useful features: lines, corners, color transitions, and textures. Deeper layers combine these into parts and object-level patterns, such as wheels, eyes, fur, or text-like strokes. The learned numerical settings, called weights, are saved and reused rather than trained again from scratch.
How it is reused
A pretrained model usually has a backbone, which extracts visual features, and a task-specific head, which turns those features into predictions. For a new task, practitioners replace or retrain the head—for example, changing an ImageNet classifier into one that distinguishes defective from non-defective products. They then choose between:
- Feature extraction: freeze the backbone and train only the new head. This is fast and works well with small datasets.
- Fine-tuning: continue training some or all pretrained layers on the new data. This adapts the model more closely to the target images.
Why it matters in vision
Pretraining sharply reduces the labeled data, computing time, and training effort needed for tasks such as medical-image segmentation, face-related recognition, or object detection in traffic video. Libraries such as PyTorch torchvision provide pretrained ResNet, EfficientNet, and Vision Transformer weights directly. Without pretraining, a small dataset can lead a large vision model to memorize training images rather than learn reliable visual rules. A relevant caveat is domain mismatch: a model trained on everyday photos needs careful fine-tuning when used on X-rays, satellite imagery, or factory-camera images.
A pretrained model is a neural network already trained on a large dataset, such as ImageNet, to learn reusable visual representations. In computer vision, it serves as a starting point for a related task by retaining its learned backbone and adapting or fine-tuning it with new data. Pretrained models reduce data, training time, and compute requirements while improving performance on smaller or specialized datasets.
Think of a pretrained model like a chef who has already learned basic cooking skills before starting a new job. They know how to chop, season, and recognize ingredients, so they do not need to begin from zero.
In AI, a pretrained model has already spent time learning useful patterns from a huge collection of images, such as edges, shapes, textures, faces, and objects. It can then be adapted for a new purpose—for example, spotting damaged products, identifying plant diseases, or sorting photos of pets.
This matters because training an AI from scratch takes lots of images, time, and computing power. A pretrained model gives it a valuable head start.