Notes

Image Classification

Image classification is the part of computer vision that answers a simple-looking question: “What is in this picture?” A model examines the image as a whole and assigns it one or more meaningful labels, such as cat, pneumonia, damaged product, or sunset.

How it works

During training, the model is shown many images paired with correct labels. A modern neural network, commonly a convolutional neural network (CNN) or a Vision Transformer (ViT), learns visual patterns that help distinguish categories. Early layers detect simple cues such as edges, colors, and textures; deeper layers combine those cues into richer concepts such as wheels, eyes, leaves, or the overall shape of a vehicle.

  • Single-label classification chooses one class: a photo is either “dog” or “cat.”
  • Multi-label classification can assign several classes: a street image can contain “car,” “pedestrian,” “traffic light,” and “bicycle.”
  • The model returns scores or probabilities for classes, then selects labels using the highest score or a chosen threshold.

What it does—and does not—locate

Classification labels an entire image; it does not identify where each object appears. For example, it can report that a photo contains a dog, but not draw a box around the dog. That requires object detection; coloring every pixel by category requires segmentation. This distinction matters when the location is part of the decision, such as finding a defect on a circuit board or identifying a tumor boundary in a scan.

Why it matters in practice

Image classification powers photo-organizing tools, medical-image triage, wildlife-camera monitoring, and visual inspection on production lines. Its reliability depends heavily on representative labeled training images: a model trained only on clean, well-lit product photos can fail when deployed with shadows, blur, new camera angles, or uncommon defect types. Libraries such as PyTorch and TensorFlow provide pretrained models including ResNet and EfficientNet, which are commonly adapted to a new classification dataset.

Image classification is the task of assigning an image to one or more predefined categories based on its visual content, such as labeling a photo “cat,” “car,” or “pneumonia.” A model produces class probabilities or labels for the entire image rather than locating individual objects. It is a core computer-vision capability that supports content search, medical-image analysis, quality inspection, and many downstream recognition systems.

Think of sorting a box of photos into labeled folders: “cats,” “dogs,” “beaches,” and “birthday parties.” Image classification is the AI version of that job. You give a computer an image, and it decides which label best describes what is in it.

For example, a phone app might classify a photo as a plant, a medical scan as normal or suspicious, or a satellite image as forest, water, or city. It does not need to name every object or show exactly where each one is. Its main job is to answer: “What kind of image is this?” This makes large collections of pictures easier to search, organize, and understand.