Methods for enabling machines to understand visual data — images and video. Spans foundational image concepts, classical hand-engineered feature techniques, convolutional and transformer-based architectures, the canonical task families (recognition, detection, segmentation), and the specialised sub-fields of video understanding and 3D vision, together with preprocessing and transfer-learning practices that make models work on real datasets.
Foundations & Concepts
Before any algorithm can interpret an image, we need a clear picture of what an image actually is to a computer and how it is described, stored, and labelled. This section lays that groundwork — the raw building blocks of digital pictures, the properties that define their quality and shape, and the conventions used to mark up regions and frames for machine learning.
At the smallest scale, every digital image is a grid of dots. A single Pixel is one such dot, holding a colour or brightness value, and stacking these values turns a picture into an Image as Tensor — the multi-dimensional array that deep-learning frameworks operate on. The depth of that array is set by its Color Channels, such as the red, green, and blue planes of a colour image, while the Bit Depth fixes how many distinct intensity levels each channel can record.
Several properties describe an image as a whole. Its Image Resolution is the number of pixels across its width and height, and the Aspect Ratio is the proportion between those two dimensions. How the pixel data is encoded and compressed on disk is its Image Format. To locate anything within the picture we use Spatial Coordinates, the row-and-column addressing of pixels, which lets us name a specific Region of Interest (ROI) to focus computation on, or pass a Sliding Window across the image to examine every location in turn.
Moving images add a time dimension. A Video Frame is a single still picture pulled from a sequence, and the Frame Rate sets how many such frames appear each second, governing how smoothly motion is captured.
Finally, supervised vision depends on telling the model what is in each image. Image Annotation is the general practice of attaching such labels. For detection, a Bounding Box Annotation Format records the rectangle around each object in an agreed coordinate convention, while for segmentation a Ground Truth Mask marks the exact pixels belonging to each region — the reference answer a model is trained to reproduce.
Image Preprocessing
Raw images rarely arrive in the form a model expects. Preprocessing reshapes, recolours, and varies them so that training is stable and the network sees consistent, diverse input. This section groups those operations into three families: geometric changes to an image's shape and framing, pixel-level changes to its colour and intensity, and augmentation strategies that deliberately introduce variety to improve generalisation.
Spatial / Geometric Transforms
Geometric transforms change where pixels sit rather than their colour. Image Resize rescales a picture to the fixed dimensions a network requires, while Center Crop takes a consistent central patch — both common steps in turning varied inputs into a uniform batch. For variety, Random Crop takes a patch from a random location, and flipping mirrors the image either left-to-right with a Horizontal Flip or top-to-bottom with a Vertical Flip. Random Rotation turns the image by a random angle, helping a model become robust to orientation.
Pixel / Colour Transforms
Colour transforms operate on pixel values. The Color Space (RGB, HSV) determines how colour is encoded, and the Channel Order (BGR vs RGB) is a practical convention that must match what a library expects or colours will be swapped. Grayscale Conversion collapses colour to a single intensity channel, and Histogram Equalization redistributes intensities to improve contrast.
Most networks also expect inputs on a standard numeric scale. Image Normalization brings pixel values into a consistent range, typically by Mean Subtraction, which centres the data by removing the average, and Pixel Scaling, which divides values down to a small fixed interval. Together these keep the input distribution stable so training converges reliably.
Augmentation Strategies
Augmentation invents new training examples by perturbing existing ones, reducing overfitting. Color Jitter randomly nudges brightness, contrast, and saturation, and Gaussian Blur (Augmentation) softens detail to simulate focus and resolution changes. Cutout masks out random patches, forcing the model not to rely on any single region.
Rather than tuning these by hand, learned policies automate the choice. AutoAugment searches for an effective combination of augmentations for a given dataset, and RandAugment offers a simpler, cheaper alternative that applies randomly chosen operations with minimal tuning while retaining most of the benefit.
Classical CV
Before deep learning dominated vision, images were understood through hand-designed algorithms that detect edges, describe distinctive points, match and track them across pictures, and aggregate evidence across scales. These classical techniques remain foundational — fast, interpretable, and still widely used. This section follows the pipeline from low-level gradients up to higher-level matching and aggregation.
Edge & Gradient Operators
Edges — places where brightness changes sharply — are among the most informative cues in an image, and they are found through gradients. The Image Gradient measures the rate and direction of intensity change at each pixel. Practical operators estimate it with small filters: the Sobel Operator and the Prewitt Operator both approximate horizontal and vertical gradients, differing in their weighting. The Laplacian of Gaussian (LoG) smooths first and then highlights regions of rapid change, locating edges where its response crosses zero. The Canny Edge Detection method combines smoothing, gradient computation, and careful thresholding into a robust multi-step detector that remains a standard.
Feature Detectors & Descriptors
Beyond edges, vision needs distinctive, repeatable points that can be recognised across different images. Corner Detection (Harris) finds corners — points where intensity changes in multiple directions — and the FAST Corner Detector does so very quickly for real-time use. To describe the appearance around such points, the Histogram of Oriented Gradients (HOG) summarises local gradient directions, a representation famous for pedestrian detection.
Richer descriptors aim for invariance to scale and rotation. The Scale-Invariant Feature Transform (SIFT) detects and describes keypoints stable across scale and orientation, and Speeded-Up Robust Features (SURF) offers a faster approximation of the same idea. Compact binary descriptors made matching cheaper still: the BRIEF Descriptor encodes a patch as a short binary string, and Oriented FAST and Rotated BRIEF (ORB) combines fast detection with a rotation-aware binary descriptor as an efficient, license-free alternative. For faces specifically, Haar cascades use simple rectangular features in a staged classifier — the core of the classic Viola-Jones Detector, the first method to do real-time face detection.
Matching & Tracking
Once features are described they can be compared and followed. Template Matching slides a reference patch over an image to find where it best fits, and Keypoint Matching pairs descriptors between two images to find corresponding points. Across video, Optical Flow estimates the apparent motion of pixels between frames, and the Lucas-Kanade Method is a classic technique for computing it for tracked points under a local-smoothness assumption.
Multi-Scale & Aggregation
Objects appear at many sizes, so classical vision reasons across scales and pools local evidence. Scale-Space represents an image at a continuum of smoothing levels so features can be found regardless of size, and an Image Pyramid realises this as a stack of progressively downsized copies. The Hough Transform aggregates votes from many pixels to detect global shapes such as lines and circles, and the Bag of Visual Words summarises an image as a histogram of recurring local features, enabling classification and retrieval.
CNN Architectures
The convolutional neural network reshaped computer vision, and its history is told through a succession of landmark architectures. Each generation introduced an idea that solved a limitation of the last — going deeper, computing more cleverly, easing the training of very deep networks, and finally doing all of this efficiently. This section walks that lineage from the foundational designs to the efficient modern families.
Foundational Architectures
The story begins with LeNet, an early convolutional network for recognising handwritten digits that established the basic pattern of convolution and pooling layers. The breakthrough came with AlexNet, which scaled this idea up, won a major image-recognition contest, and ignited the deep-learning era in vision. VGG Net showed that stacking many small uniform filters into a deep, regular network worked remarkably well, and Network in Network introduced the idea of tiny per-pixel sub-networks and global pooling that influenced many later designs.
Inception Family
The Inception line pursued efficiency through smarter layer design. Its key idea is the Inception Module, which runs filters of several sizes in parallel and combines them, capturing patterns at multiple scales at once. GoogLeNet (Inception) assembled these modules into a deep but economical network, and Inception-v3 refined the design with factorised convolutions and other optimisations. Xception pushed the principle to its limit by replacing the module with depthwise-separable convolutions for greater efficiency.
ResNet & Dense Family
As networks grew deeper they became hard to train, and this family solved that. The Residual Block adds a shortcut connection that lets a layer learn only a small change to its input, making gradients flow cleanly through great depth. ResNet stacked these blocks into networks hundreds of layers deep, a watershed in vision. ResNeXt extended the idea with parallel transformation paths, and DenseNet connected every layer to all subsequent ones, maximising feature reuse and gradient flow.
Efficient Architectures
The final theme is doing more with less, for phones and embedded devices. SqueezeNet matched earlier accuracy with far fewer parameters. MobileNet used depthwise-separable convolutions for mobile efficiency, and MobileNetV2 improved it with inverted residual structures. ShuffleNet cut cost further with grouped convolutions and channel shuffling. EfficientNet introduced a principled way to scale depth, width, and resolution together, while NASNet let an automated search design the architecture itself. Bringing the lineage up to date, ConvNeXt modernised the pure convolutional network with design choices borrowed from transformers, showing CNNs remain competitive.
Recognition Tasks
Recognition is the family of computer-vision problems that answer the question "what is in this image?" The tasks range from assigning a single label to a whole picture, through finding and identifying people, to reading text — each defined by what kind of answer it produces. This section introduces them in roughly increasing specificity, from whole-image labels to fine, structured outputs.
The most basic task is Image Classification, which assigns one category to an entire image. When a picture can belong to several categories at once, the task becomes Multi-Label Image Classification. Distinguishing between very similar subcategories — breeds of dog or models of car — is Fine-grained Recognition, which demands attention to subtle detail. Recognising the overall setting or environment of a photo is Scene Classification. A closely related capability is Image Retrieval, which searches a collection for pictures visually similar to a query rather than naming a fixed class.
A large sub-family concerns people and faces. Face Detection simply locates where faces appear, while Face Recognition goes further to determine whose face it is. Pinpointing specific facial points such as the eyes and mouth corners is Facial Landmark Detection. More generally, Keypoint Detection finds notable points on any object, and when those points are the joints of a body, the task becomes Pose Estimation, recovering how a person is positioned. Matching the same individual across different cameras or moments is Person Re-identification.
Finally, recognising and transcribing text within images is Optical Character Recognition (OCR), turning pictures of words into machine-readable characters.
Object Detection
Object detection finds and identifies multiple objects in an image, drawing a box around each and labelling it. It combines "what" and "where" into one task, and its methods split into the shared concepts every detector relies on, the accurate two-stage approaches, and the fast single-stage and transformer-based detectors. This section follows that progression.
Detection Concepts
Every detector predicts a Bounding Box — the rectangle enclosing an object. Many methods start from a set of reference rectangles called an Anchor Box, or generate candidate areas through Region Proposal. To judge how well a predicted box matches the truth, Intersection over Union (IoU) measures their overlap, and because detectors emit many overlapping guesses, Non-Maximum Suppression (NMS) prunes them to one box per object.
Performance is scored with precision-based metrics: Average Precision (AP) summarises accuracy for a single class across detection thresholds, and Mean Average Precision (mAP) averages this over all classes as the field's standard measure. To detect objects across sizes, the Feature Pyramid Network (FPN) combines features from multiple network depths so both small and large objects are well represented.
Two-Stage Detectors
Two-stage detectors first propose regions and then classify them, trading speed for accuracy. R-CNN pioneered the approach by running a classifier on externally proposed regions. Fast R-CNN sped this up by sharing computation across proposals, and Faster R-CNN folded proposal generation into the network itself for near end-to-end training. Cascade R-CNN chained several refinement stages at increasing quality thresholds to produce more precise boxes.
Single-Stage & Modern Detectors
Single-stage detectors predict boxes and labels in one pass, prioritising speed. SSD (Single Shot Detector) detects objects at multiple feature scales in a single network, and the YOLO (You Only Look Once) family reframed detection as one fast regression over a grid. The line evolved rapidly through YOLOv3, YOLOv5, YOLOv7, and YOLOv8, each improving accuracy and speed. RetinaNet made single-stage detection as accurate as two-stage methods by introducing a loss that focuses training on hard examples.
A newer paradigm removes hand-designed components entirely. DETR (Detection Transformer) casts detection as a direct set-prediction problem solved by a transformer, dispensing with anchors and suppression. Deformable DETR refined it with sparse attention over key locations, greatly speeding training and improving detection of small objects.
Segmentation
Segmentation is the most detailed recognition task: instead of a box, it labels every pixel, carving an image into meaningful regions. This section covers the different segmentation tasks and what each one outputs, the architectures that produce dense pixel labels, and the losses and metrics used to train and evaluate them.
Segmentation Tasks
The tasks differ in how finely they distinguish things. Semantic Segmentation assigns every pixel a class but does not separate individual objects of the same class; at its core it is Pixel-wise Classification, a category decision made independently for each pixel. Instance Segmentation goes further by telling apart separate objects of the same class, and Panoptic Segmentation unifies both views, labelling every pixel while also distinguishing each countable object.
Segmentation Architectures
Producing a full-resolution label map requires special network designs. The Fully Convolutional Network (FCN) replaced a classifier's dense layers with convolutions so it could output a spatial map, founding the field. U-Net added skip connections between an encoder and decoder to recover fine detail, becoming a standard especially in medical imaging, and SegNet used a similar encoder-decoder design with efficient upsampling.
To capture context at multiple scales, the DeepLab family introduced dilated convolutions and the Atrous Spatial Pyramid Pooling (ASPP) module, which probes a feature map at several dilation rates at once; the related Pyramid Pooling Module aggregates context over several region sizes. For instance-level results, Mask R-CNN extended a detector to also predict a mask for each detected object. Most recently, the Segment Anything Model (SAM) introduced a promptable, general-purpose segmenter that can isolate almost any object from a simple cue.
Segmentation Losses & Metrics
Evaluating dense predictions needs region-aware measures. Pixel Accuracy reports the fraction of correctly labelled pixels but can be misleading when classes are imbalanced, so Mean Intersection over Union (mIoU) — averaging overlap between predicted and true regions across classes — is the preferred metric. Training uses losses tailored to overlap: Dice Loss directly optimises region similarity and copes well with small foreground regions, and IoU Loss optimises the overlap measure itself.
Transfer Learning
Training a vision model from scratch needs enormous data and compute, but most practitioners never have to. Transfer learning reuses knowledge already captured by a model trained on a large dataset and adapts it to a new task. This section explains the components of that workflow — where the reusable knowledge lives, how to graft it onto a new problem, and the techniques for adapting it carefully.
The reusable knowledge sits in a Pretrained Model — a network already trained on a big general dataset, most commonly through ImageNet Pretraining, which teaches it broadly useful visual features. The feature-extracting body of such a network is its Backbone, the part that turns raw pixels into rich representations that later task-specific layers consume.
There are two broad ways to reuse a backbone. The lightest is Frozen Feature Extraction, where the pretrained weights are held fixed and only a new output layer is trained on top; holding weights fixed in this way is called Layer Freezing. The more powerful approach is Backbone Fine-tuning, which continues training the pretrained weights on the new data so they specialise to the new task.
Fine-tuning is most effective when done gently. Progressive Unfreezing thaws the network gradually, starting from the layers nearest the output and working back, so early general-purpose features are disturbed last. Pairing this with Discriminative Learning Rates — smaller updates for early layers and larger ones for later layers — protects the broadly useful features while letting task-specific ones adapt quickly. When the new task is detection, the adapted backbone feeds a task-specific Detection Head that produces boxes and labels.
Two related ideas extend the theme. Domain Adaptation addresses the case where the new data looks different from the training data — a different camera, lighting, or style — and adjusts the model to bridge that gap. Knowledge Distillation transfers knowledge in another direction, training a smaller, faster "student" model to imitate a large "teacher", preserving much of its ability at a fraction of the cost.
Vision Transformers
Transformers, originally built for language, now rival and often surpass convolutional networks in vision. They treat an image as a sequence and learn relationships between its parts through attention. This section introduces the core vision-transformer design, the variants that made it practical and powerful, and the self-supervised methods that let such models learn from unlabelled images.
Core ViT
The central model is the Vision Transformer (ViT), which applies a standard transformer directly to images. It first splits the picture into Image Patches — small fixed-size tiles treated like words in a sentence — and converts each into a vector through Patch Embedding. A special learnable CLS Token is prepended to the sequence to gather global information, and reading out its final state through Class Token Pooling yields the representation used for classification.
ViT Variants
Plain ViTs need very large datasets, and these variants address that and more. DeiT (Data-efficient Image Transformer) trains strong vision transformers on modest data using distillation. The Swin Transformer introduced a hierarchical design with Shifted Windows, computing attention within local windows that shift between layers to balance efficiency and global reach. A Hybrid CNN-Transformer combines convolutional feature extraction with attention to get the strengths of both, and BEiT brought masked-token pretraining, in the style of language models, to images.
Self-Supervised Visual Learning
Self-supervised methods learn useful features without labels. Contrastive approaches pull together different views of the same image and push apart different images: SimCLR did this with heavy augmentation and large batches, MoCo maintained a running dictionary of negatives for efficiency, and BYOL showed strong learning was possible without negative pairs at all. DINO (Self-Supervised) trained vision transformers with a self-distillation scheme that produced strikingly meaningful attention, and DINOv2 scaled this into a general-purpose visual feature extractor. Masked Autoencoder (MAE) learns by hiding most patches and reconstructing them, a simple and highly scalable objective.
A parallel line connects images with text. CLIP trained on huge image-caption pairs to align pictures and language in a shared space, enabling zero-shot recognition; ALIGN showed the same idea scales with noisy web data, and OpenCLIP provides open reproductions of these vision-language models.
Video Understanding
Video adds time to vision: a model must reason not just about what appears in a frame but about how things move and change across many frames. This section introduces the tasks that define video understanding and the architectural ideas that let networks capture motion and temporal structure.
The headline tasks parallel their image counterparts but over sequences. Video Classification assigns a single label to a whole clip, and the closely related Action Recognition identifies the activity being performed. Because clips contain far more frames than a model can process at once, Frame Sampling selects a representative subset to keep computation manageable.
To model how frames relate over time, networks use temporal building blocks. Temporal Convolution slides filters along the time axis to detect short motion patterns, while Temporal Pooling summarises information across many frames into a compact clip-level representation. Early deep approaches such as C3D extended ordinary convolutions into three dimensions to learn directly from stacks of frames. I3D (Inflated 3D ConvNet) reused successful image architectures by "inflating" their two-dimensional filters into three, inheriting strong pretrained features.
Other designs split the problem by what they model. A Two-Stream Network processes appearance and motion in separate pathways and fuses them, and SlowFast Networks use two pathways running at different frame rates to capture both slow context and fast movement efficiently.
Beyond classification, video understanding includes following objects through time. Multi-Object Tracking (MOT) keeps consistent identities for many moving objects across frames, and Video Object Segmentation delineates target objects at the pixel level throughout a clip.
3D Vision
3D vision recovers the three-dimensional structure of the world from two-dimensional images. Where ordinary recognition asks what is in a picture, 3D vision asks where things are in space and what shape they have. This section moves from estimating depth, through reconstructing scenes and tracking a camera's motion, to the data structures and models used to represent and learn 3D shape.
The first step toward 3D is recovering distance. Depth Estimation predicts how far each pixel is from the camera. A classic way to obtain it is Stereo Vision, which compares two views taken from slightly different positions; the per-pixel shift between those views is captured in a Disparity Map, from which depth follows directly.
From many images we can rebuild whole scenes and locate the camera. Structure-from-Motion (SfM) reconstructs 3D geometry from a set of overlapping photos taken from different viewpoints. Tracking the camera's own movement frame by frame is Visual Odometry, and building a map of an unknown environment while simultaneously keeping track of one's location within it is SLAM.
Three-dimensional information needs suitable representations. A Point Cloud is a set of points in space sampling an object's surface, a Voxel Grid divides space into a regular array of small cubes, and a 3D Mesh describes a surface as connected vertices and faces. To learn directly from irregular point sets, PointNet introduced a network that consumes raw point clouds while respecting their unordered nature.
Finally, generating new 3D structure is the goal of 3D Reconstruction, which builds explicit models of objects and scenes. A powerful modern approach, Neural Radiance Fields (NeRF), represents a scene implicitly as a continuous function learned from images, enabling photorealistic rendering from entirely new viewpoints.