Notes

ResNet

ResNet is a landmark image-recognition network built around a simple insight: a very deep model should not have to relearn everything from scratch at every layer. Instead, it can preserve useful visual information and learn only what needs to change.

How residual learning works

A conventional convolutional network passes an image representation through layer after layer. As networks became deeper, they became harder to train: gradients weakened during backpropagation, and extra layers could even reduce accuracy. A ResNet, short for Residual Network, addresses this with a skip connection (or shortcut). Rather than producing an entirely new output, a block learns a residual adjustment:

output = F(input) + input

Here, F(input) is the change learned by a few convolutional layers, while the original input travels along the shortcut. It is like revising a draft instead of rewriting the whole document. If no change is useful, the block can learn a near-zero residual and preserve the input.

Why depth becomes practical

The shortcut gives gradients a direct route backward through the network, making optimization more stable. This enabled successful models with 50, 101, or 152 layers, such as ResNet-50. When feature dimensions differ, the shortcut uses a learned 1×1 convolution to align them before addition.

  • ResNet-18 and ResNet-34 use simpler residual blocks.
  • ResNet-50, 101, and 152 use compact “bottleneck” blocks to control computation.
Where it is used

ResNet is widely used as a pretrained visual backbone in PyTorch’s torchvision.models.resnet50. Its learned features support object detection in street scenes, tumor segmentation in scans, face-related image analysis, and defect inspection on production lines. Even when a system’s final task is detection or segmentation rather than classification, ResNet helps turn raw pixels into increasingly meaningful shapes, textures, and object-level patterns.

ResNet (Residual Network) is a convolutional neural-network architecture that uses residual connections, or skip connections, to pass information around layers. These connections let networks learn residual changes rather than complete transformations, enabling reliable training at great depth. ResNet established a foundation for high-accuracy image classification and remains widely used as a backbone for object detection, segmentation, and transfer learning.

Imagine learning to recognize faces by adding one small helpful observation at a time: “the eyes are round,” then “there is a smile,” then “the hair is curly.” ResNet, short for “Residual Network,” helps an AI learn images in a similar step-by-step way.

It is a design for image-recognition AI that lets very deep models keep learning without losing track of earlier, useful clues. Rather than forcing every new layer of the AI to relearn everything, ResNet allows it to carry forward what it already knows and focus on what is new.

This made it practical to train much larger, more accurate vision systems for tasks such as identifying objects, reading medical scans, and helping cars understand the road.