Notes

LeNet

LeNet is one of the early neural networks that showed computers could learn to read images rather than relying entirely on hand-written rules. It became famous for recognizing handwritten digits—the kind found on bank checks and postal codes—and helped establish the design ideas behind modern image-recognition systems.

How LeNet reads an image
The best-known version, LeNet-5, takes a small grayscale image, such as a 32×32 pixel digit, and processes it through a sequence of learned layers. Its key idea is the convolutional layer: small learned filters slide across the image looking for useful visual patterns. Early filters respond to simple marks such as edges, curves, and corners; later layers combine these into digit-like parts and complete shapes.

The main building blocks

  • Convolution layers create feature maps that preserve where visual patterns occur.
  • Pooling layers shrink those maps, reducing computation and making the network less sensitive to a feature shifting by a few pixels.
  • Fully connected layers use the extracted features to choose a class, such as “3” rather than “8.”

Why it still matters
LeNet is small by current standards, but its pattern—convolutions, spatial downsampling, then classification—became the foundation for later CNNs such as AlexNet and VGG. In practice, a closely related model can classify simple grayscale images, inspect characters in scanned documents, or recognize marks in controlled manufacturing images. Modern libraries make its structure easy to build; for example, PyTorch provides Conv2d, AvgPool2d, and Linear layers for the same core pipeline. Without learned convolutional features, a system would need manually designed rules for every edge, stroke, and variation in appearance—rules that break quickly on messy real images.

LeNet is an early convolutional neural network (CNN), introduced by Yann LeCun and colleagues, designed for handwritten digit recognition. Its layered convolutions, pooling, and fully connected classifiers learned visual features directly from pixel data. LeNet demonstrated that CNNs could achieve reliable image classification and established core architectural principles used by later computer-vision models.

Think of LeNet as one of the early “reading machines” that helped computers learn to recognize what they see. Its famous job was reading handwritten digits, such as the numbers written on bank checks or postal codes. Instead of being told every possible way someone might write a 7, it learned useful visual clues from many examples.

LeNet showed that a computer could look at an image in small pieces, notice simple patterns like edges and curves, and combine them into a guess: “this is probably a 3.” It was a major early step toward today’s image-recognition AI, from phone photo search to systems that read documents automatically.