Notes

Sliding Window

A sliding window is a simple way to make a computer inspect an image piece by piece rather than trying to understand the whole image at once. Imagine moving a small picture frame across a photograph: every position reveals a different local patch for the system to examine.

How it works
A sliding-window process chooses a rectangular region of a fixed size, such as 64 × 64 pixels, and moves it across the image by a chosen number of pixels called the stride. At each location, a vision algorithm extracts features or runs a classifier to ask a question such as “does this patch contain a face?” or “is there a scratch here?” The process is repeated at several window sizes because an object can appear large when close to the camera and small when far away. This is called an image pyramid or multi-scale search.

From many patches to useful detections
A window-based detector can produce hundreds of overlapping positive results around the same object. Non-maximum suppression keeps the strongest detection and removes nearby duplicates. Classical object detectors, including early face and pedestrian detectors using HOG features with an SVM, relied heavily on this approach. It also appears in practical tasks:

  • Scanning a road image for pedestrians or traffic signs.
  • Checking microscope images for suspicious cells.
  • Finding defects across a high-resolution product photo.
  • Reading characters in a document by examining local regions.

Why it matters today
Sliding windows are intuitive and thorough: they can search every plausible location without needing prior knowledge of where an object is. Their weakness is cost—an image with many positions and scales creates a huge number of patches to evaluate. Modern detectors such as Faster R-CNN, YOLO, and SSD avoid repeatedly classifying nearly identical crops by processing the image once and predicting boxes from shared feature maps. Even so, the sliding-window idea remains foundational: convolutional neural networks effectively apply learned local filters across an image in a closely related, far more efficient way.

A sliding window is a fixed-size region moved systematically across an image or feature map to inspect local content at each position. Each window can be classified, scored, or used to compute features, such as detecting whether it contains an object. It enables localized visual analysis and underlies classical object detection, image scanning, and convolutional operations; without it, exhaustive location-based search is not defined.

Imagine looking for a friend in a huge crowd by covering most of the scene with a small picture frame, then moving that frame little by little until you have checked everywhere. A sliding window does something similar with an image.

It is a small viewing area that moves across a picture, examining one patch at a time. An AI vision system can use it to search for things such as faces, cars, or animals, including objects that may appear in different parts of the image. It matters because a computer image is just a grid of tiny colored dots; the sliding window gives the system a practical way to focus on local areas instead of treating the whole image as one undifferentiated scene.