Notes

Shifted Windows

Imagine dividing an image into small neighborhoods and letting each patch look only at nearby patches. This keeps attention efficient, but it can trap information inside those neighborhood boundaries. Shifted windows solve that problem by moving the neighborhood layout between transformer layers, allowing visual information to cross boundaries without paying the cost of global attention.

How the shift works

Shifted windows are a central idea in the Swin Transformer. In one layer, the image’s patch tokens are partitioned into fixed-size windows, such as 7 × 7 patches. Self-attention is computed independently inside each window, so a patch compares itself with only 48 nearby patches rather than every patch in the image. In the next layer, the window grid is shifted—typically by half a window width and height. Patches that were separated by a previous window border now appear together in a new window.

  • Regular window attention captures relationships within local regions.
  • Shifted window attention connects adjacent regions across the former borders.
  • An attention mask prevents patches artificially brought together by the image-edge wraparound from attending to one another.
Why this matters for images

Objects rarely align neatly with a fixed grid. A car can span several regions, a face can cross a window boundary, and the edge of a tumor can extend through many local areas. Without shifting, each window would develop a limited, isolated view. Repeated shifted-window layers let information gradually travel across the image while keeping computation manageable. This makes Swin-style models useful as backbones for object detection, semantic segmentation, and visual inspection—not only image classification.

Efficiency with broader context

Full self-attention grows rapidly as image resolution increases because every patch attends to every other patch. Windowed attention limits that growth, and shifting restores cross-window communication that purely local windows would lose. Libraries such as torchvision provide Swin Transformer implementations, where shifted-window attention is built into the model blocks rather than added as a separate image preprocessing step.

Shifted windows are a Swin Transformer attention scheme in which the window partition is offset between successive layers. Self-attention remains local and computationally efficient within each window, while the shift lets tokens interact across previous window boundaries. This enables hierarchical vision transformers to capture broader spatial relationships without the cost of global attention, improving performance on dense tasks such as object detection and semantic segmentation.

Imagine a crowd at a festival split into small groups. People can chat easily within their own group, but not with people in the next one. Now imagine the group boundaries move sideways: suddenly, each new group contains some people who were previously separated. Information can spread across the whole crowd.

Shifted windows use this idea in image AI. A model examines an image in small regions, or “windows,” to keep the task manageable. Then, in the next stage, it shifts those regions slightly. This lets details near one window’s edge connect with nearby details in another, helping the model understand larger objects and scenes without needing to examine everything at once.