Notes

Swin Transformer

A Swin Transformer is a vision model designed to look at an image in manageable neighborhoods rather than comparing every image patch with every other patch at once. This makes transformer-style image understanding practical for large, detailed images—the kind used in detection, segmentation, and video.

How it looks at an image

Like other vision transformers, Swin first divides an image into small square patches, treating each patch as a token. Its key idea is window-based self-attention: attention is calculated only among patches inside a small local window. This sharply reduces the expensive computation of standard global attention, whose cost rises rapidly as image resolution grows.

Why the windows shift

Fixed windows would isolate information: a patch near one window’s edge could not communicate with a neighboring window. Swin solves this with shifted window attention. One transformer block uses regular windows; the next shifts the window layout by part of a window. Patches then gain connections across the previous boundaries, allowing information to spread through the image while retaining efficient local computation. It is a little like examining a mural through a grid, then sliding the grid so details on opposite sides of a grid line can be considered together.

A hierarchy useful for vision tasks

Swin also progressively merges nearby patches, producing lower-resolution but richer feature maps at several scales. That resembles the multiscale structure used by convolutional networks and is especially valuable when objects have very different sizes. For example:

  • In object detection, it helps find both a distant pedestrian and a large nearby vehicle.
  • In medical image segmentation, it preserves fine local boundaries while incorporating surrounding anatomy.
  • In visual inspection, it can combine tiny surface defects with the larger context of a manufactured part.

Models such as Swin-T, Swin-S, and Swin-B scale this design to different compute budgets. Swin became a widely used backbone because it combines transformers’ flexible attention with the efficient, multiscale features that dense vision tasks require.

Swin Transformer is a hierarchical vision transformer that computes self-attention within local image windows and shifts those windows between layers to connect neighboring regions. This design reduces the quadratic cost of global attention while producing multi-scale feature maps. Swin Transformers are important because they provide efficient, scalable backbones for image classification, object detection, and semantic segmentation.

Imagine looking at a huge city map through a set of small windows. You study one neighbourhood at a time, then shift the windows slightly so you can see how streets and buildings connect across their edges. A Swin Transformer helps an AI examine images in a similar practical way.

It is designed for pictures where both tiny details and the big scene matter: spotting a pedestrian in a street photo, outlining a tumour in a scan, or identifying objects in a crowded room. Rather than treating every part of an image as equally important all at once, it can build an understanding from local areas up to the whole image. This makes it useful for tasks that need precise visual understanding, not just a single label.