Class Token Pooling
Imagine an image split into many small tiles, with each tile becoming a piece of information a model can inspect. Class token pooling gives the model one extra, learnable “summary slot” whose job is to gather what all those image tiles collectively mean.
How the summary is formed
In a standard Vision Transformer (ViT), an image is divided into fixed-size patches, such as 16×16 pixels. Each patch is converted into a patch token, and a special learned vector called the class token, written as [CLS], is added to the beginning of the sequence. All tokens then pass through transformer encoder layers together. Through self-attention, the class token can attend to any patch token: it learns which regions matter and combines their information into its own representation. After the final layer, the model uses the final class-token vector as a pooled, image-level feature and sends it to a small classification head.
Why call it pooling?
The term “pooling” reflects the class token’s role: it turns many local patch representations into one global representation. Unlike fixed operations such as average pooling or max pooling, class token pooling is learned and content-dependent. For a photo classified as “golden retriever,” the class token can place attention on the dog’s face, fur, and body while downplaying grass or sky.
Why it matters in vision tasks
A single global representation is useful when the task needs one answer for an entire image, including:
- Image classification, such as identifying a defect-free product on a production line.
- Face recognition, where an image is converted into a compact identity-related feature vector.
- Medical image classification, such as predicting whether a scan shows a particular condition.
For tasks requiring an answer at each location—such as object detection or semantic segmentation—the patch tokens are also essential, because they preserve spatial detail. The original ViT architecture uses the final class token for classification, while alternatives such as global average pooling average all patch tokens instead.
Class token pooling is the Vision Transformer method of using a learned class token to aggregate information from all image patch tokens through self-attention. After the transformer encoder, the final class-token embedding serves as a single global image representation for classification. It enables ViTs to convert a variable set of patch features into a fixed-size vector that downstream classifiers can use.
Imagine an image being examined by a group of people, each looking at one small piece. At the end, they need one spokesperson to collect the important ideas and give a single verdict: “This looks like a bicycle.” Class token pooling gives a Vision Transformer that spokesperson.
The model keeps one special placeholder that represents the whole image. After the image’s pieces have been considered together, this placeholder becomes a compact “big picture” summary. That summary can be used to classify an image, compare similar images, or learn useful visual patterns even in unsupervised learning, where no human-provided labels are available.