CLS Token
A CLS token is a small learned piece of data placed at the front of an image’s token sequence. Think of it as a designated “meeting point” where a Vision Transformer gathers what it has learned about the whole image before making a classification decision.
How it works
A Vision Transformer (ViT) splits an image into fixed-size patches, such as 16×16-pixel squares. Each patch becomes a numeric patch token, and positional embeddings tell the model where each patch came from. Before these patch tokens, ViT inserts one extra learned vector: the CLS token, short for “classification token.”
- At the beginning, the CLS token contains no image-specific information.
- Inside each transformer encoder layer, self-attention lets it attend to every image patch.
- Across many layers, it collects evidence about shapes, textures, object parts, and their relationships.
- The final CLS-token representation is passed to a small classification head, which might output labels such as “cat,” “car,” or “defective part.”
Why this is useful
The CLS token gives the model one compact vector intended to represent the image globally. For photo classification, this is convenient: instead of combining hundreds of patch outputs manually, the classifier reads one designated output. In face recognition, a related global representation can encode the identity-relevant features of a face. In industrial visual inspection, it can help produce a decision such as “acceptable” or “scratch detected.”
Important distinction
A CLS token is most natural for image-level classification. Tasks requiring a prediction at each location—such as medical-image segmentation or object detection—need detailed patch-level features too. Modern architectures may use all patch tokens, pooling methods such as global average pooling, or several task-specific query tokens instead of relying only on one CLS token. In the original ViT implementation, the CLS token is a trainable parameter, updated through learning just like other model weights.
A CLS token is a learned embedding prepended to an image’s patch-token sequence in a Vision Transformer. Through self-attention, it gathers information from all patches; its final representation is used as a global image feature, typically passed to a classification head. It enables a ViT to produce a single prediction for an entire image without explicit pooling.
Imagine giving someone a photo cut into many small puzzle pieces, then placing one extra blank note beside them. As they look over the pieces, they use that note to write down the big-picture answer: “This is a dog” or “This is a street scene.”
In a Vision Transformer, the CLS token is that extra note. It is not part of the image itself. Instead, it becomes a compact summary of what the model thinks the whole image contains. This gives the AI one convenient place to make a final decision, such as choosing an image label. It can also learn useful visual summaries from unlabeled images during self-supervised learning.