Register and [CLS] tokens yield a decoupling of local and global features in large ViTs

Alexander Lappe; Martin A. Giese

arXiv:2505.05892·cs.CV·October 27, 2025

Register and [CLS] tokens yield a decoupling of local and global features in large ViTs

Alexander Lappe, Martin A. Giese

PDF

Open Access

TL;DR

This paper investigates how register and [CLS] tokens in large Vision Transformers affect the relationship between local and global features, revealing a decoupling that impacts interpretability and suggesting ways to improve model transparency.

Contribution

The study uncovers that register and [CLS] tokens cause a decoupling of local and global features in large ViTs, affecting attention map interpretability and proposing insights for more transparent models.

Findings

01

Register tokens produce cleaner attention maps.

02

Global information is dominated by register tokens.

03

[CLS] token causes similar decoupling in models without explicit register tokens.

Abstract

Recent work has shown that the attention maps of the widely popular DINOv2 model exhibit artifacts, which hurt both model interpretability and performance on dense image tasks. These artifacts emerge due to the model repurposing patch tokens with redundant local information for the storage of global image information. To address this problem, additional register tokens have been incorporated in which the model can store such information instead. We carefully examine the influence of these register tokens on the relationship between global and local image features, showing that while register tokens yield cleaner attention maps, these maps do not accurately reflect the integration of local image information in large models. Instead, global information is dominated by information extracted from register tokens, leading to a disconnect between local and global features. Inspired by these…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Explainable Artificial Intelligence (XAI) · Multimodal Machine Learning Applications

MethodsSoftmax · Attention Is All You Need