about
Vision Transformers Need Registers (arxiv.org)
94 points by felineflock on Apr 28, 2025 | hide | past | pdf | 9 comments on HN

In plain words: Vision Transformers park high-energy spots in blank backgrounds as scratch space for internal math; adding a few input tokens gives them a place for that work. This removes the artifacts entirely and makes feature and attention maps smoother than in models without these tokens.

Abstract

Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.

Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski
arXiv:2309.16588 · cs.CV · submitted Sep 28, 2023 · updated Apr 12, 2024
abstract · pdf · html

add comment on HN
Also discussed: Oct 2023 (22 points, 0 comments)

Extremely cool!

It's interesting and honestly encouraging that this kind of thing can be discovered and understood using just "simple linear methods" and high-level analysis of patterns in layer activations.

So basically multiple CLS tokens.

Fwiw, I tried multiple global tokens in my chess neural net and didn't see any uplift compared to my baseline of just having one.

Note that it's not done for performance reason but rather to generate clear feature maps.
Has this been used widely since?
I ran a comparison of DINOv2 with and without registers on some image embedding tasks for work; DINOv2+registers saw a performance metric bump of 2-3%. Not nothing, not transformative, worth using when the only difference for inference is the model name string you're loading.
I can't speak for the whole industry, but we used it in older UForm <https://github.com/unum-cloud/uform> and saw good adoption, especially among those deploying on the Edge, where every little trick counts. It's hard to pin down exact numbers since most deployments didn't go through Hugging Face, but at the time, these models were likely among the more widely deployed by device count.
yes