about
Transformers without normalization (arxiv.org)
42 points by kaycebasques on Jul 24, 2025 | hide | past | pdf | 6 comments on HN

In plain words: Neural networks usually include layers that rescale each layer's numbers to keep training stable. This work swaps those for a simple stretch-and-squash curve on every value, and Transformers using it match or beat the usual ones across vision and language tasks, mostly without tuning.

Abstract · Transformers without Normalization

Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation $DyT($x$) = \tanh(α$x$)$, as a drop-in replacement for normalization layers in Transformers. DyT is inspired by the observation that layer normalization in Transformers often produces tanh-like, $S$-shaped input-output mappings. By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning. We validate the effectiveness of Transformers with DyT across diverse settings, ranging from recognition to generation, supervised to self-supervised learning, and computer vision to language models. These findings challenge the conventional understanding that normalization layers are indispensable in modern neural networks, and offer new insights into their role in deep networks.

Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu
arXiv:2503.10622 · cs.LG, cs.AI, cs.CL, cs.CV · submitted Mar 13, 2025 · updated Jun 14, 2025
abstract · pdf · html · CVPR 2025; Project page: https://jiachenzhu.github.io/DyT/

add comment on HN
Also discussed: Mar 2025 (2 points, 1 comment) · Mar 2025 (2 points, 0 comments) · Mar 2025 (4 points, 0 comments)

I think other than the title being a bit misleading, the paper is good. I say misleading because they replace Layer Normalization with a tanh function, which still bounds the range to [-1,1]. Plenty of people would call that normalization (an unfortunately overloaded term).

While the result isn't too surprising it has a good ablation study and helps build confidence in the mechanism. It's simple and quick to implement, but I don't find that a disadvantage. Arguably this is not novel, but sometimes it is worth revisiting things when the rest of the environment has changed and I think the study being thorough makes it useful to the community.

The project page is here[0] which will give you a very quick understanding of the paper.

[0] https://jiachenzhu.github.io/DyT/

> (an unfortunately overloaded term)

I mentioned normalization in an interview, and they had no idea what I was talking about given my context, they were thinking of database normalization, I was thinking of DATA normalization, where you uppercase all inputs for e.g. an email, so when they login, casing doesn't matter, since you'll uppercase it when you check against the database. I'm sure there's a zillion other normalization methods for different things.

I've always thought that normalization, as defined in the statistical sense, needs to be a linear transformation to preserve the shape of the distribution. tanh is definitely not normalization from that point of view. Even so, they could have been more specific and called it 'linear normalization'.
I never liked the conventional normalization, this tanh looks like it should execute faster
Depends on your context and goals.

LayerNorm isn't going to bound you strictly into [-1,1] like this will. So that can have some advantages. A strict bounding can sometimes get you in trouble as it may not be as robust to novel inputs. For a basic example if you consider the "classic" normalization where you rescale your training data so that it is bound on [0,1] this does not mean that data from your test set will be in [0,1]. Does your model know how to generalize this?

A scheme like this has the potential to land you in similar trouble. Think about the domain and range. If your training data is all in [-100, 100] then you might get a pretty wide tanh to accommodate that. But will the resultant filter be able to differentiate the value 100 from 1000? Probably not. The filter is going to optimize to the data it saw through training. Will there be some filter that has that capacity? Maybe. But also there are ways to process your data where this bound really wouldn't matter.

We're getting into the weeds here but I'm just trying to illustrate why there are so many different normalization schemes. There's no one-size-fits all process and it is best to understand where certain methods have advantages and disadvantages (there's definitely advantages in that numbers closer to the origin have higher precision as the density of addressable values in fp{16,32,64} are not evenly distributed)

Discussion (260 points, 4 months ago, 32 comments) https://news.ycombinator.com/item?id=43369633