about
Scalable-Softmax Is Superior for Attention (arxiv.org)
2 points by jw1224 on Feb 4, 2025 | hide | past | pdf | 2 comments on HN

In plain words: Softmax spreads attention flatter as context grows, losing key details; this fix scales its output with the input size to keep focus sharp. Models using it learned faster and retrieved details better in long contexts, and swapping it into a trained model helped.

Abstract

The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to flatten as the context size grows. This reduces the model's ability to prioritize key information effectively and potentially limits its length generalization. To address this problem, we propose Scalable-Softmax (SSMax), which replaces Softmax in scenarios where the input vector size varies. SSMax can be seamlessly integrated into existing Transformer-based architectures. Experimental results in language modeling show that models using SSMax not only achieve faster loss reduction during pretraining but also significantly improve performance in long contexts and key information retrieval. Furthermore, an analysis of attention scores reveals that SSMax enables the model to focus attention on key information even in long contexts. Additionally, although models that use SSMax from the beginning of pretraining achieve better length generalization, those that have already started pretraining can still gain some of this ability by replacing Softmax in the attention layers with SSMax, either during or after pretraining.

Ken M. Nakanishi
arXiv:2501.19399 · cs.CL, cs.AI, cs.LG · submitted Jan 31, 2025
abstract · pdf · html · 11 pages, 8 figures

add comment on HN

Abstract:

> The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to flatten as the context size grows. This reduces the model's ability to prioritize key information effectively and potentially limits its length generalization.

> To address this problem, we propose Scalable-Softmax (SSMax), which replaces Softmax in scenarios where the input vector size varies. SSMax can be seamlessly integrated into existing Transformer-based architectures. Experimental results in language modeling show that models using SSMax not only achieve faster loss reduction during pretraining but also significantly improve performance in long contexts and key information retrieval.

> Furthermore, an analysis of attention scores reveals that SSMax enables the model to focus attention on key information even in long contexts. Additionally, although models that use SSMax from the beginning of pretraining achieve better length generalization, those that have already started pretraining can still gain some of this ability by replacing Softmax in the attention layers with SSMax, either during or after pretraining

There is a hyperparameter `s` in scalable softmax.

SSoftMax_i = exp(s log(n) z_i) / sum (n is length of embedding).

Normal softmax (with temperature)

SoftMax_i = exp(z_i / T) / sum (T is the temperature).

Here Temperature is a hyperparameter. Having a temperature as hyperparameter does not seem too different to me than having `s` as a hyperparameter. I personally don't understand the benefits of SSoftMax. During hyperparameter search you would find the optimal `s` as you might find the optimal temperature `T`.