about
Physics in Next-Token Prediction (arxiv.org)
28 points by Anon84 on Nov 29, 2024 | hide | past | pdf | 4 comments on HN

In plain words: They show that when a language model predicts the next word, information is transferred rather than created, and that training energy is tied to how much information the model can hold. These rules line up with known scaling laws for size, knowledge, and precision.

Abstract · Physics in Next-token Prediction

We discovered the underlying physics in Next-token Prediction (NTP). We identified the law of information conservation within NTP and proposed the First Law of Information Capacity (IC-1), demonstrating that the essence of intelligence emergence in auto-regressive models is fundamentally a process of information transfer. We also introduced Landauer's Principle into NTP, formulating the Second Law of Information Capacity (IC-2), which establishes the relationship between auto-regressive model training and energy consumption. Additionally, we presented several corollaries, which hold practical significance for production practices. Finally, we demonstrate the consistency between the Law of Information Capacity and the Scaling Law for Neural Language Models, the Knowledge Capacity Scaling Laws, and the Scaling Laws for Precision.

Hongjun An, Yiliang Song, Xuelong Li
arXiv:2411.00660 · cs.LG, cs.AI · submitted Nov 1, 2024 · updated Nov 16, 2024
abstract · pdf · html · Second Submit

add comment on HN

Reminds me of the Physics of LM papers (Part 1 here https://arxiv.org/abs/2305.13673)
Was hoping for something slightly stronger and can't help but feel put off by a big over sized box on the first page with some quantities which feel more "derived post-facto" in practice.

Maybe very useful, but still feels more qualitative than quantitative, or maybe I'm missing something (wouldn't shock me)...

Anyone with a physics and a ML/DL background analyze this? Any insights
Just as a preface: This paper is an extremely basic derivation that totally ignores all current architectures and training algorithms. If someone had actually done this for a realistic, modern model, it would be amazing - but that is extremely challenging mathematically and I don't think anyone has come close to cracking that.

So their main proposal makes no assumptions about the model and essentially says that any model "absorbs" information equal to the difference between the training data set's entropy and the final cross entropy loss after training. They state that this information difference must have gone somewhere and thereby it must have been encoded in the model. They argue about this in terms of transmission, but it seems vague enough to be generally applicable. So nothing too crazy here.

But: This feels pretty similar to MOND research in physics, in the sense that it would have been super cool if someone had predicted this before we knew all those laws from empirics. But since it came out post factum, it leaves the bitter taste of someone just trying to mold their world view to the available data.