In plain words: Numbers are usually chopped into digit pieces, so models handle them awkwardly. This approach turns each number into one slot sized by its value, letting scientific data be trained as text; it generalized better to unfamiliar numbers and ran more efficiently than other encodings.
Abstract · xVal: A Continuous Numerical Tokenization for Scientific Language Models
Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.
Siavash Golkar, Mariel Pettee, Michael Eickenberg, Alberto Bietti, Miles Cranmer, Geraud Krawezik, Francois Lanusse, Michael McCabe, Ruben Ohana, Liam Parker, Bruno Régaldo-Saint Blancard, Tiberiu Tesileanu, et al.
arXiv:2310.02989 · stat.ML, cs.AI, cs.CL, cs.LG · submitted Oct 4, 2023 · updated Dec 15, 2024
abstract · pdf · html · 15 pages, 12 figures. Appendix: 8 pages, 2 figures. Accepted contribution at the NeurIPS Workshop on ML for the Physical Sciences
The big idea is to take any corpus and instead of making tokenizations for numbers that are digits (GPT-2 era) or weirdo floating point range things, instead every number goes to a single token: [NUM]. They then keep a shadow tensor/vector where the [NUM] is given an actual floating point (maybe fixed point?) number.
When the model predicts a [NUM] token, there's a number prediction layer that chooses a number. This lets them train against predicting whether or not a number will be next in a generated text (the model turns out to be really good at this, not surprising). And, how cool -- their loss function can check the actual number created by the number layer, and give a high quality result as to how good the number guess was.
This works really really well out of the box for a bunch of math problems up to 5 digit multiplication, like super human levels of accuracy, and vastly beats SOTA from other Foundation models.
They then go on to throw fine tuning tasks of scientific data at it, and it absolutely does not shit the bed. Which is extremely profound to me, and something they do not make a big deal of. I would expect some scaled up models using this tech could be generally highly useful for a broad range of numeric / stats / prediction in the same way that GPT-3 and on have been highly useful as "text calculators".
Worth a read.