In plain words: Plans tiny language models with 10 to 100 million parameters that read bytes instead of a word-chopping step, pool nearby bytes, and share weights between input and output layers. The aim is far fewer parameters without losing performance, though results are not yet reported.
Abstract · Super Tiny Language Models
The rapid advancement of large language models (LLMs) has led to significant improvements in natural language processing but also poses challenges due to their high computational and energy demands. This paper introduces a series of research efforts focused on Super Tiny Language Models (STLMs), which aim to deliver high performance with significantly reduced parameter counts. We explore innovative techniques such as byte-level tokenization with a pooling mechanism, weight tying, and efficient training strategies. These methods aim to significantly reduce reduce the parameter count compared to traditional models -- in future works, we aim to build on these in a way that maintains and improves upon the performance of base transformer models. This series of papers will explore into various subproblems, including tokenizer-free models, self-play based training, and alternative training objectives. We will target models with 10M, 50M, and 100M parameters. Our ultimate goal is to make high-performance language models more accessible and practical for a wide range of applications.
Dylan Hillier, Leon Guertler, Cheston Tan, Palaash Agrawal, Chen Ruirui, Bobby Cheng
arXiv:2405.14159 · cs.CL, cs.AI · submitted May 23, 2024 · updated Jun 26, 2024
abstract · pdf · html · 11 pages, 4 figures
https://github.com/leonguertler/supertinylanguagemodels?tab=...
I liked this paper for attempting to make LM’s small enough to pretrain on commodity hardware. I’d like to see more work in this area. My questions are:
1. What real-world uses exist for both tiny and small models?
2. What benchmarks can we reduce in some way where tiny models can make progress on them? As in, can we design tiny models as a proxy for assessing performance of larger models?
3. Can we design tiny models in a way where we can experiment on architecture, hyperparameters, optimization algorithms, etc.? And where it tells us something useful for applying those to larger models?
4. If tiny models are useful, can we convert them to digital or analog hardware to run on cheap, low-power ASIC’s? Or just FPGA’s?