about
H2O-Danube-1.8B Technical Report (arxiv.org)
7 points by tosh on Apr 7, 2024 | hide | past | pdf | discuss on HN

In plain words: Built a compact 1.8-billion-parameter language model trained on huge amounts of text using the same recipe as today's leading open models, plus a chat version tuned to follow user preferences. It ranked first among similarly small open models on a public leaderboard.

Abstract

We present H2O-Danube, a series of small 1.8B language models consisting of H2O-Danube-1.8B, trained on 1T tokens, and the incremental improved H2O-Danube2-1.8B trained on an additional 2T tokens. Our models exhibit highly competitive metrics across a multitude of benchmarks and, as of the time of this writing, H2O-Danube2-1.8B achieves the top ranking on Open LLM Leaderboard for all models below the 2B parameter range. The models follow core principles of LLama 2 and Mistral, and we leverage and refine various techniques for pre-training large language models. We additionally release chat models trained with supervised fine-tuning followed by direct preference optimization. We make all models openly available under Apache 2.0 license further democratizing LLMs to a wider audience economically.

Philipp Singer, Pascal Pfeiffer, Yauhen Babakhin, Maximilian Jeblick, Nischay Dhankhar, Gabor Fodor, Sri Satish Ambati
arXiv:2401.16818 · cs.CL, cs.LG · submitted Jan 30, 2024 · updated Apr 15, 2024
abstract · pdf · html

add comment on HN
Also discussed: Apr 2024 (3 points, 0 comments)