In plain words: Two AI language models were trained mostly on non-English text, using a word-splitting system built to handle all 24 official EU languages. They performed strongly on European versions of standard knowledge and reasoning tests.
Abstract · Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing a custom multilingual tokenizer, our models address the limitations of existing LLMs that predominantly focus on English or a few high-resource languages. We detail the models' development principles, i.e., data composition, tokenizer optimization, and training methodologies. The models demonstrate strong performance across multilingual benchmarks, as evidenced by their performance on European versions of ARC, HellaSwag, and TruthfulQA.
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, Katrin Klug, Jasper Schulze Buschhoff, et al.
arXiv:2410.03730 · cs.CL, cs.AI, cs.LG · submitted Sep 30, 2024 · updated Aug 21, 2025
abstract · pdf · html