about
Comparative Study of Large Language Model Architectures on Frontier (arxiv.org)
1 point by mfiguiere on Mar 18, 2024 | hide | past | pdf | discuss on HN

In plain words: Two popular GPT designs were trained on the same materials science text with identical steps on the Frontier supercomputer, so their differences could be measured fairly. The resulting models set the best score yet on a hard materials science test.

Abstract

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world's first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.

Junqi Yin, Avishek Bose, Guojing Cong, Isaac Lyngaas, Quentin Anthony
arXiv:2402.00691 · cs.DC · submitted Feb 1, 2024
abstract · pdf · html

add comment on HN