about
Can Large Language Models Write Parallel Code? (arxiv.org)
1 point by danielnichols on Jan 29, 2024 | hide | past | pdf | discuss on HN

In plain words: Tested several large language models on 420 scientific computing tasks needing code split across many processors, across 12 problem types and six parallel coding styles. Rather than only checking if the code looks right, the study measures how fast and correctly the generated programs run.

Abstract

Large language models are increasingly becoming a popular tool for software development. Their ability to model and generate source code has been demonstrated in a variety of contexts, including code completion, summarization, translation, and lookup. However, they often struggle to generate code for complex programs. In this paper, we study the capabilities of state-of-the-art language models to generate parallel code. In order to evaluate language models, we create a benchmark, ParEval, consisting of prompts that represent 420 different coding tasks related to scientific and parallel computing. We use ParEval to evaluate the effectiveness of several state-of-the-art open- and closed-source language models on these tasks. We introduce novel metrics for evaluating the performance of generated code, and use them to explore how well each large language model performs for 12 different computational problem types and six different parallel programming models.

Daniel Nichols, Joshua H. Davis, Zhaojun Xie, Arjun Rajaram, Abhinav Bhatele
arXiv:2401.12554 · cs.DC, cs.AI · submitted Jan 23, 2024 · updated May 14, 2024
abstract · pdf · html

add comment on HN