about
Evaluating AI-Generated Code for C++, Fortran, Go, Java, Julia, Matlab, etc. (arxiv.org)
1 point by Bostonian on May 29, 2024 | hide | past | pdf | 2 comments on HN

In plain words: ChatGPT 3.5 and 4 were asked to write three scientific programs in nine languages, then checked for compiling, speed, and correct answers. Both produced working code with some help, but some languages were easier than others and parallel code was hard to get right.

Abstract · Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust

This study evaluates the capabilities of ChatGPT versions 3.5 and 4 in generating code across a diverse range of programming languages. Our objective is to assess the effectiveness of these AI models for generating scientific programs. To this end, we asked ChatGPT to generate three distinct codes: a simple numerical integration, a conjugate gradient solver, and a parallel 1D stencil-based heat equation solver. The focus of our analysis was on the compilation, runtime performance, and accuracy of the codes. While both versions of ChatGPT successfully created codes that compiled and ran (with some help), some languages were easier for the AI to use than others (possibly because of the size of the training sets used). Parallel codes -- even the simple example we chose to study here -- also difficult for the AI to generate correctly.

Patrick Diehl, Noujoud Nader, Steve Brandt, Hartmut Kaiser
arXiv:2405.13101 · cs.SE, cs.AI · submitted May 21, 2024 · updated Jul 5, 2024
abstract · pdf · html · 9 pages, 3 figures

add comment on HN

The insight here is something like C++ might be more amenable to Co-Pilot help due to the sheer volume of C++ code out there. In essence, this might turn out to make C++ a lot more usable with AI Co-Pilots than something like Python (which has a lot of code but poor quality) or Rust (not that much code).
The title was truncated to fit -- etc. stands for Python, R, and Rust.