about
An AI system to help scientists write expert-level empirical software (arxiv.org)
8 points by simonpure on Sep 9, 2025 | hide | past | pdf | 1 comment on HN

In plain words: An AI system writes scientific software by trying code changes and keeping the ones that score highest on a quality measure, searching widely instead of writing one draft. It produced 40 new single-cell analysis methods that beat the best human-made ones on a public leaderboard.

Abstract

The cycle of scientific discovery is frequently bottlenecked by the slow, manual creation of software to support computational experiments\cite{hannay2009how}. To address this, we present Empirical Research Assistance (ERA), an AI system that creates expert-level scientific software whose goal is to maximize a quality metric. The system uses a Large Language Model (LLM) and Tree Search (TS)\cite{silver2016mastering} to systematically improve the quality metric and intelligently navigate the large space of possible solutions. ERA achieves expert-level results when it explores and integrates complex research ideas from external sources. The effectiveness of tree search is demonstrated across a diverse range of tasks. In bioinformatics, ERA discovered 40 novel methods for single-cell data analysis that outperformed the top human-developed methods on a public leaderboard. In epidemiology, ERA generated 14 models that outperformed the CDC ensemble and all other individual models for forecasting COVID-19 hospitalizations. ERA also produced expert-level software for geospatial analysis, neural activity prediction in zebrafish, and numerical solution of integrals, and a novel rule-based construction for time series forecasting. By devising and implementing novel solutions to diverse tasks, ERA represents a significant step towards accelerating scientific progress.

Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Jake Garrison, Renee Johnston Anton Kast, Cory Y. McLean, Peter Norgaard, Zahra Shamsi, David Smalling, James Thompson, et al.
arXiv:2509.06503 · cs.AI, q-bio.QM · submitted Sep 8, 2025 · updated May 21, 2026
abstract · pdf · html · 78 pages, 31 figures, 22 tables

add comment on HN
Also discussed: May 2026 (2 points, 0 comments) · Sep 2025 (3 points, 1 comment)

I think an interesting, yet unsurprising, result is that when asked to make a general predictor the ultimate result of all of this is that gradient boosting is still a winner.

> The discovered solutions showed strong convergence towards gradient boosting

and figure 17 which shows that about 2/3 of all solutions arrive at using some kind of gradient boosting or ensemble method.