about
Toward Autonomous Long-Horizon Engineering for ML Research (arxiv.org)
1 point by Anon84 169 days ago | hide | past | pdf | discuss on HN

In plain words: A team of AI agents turns a research goal into a working machine-learning system by sharing one persistent folder, so each agent builds on past decisions. It beat the strongest comparable setups by about 10 points; deleting that folder cut its score by 6.

Abstract

Agentic systems increasingly automate pieces of AI research. Yet turning underspecified research objectives into runnable, experimentally validated ML systems remains a central bottleneck. We study this operational setting as \emph{long-horizon ML research engineering}: converting a research specification into a runnable ML system through repeated implementation, experimentation, and refinement. The central challenge is to sustain cumulative project progress across heterogeneous stages under delayed, confounded feedback. We introduce AiScientist, a multi-agent system built around thin control over thick state: a lightweight hierarchical research team coordinates through a File-as-Bus workspace that preserves decision-relevant artifacts across roles and invocations. On PaperBench, AiScientist improves over the strongest matched baselines by 9.92 and 11.15 points with Gemini-3-Flash and GLM-5, respectively. On MLE-Bench Lite, it reaches 81.82 Any Medal\% under both backbones, improving over the strongest matched baselines by 4.55 and 16.67 points, and exceeding a Codex/GPT-5.5 xhigh frontier harness reference by 13.64 Any Medal points. Ablations and process analyses show that durable project state is central to later-round refinement: removing File-as-Bus lowers PaperBench score by 6.41 points and MLE-Bench Lite Any Medal\% by 31.82 points. These results suggest that long-horizon AI research is not only a problem of stronger local reasoning, but a systems problem of maintaining cumulative, inspectable project progress.

Guoxin Chen, Jie Chen, Lei Chen, Jiale Zhao, Fanzhe Meng, Wayne Xin Zhao, Ruihua Song, Cheng Chen, Ji-Rong Wen, Kai Jia
arXiv:2604.13018 · cs.CL · submitted Apr 14, 2026 · updated May 26, 2026
abstract · pdf · html · Repo: https://github.com/AweAI-Team/AiScientist

add comment on HN
Also discussed: Apr 2026 (1 point, 0 comments)