about
Aide: AI-Driven Exploration in the Space of Code (Arxiv) (arxiv.org)
1 point by WecoAI on Feb 19, 2025 | hide | past | pdf | 1 comment on HN

In plain words: AIDE is an AI agent that treats building machine learning models as code optimization, trying solutions as branches of a search tree and reusing the most promising ones instead of starting over. It beat the best results on several machine learning engineering benchmarks.

Abstract · AIDE: AI-Driven Exploration in the Space of Code

Machine learning, the foundation of modern artificial intelligence, has driven innovations that have fundamentally transformed the world. Yet, behind advancements lies a complex and often tedious process requiring labor and compute intensive iteration and experimentation. Engineers and scientists developing machine learning models spend much of their time on trial-and-error tasks instead of conceptualizing innovative solutions or research hypotheses. To address this challenge, we introduce AI-Driven Exploration (AIDE), a machine learning engineering agent powered by large language models (LLMs). AIDE frames machine learning engineering as a code optimization problem, and formulates trial-and-error as a tree search in the space of potential solutions. By strategically reusing and refining promising solutions, AIDE effectively trades computational resources for enhanced performance, achieving state-of-the-art results on multiple machine learning engineering benchmarks, including our Kaggle evaluations, OpenAI MLE-Bench and METRs RE-Bench.

Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, Yuxiang Wu
arXiv:2502.13138 · cs.AI, cs.LG · submitted Feb 18, 2025
abstract · pdf · html

add comment on HN

We’re excited to share our work on AIDE, an LLM-powered agent that automates machine learning engineering through systematic trial-and-error. Unlike conventional AutoML, AIDE searches directly in the space of code, iteratively refining solutions using a structured tree search approach.

Key Highlights:

- State-of-the-art performance: AIDE outperforms human competitors in Kaggle-style ML tasks, as shown in OpenAI’s MLE-Bench. - Scalability & efficiency: By reusing and refining promising solutions, AIDE achieves 3.5x performance gains over o1 alone. - Beyond tabular ML: AIDE extends to deep learning and even AI research tasks, showing human-level capabilities in structured R&D. - We’re open-sourcing the code and sharing our research paper to help the community build on top of it. Would love to hear your thoughts!

Paper: https://www.arxiv.org/abs/2502.13138 Code: https://github.com/WecoAI/aideml