about
Agentless: Demystifying LLM-Based Software Engineering Agents (arxiv.org)
4 points by dmezzetti on Jul 8, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of letting an AI run commands and plan its steps, this tool finds the buggy code, writes a fix, and checks the patch in three steps. On real GitHub bug reports it fixed 96 issues (32%), beating every open-source agent while costing $0.70.

Abstract · Agentless: Demystifying LLM-based Software Engineering Agents

Recent advancements in large language models (LLMs) have significantly advanced the automation of software development tasks, including code synthesis, program repair, and test generation. More recently, researchers and industry practitioners have developed various autonomous LLM agents to perform end-to-end software development tasks. These agents are equipped with the ability to use tools, run commands, observe feedback from the environment, and plan for future actions. However, the complexity of these agent-based approaches, together with the limited abilities of current LLMs, raises the following question: Do we really have to employ complex autonomous software agents? To attempt to answer this question, we build Agentless -- an agentless approach to automatically solve software development problems. Compared to the verbose and complex setup of agent-based approaches, Agentless employs a simplistic three-phase process of localization, repair, and patch validation, without letting the LLM decide future actions or operate with complex tools. Our results on the popular SWE-bench Lite benchmark show that surprisingly the simplistic Agentless is able to achieve both the highest performance (32.00%, 96 correct fixes) and low cost ($0.70) compared with all existing open-source software agents! Furthermore, we manually classified the problems in SWE-bench Lite and found problems with exact ground truth patch or insufficient/misleading issue descriptions. As such, we construct SWE-bench Lite-S by excluding such problematic issues to perform more rigorous evaluation and comparison. Our work highlights the current overlooked potential of a simple, interpretable technique in autonomous software development. We hope Agentless will help reset the baseline, starting point, and horizon for autonomous software agents, and inspire future work along this crucial direction.

Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming Zhang
arXiv:2407.01489 · cs.SE, cs.AI, cs.CL, cs.LG · submitted Jul 1, 2024 · updated Oct 29, 2024
abstract · pdf · html

add comment on HN