In plain words: OpenHands is a free, open platform where AI agents work like human developers—writing code, running commands, and browsing the web—inside safe, isolated sandboxes so their actions can't damage your computer. It tested agents on 15 tough tasks, from fixing real software bugs to navigating websites.
Abstract · OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Software is one of the most powerful tools that we humans have at our disposal; it allows a skilled programmer to interact with the world in complex and profound ways. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and affect change in their surrounding environments. In this paper, we introduce OpenHands (f.k.a. OpenDevin), a platform for the development of powerful and flexible AI agents that interact with the world in similar ways to those of a human developer: by writing code, interacting with a command line, and browsing the web. We describe how the platform allows for the implementation of new agents, safe interaction with sandboxed environments for code execution, coordination between multiple agents, and incorporation of evaluation benchmarks. Based on our currently incorporated benchmarks, we perform an evaluation of agents over 15 challenging tasks, including software engineering (e.g., SWE-BENCH) and web browsing (e.g., WEBARENA), among others. Released under the permissive MIT license, OpenHands is a community project spanning academia and industry with more than 2.1K contributions from over 188 contributors.
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, et al.
arXiv:2407.16741 · cs.SE, cs.AI, cs.CL · submitted Jul 23, 2024 · updated Apr 18, 2025
abstract · pdf · html · Accepted by ICLR 2025; Code: https://github.com/All-Hands-AI/OpenHands
I gave it one example and then asked it to do the work for the other files.
It was able to do about half the files correctly. But it ended up taking an hour, costing >$50 in OpenAI credits, and took me longer to debug, fix, and verify the work than it would have to do the work manually.
My take: good glimpse of the future after a few more Moore’s Law doublings and model improvement cycles make it 10x better, 10x faster, and 10x cheaper. But probably not yet worth trying to use for real work vs playing with it for curiosity, learning, and understanding.
Edit: writing the tests in this PR given the code + one test as an example was the task: https://github.com/roboflow/inference/pull/533
This commit was the manual example: https://github.com/roboflow/inference/pull/533/commits/93165...
This commit adds the partially OpenDevin written ones: https://github.com/roboflow/inference/pull/533/commits/65f51...