In plain words: Success drops at a steady rate for every extra minute of work a human would need, so each agent has a task length where its success rate halves. This simple rule matched measured success rates across research-engineering task lengths.
Abstract
Building on the recent empirical work of Kwa et al. (2025), I show that within their suite of research-engineering tasks the performance of AI agents on longer-duration tasks can be explained by an extremely simple mathematical model -- a constant rate of failing during each minute a human would take to do the task. This implies an exponentially declining success rate with the length of the task and that each agent could be characterised by its own half-life. This empirical regularity allows us to estimate the success rate for an agent at different task lengths. And the fact that this model is a good fit for the data is suggestive of the underlying causes of failure on longer tasks -- that they involve increasingly large sets of subtasks where failing any one fails the task. Whether this model applies more generally on other suites of tasks is unknown and an important subject for further work.
Toby Ord
arXiv:2505.05115 · cs.AI · submitted May 8, 2025
abstract · pdf · 9 pages, 5 figures