about
Remote Labor Index: Measuring AI Automation of Remote Work (arxiv.org)
2 points by belter 337 days ago | hide | past | pdf | discuss on HN

In plain words: A test made of real, paid projects across many kinds of jobs lets AI agents try to finish whole tasks from start to end. The best agent automated just 2.5% of the work, far below what smart-test scores suggest.

Abstract

AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To measure this, we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings. AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%. These results help ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts and enabling stakeholders to proactively navigate AI-driven labor automation.

Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, et al.
arXiv:2510.26787 · cs.LG, cs.AI, cs.CL · submitted Oct 30, 2025
abstract · pdf · html · Website: https://www.remotelabor.ai

add comment on HN
Also discussed: Feb 2026 (2 points, 0 comments)