about
Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges (arxiv.org)
1 point by wek 210 days ago | hide | past | pdf | discuss on HN

In plain words: Unlike typical benchmarks made of invented tasks, this test suite gathers 104 real-world scenarios from social media and product questions, checking each for realism, difficulty, and verifiable answers. Testing AI agents on it exposed clear gaps between their abilities and real user needs.

Abstract · LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges

As large language models grow more capable, general AI agents have become increasingly prevalent in practical applications. However, existing benchmarks face significant limitations, failing to represent real-world user tasks accurately. To address this gap, we present LiveAgentBench, a comprehensive benchmark with 104 scenarios that reflect real user requirements. It is constructed from publicly sourced questions on social media and real-world products. Central to our approach is the Social Perception-Driven Data Generation (SPDG) method, a novel process we developed to ensure each question's real-world relevance, task complexity, and result verifiability. We evaluate various models, frameworks, and commercial products using LiveAgentBench, revealing their practical performance and identifying areas for improvement. This release includes 374 tasks, with 125 for validation and 249 for testing. The SPDG process enables continuous updates with fresh queries from real-world interactions.

Hao Li, Huan Wang, Jinjie Gu, Wenjie Wang, Chenyi Zhuang, Sikang Bian
arXiv:2603.02586 · cs.AI · submitted Mar 3, 2026
abstract · pdf · html

add comment on HN