about
Towards a Science of AI Agent Reliability (arxiv.org)
2 points by smartmic 222 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of one success score, this study rates agent reliability with twelve checks: does it act the same each run, shrug off small changes, fail predictably, and keep mistakes small? Across 15 models and two task sets, newer, more accurate agents improved only slightly.

Abstract

AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a fundamental limitation of current evaluations: compressing agent behavior into a single success metric obscures critical operational flaws. Notably, it ignores whether agents behave consistently across runs, withstand perturbations, fail predictably, or have bounded error severity. Grounded in safety-critical engineering, we provide a holistic performance profile by proposing twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety. Evaluating 15 models across two complementary benchmarks, we find that recent capability gains have only yielded small improvements in reliability. By exposing these persistent limitations, our metrics complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.

Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan
arXiv:2602.16666 · cs.AI, cs.CY, cs.LG · submitted Feb 18, 2026 · updated Jun 2, 2026
abstract · pdf · html · Accepted at ICML 2026. Interactive dashboard available at: https://hal.cs.princeton.edu/reliability

add comment on HN