about
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence (arxiv.org)
2 points by doener 187 days ago | hide | past | pdf | discuss on HN

In plain words: A new benchmark puts agents in turn-based puzzle worlds with no instructions, so they must poke around, work out the goal and rules, then plan moves. Humans solved every puzzle, while the best AI systems solved under 1%.

Abstract

We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.

ARC Prize Foundation
arXiv:2603.24621 · cs.AI · submitted Mar 24, 2026 · updated Apr 17, 2026
abstract · pdf · html

add comment on HN
Also discussed: Aug 2026 (2 points, 0 comments)