In plain words: An agent must figure out a hidden rule machine by asking whether strings are accepted and guessing the whole machine, like reverse-engineering a program. Language models solved small machines but failed as these grew, though they knew the standard algorithm and could code it.
Abstract · Can Agents Infer Environment from Interaction? Evidence from Agentic Automata Learning
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction, a capability increasingly required in agentic tasks (e.g., reproducing an executable without access to its source code by interacting with it). In our setup, an agent should uncover a hidden deterministic finite automaton (DFA) by interacting with an oracle through (1) membership queries ("Does this string belong to the target language?'') and (2) equivalence queries ("Is this the target DFA?''). Agentic automata learning yields a scalable testbed with controlled task complexity, measurable interaction efficiency, and strong algorithms to compare against from the classic automata-learning literature. Evaluating state-of-the-art LLMs with a multi-turn agent scaffold, we find that while they are able to recover simple DFAs, their performance drops sharply as DFA size increases. Results improve with a more elaborate ReAct-style state-tracking scaffold, yet strong models still struggle with complex instances. Trajectory analyses reveal recurring failures in query planning, evidence integration, and hypothesis construction. These failures occur even though the models mention classic automata learning algorithms in their reasoning, and can implement and execute them when given access to a coding environment. Our results suggest that for state-of-the-art LLMs, identifying a solution to the problem is insufficient, and reliably executing the resulting plan still poses a distinct challenge.
Reef Menaged, Gili Lior, Shauli Ravfogel, Roee Aharoni, Gabriel Stanovsky
arXiv:2606.16576 · cs.CL · submitted Jun 15, 2026 · updated Sep 27, 2026
abstract · pdf · html