In plain words: Instead of picking one action at a time or searching for a plan, this system runs its world model backwards to emit the full sequence in one pass. On nine maze tasks it beat or matched planners by 24% on average, 10-100 times faster.
Abstract
We present a neuro-inspired framework for embodied planning and control. Building on three principles that enable fast and highly effective goal-directed behavior in the mammalian brain - paired forward/inverse internal models, open-loop multi-step motor commands, and sequential, hierarchical organization of action - our Inverter framework uses learned components, trained end-to-end through Inverse Learning (IL) and supplemented where natural by analytic or algorithmic modules; we formalize IL and delineate it from supervised, reinforcement, and imitation learning. IL bridges Reinforcement Learning (RL)-style amortization, which runs in a single forward pass but emits only one action at a time, and Optimal Control (OC)-style sequence planning over whole trajectories, but with iterative test-time computation. Single Inverters or hierarchical n=2 Inverter stacks match or improve on offline-RL and diffusion-planner baselines on all 3 maze2d and 6 antmaze D4RL variants by an average of +24.2% (range -1.9% to +78.2%), at one-to-two orders of magnitude less inference compute time. Distinctively, optimizing through the Forward Model (FoM) over the entire T-step action sequence - rather than per step - lets Inverters produce smooth, goal-coherent, trajectory-wide structure and reach control policies closer to the analytic optimum than the policy underlying the training data itself. We also identify a failure mode of IL: FoM hacking under narrow training-data coverage, which we mitigate by using random training data with broader coverage. As an application example, a Pulse Inverter synthesizes arbitrary single-qubit quantum gates with fidelity matching the standard iterative numerical baseline (GRAPE), at more than 1000x lower per-gate compute time. In summary, we conclude that IL enables a versatile class of world-interfaces, especially for latency- and resource-critical embodied AI.
Maryna Kapitonova, Tonio Ball
arXiv:2605.24152 · cs.AI · submitted May 22, 2026 · updated May 26, 2026
abstract · pdf · html · Version 2, minor fix in online version of the abstract, pdf unchanged
Our findings show that the Inverter framework, built on the same three principles, enables fast and effective planning and control through a feedforward, sequence-level FoM-and-IM core that emits entire action sequences in single forward passes. We find that Inverters offer consistently high task performance at a fraction of the inference compute time used by step-wise RL or iterative planners.”