about
We Automated RL Environment Engineering for $10 (arxiv.org)
2 points by milkkarten 206 days ago | hide | past | pdf | discuss on HN

In plain words: A system automatically rewrites training worlds for AI agents into fast code, then checks the rewrite behaves identically using tests and the same agent in both. Across five worlds it matched the originals, with simulation taking under 4% of training time at 200M parameters.

Abstract · Automatic Generation of High-Performance RL Environments

Translating complex reinforcement learning (RL) environments into high-performance implementations has traditionally required months of specialized engineering. We present a closed-loop methodology that produces equivalent high-performance environments for minimal compute cost. Our method uses a generic prompt template, hierarchical verification (property, interaction, and rollout tests), iterative repair, and cross-backend policy transfer to verify no sim-to-sim gap. We demonstrate three distinct workflows across five environments: (1) Direct translation (no prior performance implementation exists) from Game Boy emulator PyBoy to our EmuRust (via Rust IPC) and from Pokemon Showdown to our PokeJAX (via JAX); (2) Translation verified against existing performance implementations via throughput parity with Puffer Pong, MJX and Brax at matched GPU batch sizes; and (3) New environment creation: TCGJax, the first Pokemon TCG Pocket environment, created from a web-extracted specification. At 200M parameters, the environment overhead drops below 4% of training time. Our closed-loop methodology confirms equivalence for all five environments. TCGJax, synthesized from a private reference absent from public repositories, serves as a contamination control for agent pretraining data concerns.

Seth Karten, Rahul Dev Appapogu, Chi Jin
arXiv:2603.12145 · cs.LG, cs.AI, cs.SE · submitted Mar 12, 2026 · updated May 17, 2026
abstract · pdf · html · 20 pages, 5 figures

add comment on HN