about
Diffusion Models Are Real-Time Game Engines (arxiv.org)
2 points by LordNibbler on Aug 29, 2024 | hide | past | pdf | discuss on HN

In plain words: A neural net trained on recorded play draws a new frame from past frames and actions, instead of hand-coded rules. It runs DOOM at 20 frames per second on one chip, stays stable for minutes, and human raters barely tell it from the real game.

Abstract

We present GameNGen, the first game engine powered entirely by a neural model that also enables real-time interaction with a complex environment over long trajectories at high quality. When trained on the classic game DOOM, GameNGen extracts gameplay and uses it to generate a playable environment that can interactively simulate new trajectories. GameNGen runs at 20 frames per second on a single TPU and remains stable over extended multi-minute play sessions. Next frame prediction achieves a PSNR of 29.4, comparable to lossy JPEG compression. Human raters are only slightly better than random chance at distinguishing short clips of the game from clips of the simulation, even after 5 minutes of auto-regressive generation. GameNGen is trained in two phases: (1) an RL-agent learns to play the game and the training sessions are recorded, and (2) a diffusion model is trained to produce the next frame, conditioned on the sequence of past frames and actions. Conditioning augmentations help ensure stable auto-regressive generation over long trajectories, and decoder fine-tuning improves the fidelity of visual details and text.

Dani Valevski, Yaniv Leviathan, Moab Arar, Shlomi Fruchter
arXiv:2408.14837 · cs.LG, cs.AI, cs.CV · submitted Aug 27, 2024 · updated Apr 24, 2025
abstract · pdf · html · ICLR 2025. Project page: https://gamengen.github.io/

add comment on HN
Also discussed: Sep 2024 (1 point, 0 comments)