about
Next Embedding Prediction Makes World Models Stronger (arxiv.org)
1 point by lucrbvi 212 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of rebuilding every pixel of the next frame, this agent predicts how the scene's compressed representation will change, using past states to plan actions. It matches or beats top agents on standard control tasks and gains the most on memory-heavy 3D games.

Abstract

Capturing temporal dependencies is critical for model-based reinforcement learning (MBRL) in partially observable, high-dimensional domains. We introduce NE-Dreamer, a decoder-free MBRL agent that leverages a temporal transformer to predict next-step encoder embeddings from latent state sequences, directly optimizing temporal predictive alignment in representation space. This approach enables NE-Dreamer to learn coherent, predictive state representations without reconstruction losses or auxiliary supervision. On the DeepMind Control Suite, NE-Dreamer matches or exceeds the performance of DreamerV3 and leading decoder-free agents. On a challenging subset of DMLab tasks involving memory and spatial reasoning, NE-Dreamer achieves substantial gains. These results establish next-embedding prediction with temporal transformers as an effective, scalable framework for MBRL in complex, partially observable environments.

George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov
arXiv:2603.02765 · cs.LG, cs.AI · submitted Mar 3, 2026
abstract · pdf · html

add comment on HN