about
Causal World Modeling for Robot Control (arxiv.org)
1 point by mountainview 237 days ago | hide | past | pdf | discuss on HN

In plain words: Unlike usual robot models that only pick actions, this one also imagines the next video frames, learning how its moves change what it sees. It handled long, multi-step manipulation with less extra training and worked on new setups it hadn't seen.

Abstract

This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by understanding the causality between actions and visual dynamics. Inspired by this, we introduce LingBot-VA, an autoregressive diffusion framework that learns frame prediction and policy execution simultaneously. Our model features three carefully crafted designs: (1) a shared latent space, integrating vision and action tokens, driven by a Mixture-of-Transformers (MoT) architecture, (2) a closed-loop rollout mechanism, allowing for ongoing acquisition of environmental feedback with ground-truth observations, (3) an asynchronous inference pipeline, parallelizing action prediction and motor execution to support efficient control. We evaluate our model on both simulation benchmarks and real-world scenarios, where it shows significant promise in long-horizon manipulation, data efficiency in post-training, and strong generalizability to novel configurations. The code and model are made publicly available to facilitate the community.

Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu
arXiv:2601.21998 · cs.CV, cs.RO · submitted Jan 29, 2026 · updated Mar 22, 2026
abstract · pdf · html · Project page: https://technology.robbyant.com/lingbot-va Code: https://github.com/robbyant/lingbot-va

add comment on HN