about
MotionLM: Multi-Agent Motion Forecasting as Language Modeling (arxiv.org)
84 points by bpierre on Oct 2, 2023 | hide | past | pdf | 2 comments on HN

In plain words: It turns future paths into word-like tokens and predicts them like text, generating everyone's paths together in one pass instead of one at a time and scoring how they interact. It ranked first in a leading public contest for predicting how cars move together.

Abstract

Reliable forecasting of the future behavior of road agents is a critical component to safe planning in autonomous vehicles. Here, we represent continuous trajectories as sequences of discrete motion tokens and cast multi-agent motion prediction as a language modeling task over this domain. Our model, MotionLM, provides several advantages: First, it does not require anchors or explicit latent variable optimization to learn multimodal distributions. Instead, we leverage a single standard language modeling objective, maximizing the average log probability over sequence tokens. Second, our approach bypasses post-hoc interaction heuristics where individual agent trajectory generation is conducted prior to interactive scoring. Instead, MotionLM produces joint distributions over interactive agent futures in a single autoregressive decoding process. In addition, the model's sequential factorization enables temporally causal conditional rollouts. The proposed approach establishes new state-of-the-art performance for multi-agent motion prediction on the Waymo Open Motion Dataset, ranking 1st on the interactive challenge leaderboard.

Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S. Refaat, Rami Al-Rfou, Benjamin Sapp
arXiv:2309.16534 · cs.CV, cs.AI, cs.LG, cs.RO · submitted Sep 28, 2023
abstract · pdf · html · To appear at the International Conference on Computer Vision (ICCV) 2023

add comment on HN
Also discussed: Sep 2023 (2 points, 0 comments)

How well does the model handle domain shift in real time? A big issue with a lot of AV models is the ability to adapt to domain shift in real time and doing so quickly! I wonder if the extensive world knowledge in pre-trained LLM’s could overcome this challenge and remove the need for online domain adaptation, where the model needs to be trained in real-time and thus have slow performance!
Has anyone utilized multi agents in production curious the use cases