about
Latent-Space Communication in Heterogeneous Multi-Agent Systems (arxiv.org)
7 points by ekaesmem 217 days ago | hide | past | pdf | 1 comment on HN

In plain words: Instead of translating internal states between different AI models, this system packs one model's ongoing thoughts into a fixed-size picture-like message that another reads through its normal image input. On nine reasoning tests it beat text-based messaging by 6.0 points and ran 1.69 times faster.

Abstract · Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with $O(N)$ components for $N$ models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69$\times$ geometric-mean speedup in batch-normalized end-to-end runtime.

Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
arXiv:2602.15382 · cs.CL, cs.CV, cs.LG · submitted Feb 17, 2026 · updated Sep 28, 2026
abstract · pdf · html · 32 pages, 9 figures, 16 tables

add comment on HN

Literally just yesterday I asked why this wasn't being done:

https://news.ycombinator.com/item?id=47195212

  > "Models have unlocked advanced collaborative reasoning, yet they remain shackled by the inefficiency of discrete text communication, which imposes significant runtime overhead and information quantization loss."