about
Words That Make Language Models Perceive (arxiv.org)
3 points by tbruckner 299 days ago | hide | past | pdf | discuss on HN

In plain words: Telling a text-only language model to "see" or "hear" before answering nudges its inner states to act as if it had visual or auditory evidence. This simple prompting reliably lined its representations up with dedicated vision and audio models, unlike plain prompts.

Abstract

Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text-only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next-token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality-appropriate representations in purely text-trained LLMs.

Sophie L. Wang, Phillip Isola, Brian Cheung
arXiv:2510.02425 · cs.CL, cs.CV, cs.LG · submitted Oct 2, 2025
abstract · pdf · html

add comment on HN
Also discussed: Oct 2025 (2 points, 0 comments)