In plain words: Comparing the model's inner activity on two kinds of physics simulations reveals a single direction encoding a physical feature; pushing it back in adds or removes that feature in new simulations. This shows the model learned real physical principles, not just surface patterns.
Abstract · Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
Recent advances in mechanistic interpretability have revealed that large language models (LLMs) develop internal representations corresponding not only to concrete entities but also distinct, human-understandable abstract concepts and behaviour. Moreover, these hidden features can be directly manipulated to steer model behaviour. However, it remains an open question whether this phenomenon is unique to models trained on inherently structured data (ie. language, images) or if it is a general property of foundation models. In this work, we investigate the internal representations of a large physics-focused foundation model. Inspired by recent work identifying single directions in activation space for complex behaviours in LLMs, we extract activation vectors from the model during forward passes over simulation datasets for different physical regimes. We then compute "delta" representations between the two regimes. These delta tensors act as concept directions in activation space, encoding specific physical features. By injecting these concept directions back into the model during inference, we can steer its predictions, demonstrating causal control over physical behaviours, such as inducing or removing some particular physical feature from a simulation. These results suggest that scientific foundation models learn generalised representations of physical principles. They do not merely rely on superficial correlations and patterns in the simulations. Our findings open new avenues for understanding and controlling scientific foundation models and has implications for AI-enabled scientific discovery.
Rio Alexa Fear, Payel Mukhopadhyay, Michael McCabe, Alberto Bietti, Miles Cranmer
arXiv:2511.20798 · cs.LG, cs.AI, physics.comp-ph · submitted Nov 25, 2025 · updated Nov 28, 2025
abstract · pdf · html · 16 Pages, 9 Figures. Code available soon at https://github.com/DJ-Fear/walrus_steering
Today they posted a follow up paper examining how the model represents the physical world.
Basically they look at the model's "thoughts" and show that it understands very abstract physical ideas (like speed, or diffusion).
To be clear Walrus itself isn’t yet a fully general purpose “physics AI” because it only works with continuum (fluid-like) data, but it feels like a big stepping stone because it is able to handle anything that is vaguely fluid like (e.g. plasma, gasses, acoustics, turbulence, astrophysics etc). It appears the model is looking at all these different systems and finding general laws/principles that underly everything.