about
Emergent World Representations in Neural Networks (arxiv.org)
3 points by mjburgess on Feb 15, 2023 | hide | past | pdf | 1 comment on HN

In plain words: A network trained only to predict legal Othello moves, with no rules given, builds an internal picture of the board instead of just memorizing move patterns. Editing that hidden picture changed the moves it predicted and produced human-readable explanations of its choices.

Abstract · Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task

Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics, or do they rely on internal representations of the process that generates the sequences they see? We investigate this question by applying a variant of the GPT model to the task of predicting legal moves in a simple board game, Othello. Although the network has no a priori knowledge of the game or its rules, we uncover evidence of an emergent nonlinear internal representation of the board state. Interventional experiments indicate this representation can be used to control the output of the network and create "latent saliency maps" that can help explain predictions in human terms.

Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, Martin Wattenberg
arXiv:2210.13382 · cs.LG, cs.AI, cs.CL · submitted Oct 24, 2022 · updated Jun 26, 2024
abstract · pdf · html · ICLR 2023 oral (notable-top-5%): https://openreview.net/forum?id=DeG07_TcZvT ; code: https://github.com/likenneth/othello_world

add comment on HN

I've just read over this paper which is one of the best ones I've read coming out of the field.

However I'm inclined to dispute their conclusions.

They seem not to address the obvious (?) reply: that the board state is just a function of the game moves.

To say that the NN has built a representation of the board state from moves is trivial, if moves are just another phrasing of the board state.

That there's a correlation between trained weights and given board states is then expected, since to be trained on moves is to be trained on board states.

Consider the difference between a network seeming to obtaining a 3D model of a simple object from a handful of 2D photographs, vs. it doing so from millions at every angle (in every lighting condition, etc.).

In the latter case of course the weights of the trained network will correlated with the actual 3D model, because there's almost no information gap between the actual 3D and the millions of 2D images provided. (There is still an exploitable gap though, which could be used to show the network hadnt learnt the model).

The paper seems to address a strawman claim that NNs are "just remembering their inputs" as literally tokenised. That isn't the claim. It's that they remember their inputs as phrased in a transformed space.

Here: if board moves are just an alternative way of specifying the board state, "phrased in a transformed space" -- then the claim stands.

The claim that NNs are just "remembering their inputs" (with an extra step).