about
Transformers Are Graph Neural Networks (arxiv.org)
33 points by Anon84 on Jun 30, 2025 | hide | past | pdf | 2 comments on HN

In plain words: A Transformer can be seen as a graph network passing messages between every pair of words, with attention deciding how much each matters and position hints giving order. It matches graph networks mathematically but runs faster because dense matrix math suits today's chips.

Abstract · Transformers are Graph Neural Networks

We establish connections between the Transformer architecture, originally introduced for natural language processing, and Graph Neural Networks (GNNs) for representation learning on graphs. We show how Transformers can be viewed as message passing GNNs operating on fully connected graphs of tokens, where the self-attention mechanism capture the relative importance of all tokens w.r.t. each-other, and positional encodings provide hints about sequential ordering or structure. Thus, Transformers are expressive set processing networks that learn relationships among input elements without being constrained by apriori graphs. Despite this mathematical connection to GNNs, Transformers are implemented via dense matrix operations that are significantly more efficient on modern hardware than sparse message passing. This leads to the perspective that Transformers are GNNs currently winning the hardware lottery.

Chaitanya K. Joshi
arXiv:2506.22084 · cs.LG, cs.AI · submitted Jun 27, 2025
abstract · pdf · html · This paper is a technical version of an article in The Gradient at https://thegradient.pub/transformers-are-graph-neural-networks/

add comment on HN

This was a hot topic in 2019-2020 and I just thought I‘ve read an article like this back then from NYU and another from the National University of Singapore, because I and many others were at the time independently working on graph problems using Transformers [0,1].

Turns out this article _is_ the one from NYU, just on arxiv now:

https://graphdeeplearning.github.io/post/transformers-are-gn...

[0] e.g. https://arxiv.org/abs/2003.10536, [1] ProteinMPNN, Dauparas et al.

Glad someone formalized this.