about
Inferring the Phylogeny of Large Language Models (arxiv.org)
69 points by weinzierl on Apr 19, 2025 | hide | past | pdf | 6 comments on HN

In plain words: Like tracing a species' family tree, this tool compares models' answers to the same questions to measure how closely related they are and draw a family tree. The trees matched known relationships across 156 models and predicted benchmark scores without running the benchmarks.

Abstract · PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks

This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their performance characteristics. Our method calculates a phylogenetic distance metric based on the similarity of LLMs' output. The resulting metric is then used to construct dendrograms, which satisfactorily capture known relationships across a set of 111 open-source and 45 closed models. Furthermore, our phylogenetic distance predicts performance in standard benchmarks, thus demonstrating its functional validity and paving the way for a time and cost-effective estimation of LLM capabilities. To sum up, by translating population genetic concepts to machine learning, we propose and validate a tool to evaluate LLM development, relationships and capabilities, even in the absence of transparent training information.

Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri
arXiv:2404.04671 · cs.CL, cs.LG, q-bio.PE · submitted Apr 6, 2024 · updated Dec 8, 2025
abstract · pdf · html · The project code is available at https://github.com/Nicolas-Yax/PhyloLM . Published as https://iclr.cc/virtual/2025/poster/28195 at ICLR 2025. A code demo is available at https://colab.research.google.com/drive/1agNE52eUevgdJ3KL3ytv5Y9JBbfJRYqd

add comment on HN

Intuitive and expected result (maybe without the prediction of performance). I'm glad somebody did the hard work of proving it.

Though, if this is so clearly seen, how come AI detectors perform so badly?

This experiment involves each LLM responding to 128 or 256 prompts. AI detection is generally focused on determining the writer of a single document, not comparing two analagous sets of 128 documents and determining if the same person/tool wrote both. Totally different problem.
It might be because detecting if output is AI generated and mapping output which is known to be from an LLM to a specific LLM or class of LLMs are different problems.
They're discovering the wrong thing. And the analogy with biology doesn't hold.

They're sensitive not to architecture but to training data. That's like grouping animals by what environment they lived in, so lions and alligators are closer to one another than lions and cats.

The real trick is to infer the underlying architecture and show the relationships between architectures.

That's not something you can tell easily by just looking at the name of the model. And that would actually be useful. This is pretty useless.

You are the one making a wrong biological analogy. Architecture isn't comparable to genes any more than training data is comparable to genes, and training data isn't comparable to environment, doing these kind of analogies brings you nothing but false confidence and misunderstanding.

What they do in the paper on the other hands is to apply the methods of biology, and get a result that is akin to phylogeny, not from a biological analogy but from a biologically-inspired method.

This is provocative but off-base in order to be so: why would we need to work backwards to determine architecture?

Similarly, "you can tell easily by just looking at the name of the model" -- that's an unfounded assertion. No, you can't. It's perfectly cromulent, accepted, and quite regular to have a fine-tuned model that has nothing in its name indicating what it was fine-tuned on. (we can observe the effects of this even if we aren't so familiar with domain enough to know this, i.e. Meta in Llama 4 making it a requirement to have it in the name)