about
EvoBPE: Evolutionary Protein Sequence Tokenization (arxiv.org)
1 point by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss on HN

In plain words: It learns protein "words" from raw sequences, then groups them into families of mutation variants using known substitution rates, like word spellings sharing an ancestor. Mutations that stay inside a family turn out to be harmless more often than the substitution rates alone predict.

Abstract · PUMA: Learning a Mutation-Aware Vocabulary of Protein Units

Modeling protein sequences as a language has made language models a powerful tool in computational biology, yet the language itself remains poorly understood. A key step toward understanding it is identifying its constituent units. In natural languages, morphemes can occur in multiple forms; similarly, in proteins, mutations can give rise to variations of a unit that persist through evolution, forming families of related units. We introduce PUMA (Protein Units via Mutation-Aware Merging), an algorithm that learns protein units from sequence and explores their mutational variants using substitution matrices, forming a genealogy of unit families. Our results show that mutations remaining within a PUMA family are more often benign than the substitution matrix alone predicts, and that PUMA genealogy improves molecular function representations compared to treating units independently. A case study of a unit family demonstrates relatedness beyond homology. PUMA achieves competitive performance on downstream tasks when used as a protein language model tokenizer. Moreover, collapsing units into families results in a smaller embedding table and faster training. Together, these results support PUMA as a biologically grounded protein vocabulary that organizes protein units into plausible families of mutational variants. The source code is available at https://github.com/boun-tabi-lifelu/PUMA.

Burak Suyunu, Özdeniz Dolu, Ibukunoluwa Abigail Olaosebikan, Hacer Karatas Bristow, Arzucan Özgür
arXiv:2503.08838 · cs.CL, q-bio.QM · submitted Mar 11, 2025 · updated Oct 1, 2026
abstract · pdf · html · 23 pages, 10 figures, 9 tables, 1 algorithm

add comment on HN