about
HyenaDNA: Long-Range Genomic Sequence Modeling (context length of 1M tokens) (arxiv.org)
22 points by famouswaffles on Jun 29, 2023 | hide | past | pdf | 3 comments on HN

In plain words: A DNA model reads the genome one letter at a time using long, cheap filters instead of attention, so it can hold up to 1 million letters. It beat the best prior results on 12 of 18 gene-task tests with far fewer parameters.

Abstract · HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution

Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizable features from unlabeled genome data that can then be fine-tuned for downstream tasks such as identifying regulatory elements. Due to the quadratic scaling of attention, previous Transformer-based genomic models have used 512 to 4k tokens as context (<0.001% of the human genome), significantly limiting the modeling of long-range interactions in DNA. In addition, these methods rely on tokenizers or fixed k-mers to aggregate meaningful DNA units, losing single nucleotide resolution where subtle genetic variations can completely alter protein function via single nucleotide polymorphisms (SNPs). Recently, Hyena, a large language model based on implicit convolutions was shown to match attention in quality while allowing longer context lengths and lower time complexity. Leveraging Hyena's new long-range capabilities, we present HyenaDNA, a genomic foundation model pretrained on the human reference genome with context lengths of up to 1 million tokens at the single nucleotide-level - an up to 500x increase over previous dense attention-based models. HyenaDNA scales sub-quadratically in sequence length (training up to 160x faster than Transformer), uses single nucleotide tokens, and has full global context at each layer. We explore what longer context enables - including the first use of in-context learning in genomics. On fine-tuned benchmarks from the Nucleotide Transformer, HyenaDNA reaches state-of-the-art (SotA) on 12 of 18 datasets using a model with orders of magnitude less parameters and pretraining data. On the GenomicBenchmarks, HyenaDNA surpasses SotA on 7 of 8 datasets on average by +10 accuracy points. Code at https://github.com/HazyResearch/hyena-dna.

Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli, Yoshua Bengio, Stefano Ermon, Stephen A. Baccus, et al.
arXiv:2306.15794 · cs.LG, q-bio.GN · submitted Jun 27, 2023 · updated Nov 14, 2023
abstract · pdf · html · NeurIPS 2023 (Spotlight)

add comment on HN

Part of what’s interesting here is that there haven’t been any robust featurizers for DNA in the same sense that we have robust featurizers for proteins (like ProtBERT et al.) that just work out of the box for amino acid sequences. Since many genes have >100k bp vs. proteins that are often less than 1k AAs, the much longer context window is needed. ProtBERT gets ~200k downloads per month on huggingface and DNABERT gets ~10s.
This is really cool...these LLM-like models in technical disciplines are going to open up he new frontiers in science.
And if you accidentially created a monster, because it switched some chars, you can always tell it: That was good, but you can do better and it will finish the job correctly.