REVIEW 20 cited by
Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large-scale sequence modeling has sparked rapid advances that now extend into biology and genomics. However, modeling genomic sequences introduces challenges such as the need to model long-range token interactions, the effects of upstream and downstream regions of the genome, and the reverse complementarity (RC) of DNA. Here, we propose an architecture motivated by these challenges that builds off the long-range Mamba block, and extends it to a BiMamba component that supports bi-directionality, and to a MambaDNA block that additionally supports RC equivariance. We use MambaDNA as the basis of Caduceus, the first family of RC equivariant bi-directional long-range DNA language models, and we introduce pre-training and fine-tuning strategies that yield Caduceus DNA foundation models. Caduceus outperforms previous long-range models on downstream benchmarks; on a challenging long-range variant effect prediction task, Caduceus exceeds the performance of 10x larger models that do not leverage bi-directionality or equivariance.
Forward citations
Cited by 20 Pith papers
-
Geometric Hyena Networks for Large-scale Equivariant Learning
Geometric Hyena is an equivariant long-convolutional architecture that captures global geometric context with sub-quadratic complexity and outperforms equivariant transformer baselines on several RNA and protein predi...
-
pLSTM: parallelizable Linear Source Transition Mark networks
pLSTM extends linear recurrent networks to general directed acyclic graphs with a parallelizable scheme and two stabilization modes for long-range propagation.
-
Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences
A new benchmark shows that genomic language models mostly fail to detect loss-of-function mutations in synthetic, evolutionarily implausible DNA, with accuracy tied to how likely the model finds the sequence.
-
Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
A learnable tokenization module with mixture of convolution experts and deformable convolution improves DNA foundation model performance on Genomic and Nucleotide Transformer Benchmarks.
-
Generating Synthetic Genotypes using Diffusion Models
A diffusion model trained on PCA embeddings generates realistic full-length synthetic human genotypes that support disease and population classifiers with near-real-data accuracy.
-
NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable Representations
NucEL shows ELECTRA-style replaced-token pretraining on single-nucleotide DNA tokens reaches state-of-the-art regulatory genomics performance with far fewer parameters.
-
SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation Model
SPACE shows that supervised prediction of genomic profiles such as chromatin accessibility and histone marks produces competitive DNA representations, with a mixture-of-experts architecture that improves cross-species...
-
HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling
A compact hybrid GDN+attention model distilled from Nucleotide Transformer v2 outperforms similarly sized models and, on several tasks, its 500x larger teacher.
-
OmniGenBench: A Modular Platform for Reproducible Genomic Foundation Models Benchmarking
OmniGenBench packages five genomic benchmark suites, 31+ foundation models, automated evaluation, and interpretability tools into standardized one-command workflows.
-
Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning
Autoregressive DNA language models fine-tuned jointly on classification, text generation, and image generation achieve strong benchmark results and open-ended cross-modal genomic tasks.
-
METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring
A 7B transformer pretrained on 1.5T base pairs of wastewater metagenomic reads achieves strong pathogen detection and embedding scores, though some new benchmarks are partly in-distribution.
-
Evaluation of Coding Schemes for Transformer-based Gene Sequence Modeling
BPE tokenization and rotary position embeddings usually outperform k-mers and other positional encodings in from-scratch Transformer DNA classifiers, but the advantage is task-dependent.
-
When repeats drive the vocabulary: a Byte-Pair Encoding analysis of T2T primate genomes
BPE tokenizers trained on nine T2T primate genomes share only 11,569 of 512,000 tokens, and the vocabulary is dominated by short repeats rather than phylogenetic signal.
-
DNAZEN: Enhanced Gene Sequence Representations via Mixed Granularities of Coding Units
DNAZEN augments a Transformer genomic model with PMI-extracted G-gram representations and whole G-gram masking, improving MCC on 21 of 28 GUE datasets against DNABERT-2 after further pre-training.
-
Brain-to-Text Benchmark '24: Lessons Learned
An ensemble of neural decoders merged by a fine-tuned large language model reduced speech-decoding word error rate from 9.7% to 5.8% in the Brain-to-Text Benchmark '24.
-
BarcodeMamba: State Space Models for Biodiversity Analysis
A Mamba-2 state space model pretrained on DNA barcodes matches or beats BarcodeBERT on species and genus classification with far fewer parameters, reaching 99.2% seen-species linear probe and 70.2% unseen-species 1-NN...
-
Simple Guidance Mechanisms for Discrete Diffusion Models
Uniform-noise discrete diffusion trained with a continuous-time variational bound (UDLM) plus discrete classifier-free and classifier-based guidance improves controllable generation over autoregressive baselines on ge...
-
Fast and Scalable Gene Embedding Search: A Comparative Study of FAISS and ScaNN
On 400 bp microbial gene fragments, embedding-based nearest-neighbor search with FAISS and ScaNN outperforms nucleotide MMseqs2 in speed and accuracy; tuned FAISS configurations also beat ScaNN.
-
Improving Genomic Models via Task-Specific Self-Pretraining
Pretraining a small DNA model on unlabeled task-related sequences improves gene finding and other BEND tasks compared to training from scratch, with less data than genome-scale pretraining.
-
Artificial Intelligence for Central Dogma-Centric Multi-Omics: Challenges and Breakthroughs
A literature review that maps AI and deep learning methods for central-dogma-centric multi-omics integration and disease modeling.
Discussion (0). Continue with ORCID to comment.