Pith. sign in

REVIEW 2 cited by

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02045 v1 pith:PURKXMCY submitted 2025-01-03 q-bio.GN cs.AIcs.CLcs.LG

classification q-bio.GNcs.AIcs.CLcs.LG
keywords metagenomicmodeldatasetgenomicmetagene-1detectionmonitoringpandemic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We pretrain METAGENE-1, a 7-billion-parameter autoregressive transformer model, which we refer to as a metagenomic foundation model, on a novel corpus of diverse metagenomic DNA and RNA sequences comprising over 1.5 trillion base pairs. This dataset is sourced from a large collection of human wastewater samples, processed and sequenced using deep metagenomic (next-generation) sequencing methods. Unlike genomic models that focus on individual genomes or curated sets of specific species, the aim of METAGENE-1 is to capture the full distribution of genomic information present within this wastewater, to aid in tasks relevant to pandemic monitoring and pathogen detection. We carry out byte-pair encoding (BPE) tokenization on our dataset, tailored for metagenomic sequences, and then pretrain our model. In this paper, we first detail the pretraining dataset, tokenization strategy, and model architecture, highlighting the considerations and design choices that enable the effective modeling of metagenomic data. We then show results of pretraining this model on our metagenomic dataset, providing details about our losses, system metrics, and training stability over the course of pretraining. Finally, we demonstrate the performance of METAGENE-1, which achieves state-of-the-art results on a set of genomic benchmarks and new evaluations focused on human-pathogen detection and genomic sequence embedding, showcasing its potential for public health applications in pandemic monitoring, biosurveillance, and early detection of emerging health threats.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences

    q-bio.QM 2025-06 conditional novelty 7.0 of 10

    A new benchmark shows that genomic language models mostly fail to detect loss-of-function mutations in synthetic, evolutionarily implausible DNA, with accuracy tied to how likely the model finds the sequence.

  2. AutomataGPT: Forecasting and Ruleset Inference for Two-Dimensional Cellular Automata

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A transformer pretrained on 100 cellular automaton rules forecasts unseen rules at 98.5% one-step accuracy and infers new rules with up to 96% functional accuracy.

Pith tools