Pith. sign in

REVIEW 5 cited by

DNAGPT: A Generalized Pre-trained Tool for Versatile DNA Sequence Analysis Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.05628 v3 pith:YSNGWFB5 submitted 2023-07-11 q-bio.GN cs.LG

classification q-bio.GNcs.LG
keywords tasksdnagptmodelsequenceanalysisdatadesignedgeneralized
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained large language models demonstrate potential in extracting information from DNA sequences, yet adapting to a variety of tasks and data modalities remains a challenge. To address this, we propose DNAGPT, a generalized DNA pre-training model trained on over 200 billion base pairs from all mammals. By enhancing the classic GPT model with a binary classification task (DNA sequence order), a numerical regression task (guanine-cytosine content prediction), and a comprehensive token language, DNAGPT can handle versatile DNA analysis tasks while processing both sequence and numerical data. Our evaluation of genomic signal and region recognition, mRNA abundance regression, and artificial genomes generation tasks demonstrates DNAGPT's superior performance compared to existing models designed for specific downstream tasks, benefiting from pre-training using the newly designed model structure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging

    cs.AI 2026-08 conditional novelty 6.0 of 10

    GeneFuse fuses genomic language model embeddings with MRI via multi-scale FiLM modulation and uncertainty-gated residuals, improving dementia screening AUROC by 0.12 over image-only on the ADNI cohort.

  2. GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance

    cs.CR 2025-05 conditional novelty 6.0 of 10

    GeneBreaker, a new attack framework, steers DNA language models to generate sequences with over 90% identity to human pathogens, with success rates up to 60% on the largest Evo2 model.

  3. From Prompt to Pipeline: Large Language Models for Scientific Workflow Development in Bioinformatics

    cs.SE 2025-07 conditional novelty 5.0 of 10

    A qualitative study of ten bioinformatics workflows finds LLMs can generate usable Galaxy and Nextflow pipelines, with Gemini best for Galaxy and DeepSeek-V3 best for Nextflow.

  4. Evaluation of Coding Schemes for Transformer-based Gene Sequence Modeling

    cs.CL 2025-07 conditional novelty 5.0 of 10

    BPE tokenization and rotary position embeddings usually outperform k-mers and other positional encodings in from-scratch Transformer DNA classifiers, but the advantage is task-dependent.

  5. DNAZEN: Enhanced Gene Sequence Representations via Mixed Granularities of Coding Units

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DNAZEN augments a Transformer genomic model with PMI-extracted G-gram representations and whole G-gram masking, improving MCC on 21 of 28 GUE datasets against DNABERT-2 after further pre-training.

Pith tools