Pith. sign in

REVIEW 6 cited by

ProGen: Language Modeling for Protein Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.03497 v1 pith:KMYKXUJ2 submitted 2020-03-08 q-bio.BM cs.LGstat.ML

classification q-bio.BMcs.LGstat.ML
keywords proteinprogensequenceengineeringgenerationlanguagemodelingaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative modeling for protein engineering is key to solving fundamental problems in synthetic biology, medicine, and material science. We pose protein engineering as an unsupervised sequence generation problem in order to leverage the exponentially growing set of proteins that lack costly, structural annotations. We train a 1.2B-parameter language model, ProGen, on ~280M protein sequences conditioned on taxonomic and keyword tags such as molecular function and cellular component. This provides ProGen with an unprecedented range of evolutionary sequence diversity and allows it to generate with fine-grained control as demonstrated by metrics based on primary sequence similarity, secondary structure accuracy, and conformational energy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 34 citations worldwide. Full citation record

  1. DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing

    q-bio.QM 2025-06 conditional novelty 7.0 of 10

    DisProtEdit learns disentangled protein representations from separate structural and functional text descriptions, enabling controllable single- and multi-attribute protein editing via latent interpolation.

  2. Steering Protein Language Models

    q-bio.BM 2025-07 reject novelty 6.0 of 10

    Activation steering can guide protein language models to generate and optimize sequences with higher predicted thermostability, solubility, or GFP brightness, but only in surrogate-based evaluation.

  3. Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction

    q-bio.QM 2025-06 conditional novelty 6.0 of 10

    Protriever trains a retriever and a protein language model together so the model learns which homologs to retrieve, reaching state-of-the-art zero-shot fitness prediction on ProteinGym with much faster retrieval.

  4. Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

    cs.LG 2026-06 conditional novelty 5.0 of 10

    A two-stage recipe — camelid-sequence continual pretraining plus DPO-style preference fine-tuning with weak positive pairs — improves antibody expression ranking over supervised baselines on an internal 1254-sequence ...

  5. Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

    cs.CL 2025-11 reject novelty 4.0 of 10

    Across the 68 papers it surveys, domain-specialized generative models usually outperform general-purpose LLMs on biological tasks, and agentic/conversational workflows are the least-covered topics.

  6. AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization

    q-bio.BM 2025-06 reject novelty 4.0 of 10

    DPO with contrastive sequence-annotation alignment improves GO term prediction by 2 to 4 percent relative F1-Max over supervised fine-tuning alone.

Pith tools