Pith. sign in

REVIEW 4 cited by

RITA: a Study on Scaling Up Generative Protein Sequence Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.05789 v2 pith:OONDHMZR submitted 2022-05-11 q-bio.QM cs.LG

classification q-bio.QMcs.LG
keywords modelsproteinritagenerativeautoregressivepredictionsequencesaccelerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we introduce RITA: a suite of autoregressive generative models for protein sequences, with up to 1.2 billion parameters, trained on over 280 million protein sequences belonging to the UniRef-100 database. Such generative models hold the promise of greatly accelerating protein design. We conduct the first systematic study of how capabilities evolve with model size for autoregressive transformers in the protein domain: we evaluate RITA models in next amino acid prediction, zero-shot fitness, and enzyme function prediction, showing benefits from increased scale. We release the RITA models openly, to the benefit of the research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Directed Evolution of Proteins via Bayesian Optimization in Embedding Space

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A Gaussian-process Bayesian optimizer running in protein language model embedding space outperforms regression-based directed evolution baselines on two in silico fitness landscapes.

  2. PDFBench: A Benchmark for De novo Protein Design from Function

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The paper presents PDFBench, a unified benchmark with 16 metrics and a new post-2025 protein test set, and finds that evaluation choices such as retrieval strategy or supported keywords can dominate model rankings.

  3. Scaling and Data Saturation in Protein Language Models

    q-bio.QM 2025-07 conditional novelty 5.0 of 10

    Protein language models trained on increasingly large UniRef100 snapshots show non-monotonic fitness prediction performance, with no evidence of data saturation.

  4. Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics

    cs.CL 2025-06 unverdicted novelty 1.0 of 10

    A review maps how NLP architectures from word2vec to Evo 2 are applied to DNA, RNA, protein, and genome sequences, with tokenization and context choices shaping performance.

Pith tools