REVIEW 4 cited by
RITA: a Study on Scaling Up Generative Protein Sequence Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work we introduce RITA: a suite of autoregressive generative models for protein sequences, with up to 1.2 billion parameters, trained on over 280 million protein sequences belonging to the UniRef-100 database. Such generative models hold the promise of greatly accelerating protein design. We conduct the first systematic study of how capabilities evolve with model size for autoregressive transformers in the protein domain: we evaluate RITA models in next amino acid prediction, zero-shot fitness, and enzyme function prediction, showing benefits from increased scale. We release the RITA models openly, to the benefit of the research community.
Forward citations
Cited by 4 Pith papers
-
Directed Evolution of Proteins via Bayesian Optimization in Embedding Space
A Gaussian-process Bayesian optimizer running in protein language model embedding space outperforms regression-based directed evolution baselines on two in silico fitness landscapes.
-
PDFBench: A Benchmark for De novo Protein Design from Function
The paper presents PDFBench, a unified benchmark with 16 metrics and a new post-2025 protein test set, and finds that evaluation choices such as retrieval strategy or supported keywords can dominate model rankings.
-
Scaling and Data Saturation in Protein Language Models
Protein language models trained on increasingly large UniRef100 snapshots show non-monotonic fitness prediction performance, with no evidence of data saturation.
-
Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
A review maps how NLP architectures from word2vec to Evo 2 are applied to DNA, RNA, protein, and genome sequences, with tokenization and context choices shaping performance.
Discussion (0). Continue with ORCID to comment.