Pith. sign in

REVIEW 1 cited by

PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17415 v1 pith:VI6YR47B submitted 2023-10-26 cs.CL cs.AIq-bio.BM

classification cs.CLcs.AIq-bio.BM
keywords languagemodelproteinmodelssizesvocabularydatasetsdownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large protein language models are adept at capturing the underlying evolutionary information in primary structures, offering significant practical value for protein engineering. Compared to natural language models, protein amino acid sequences have a smaller data volume and a limited combinatorial space. Choosing an appropriate vocabulary size to optimize the pre-trained model is a pivotal issue. Moreover, despite the wealth of benchmarks and studies in the natural language community, there remains a lack of a comprehensive benchmark for systematically evaluating protein language model quality. Given these challenges, PETA trained language models with 14 different vocabulary sizes under three tokenization methods. It conducted thousands of tests on 33 diverse downstream datasets to assess the models' transfer learning capabilities, incorporating two classification heads and three random seeds to mitigate potential biases. Extensive experiments indicate that vocabulary sizes between 50 and 200 optimize the model, whereas sizes exceeding 800 detrimentally affect the model's representational performance. Our code, model weights and datasets are available at https://github.com/ginnm/ProteinPretraining.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods

    cs.CL 2024-11 reject novelty 6.0 of 10

    BPE, WordPiece, and SentencePiece segment protein sequences differently, all struggle to preserve protein domain boundaries, and the apparent deviations from linguistic laws may be tokenizer artifacts.

Pith tools