BPE, WordPiece, and SentencePiece segment protein sequences differently, all struggle to preserve protein domain boundaries, and the apparent deviations from linguistic laws may be tokenizer artifacts.
Proteinbert: a universal deep-learning model of protein sequence and fun ction,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods
BPE, WordPiece, and SentencePiece segment protein sequences differently, all struggle to preserve protein domain boundaries, and the apparent deviations from linguistic laws may be tokenizer artifacts.