{"id":"299306ef-03f5-4bb7-845c-c6e321f936be","arxiv_id":"2502.06881","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey paper that catalogs protein language models, their architectures, training data, benchmarks, and tools, but lacks a systematic methodology and contains several factual errors.","lead":"This paper is a review that organizes protein language models by architecture, position encoding, datasets, benchmarks, and applications. It is a starting point for researchers new to the field, but it contains accuracy issues and is not a systematic review despite its title.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Systematic claim undercut: no documented selection methodology and concrete factual errors, e.g., ESM-1b listed as Pfam-trained though the primary source uses UniRef50.","rationale":"The reader's verdict is CONDITIONAL, and this stress-test independently supports that conclusion. The strongest load-bearing condition for a survey is the accuracy and reproducibility of its catalog. The absence of any documented selection methodology already makes the 'systematic' claim unfalsifiable. Adding demonstrable factual errors, such as the ESM-1b pretraining dataset misattribution in Section V.A.1, transforms that concern from a methodological preference into a concrete correctness failure. A review of this kind is meant to be a reliable reference; when a flagship model's training data is mischaracterized, the resource's core value is undermined. The paper can be repaired via errata and a documented methodology, so rejection is not necessary; the existing CONDITIONAL verdict appropriately signals that revision is required. I agree with the reader's overall assessment but would emphasize that the accuracy failures are currently the more decisive issue, since they are directly checkable and already evident from the text.","tokens_in":25017,"tokens_out":6786,"duration_ms":65314,"concrete_test":"Read the methods section of Rives et al. 2021 (ref [5]) and confirm the pretraining dataset. If it states UniRef50, then Section V.A.1 of the review is factually incorrect, providing a direct counterexample to the accuracy assumption behind the 'comprehensive' claim. Extend the same check to the 'Pretraining Dataset' column for all entries in Tables I–IV against their primary sources; if even one flagship model entry is mislabeled, the review's central reference value fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to provide a 'systematic' and 'comprehensive' review of protein language models. This claim depends on the catalog being accurate and the literature selection being reproducible. Neither condition is met. No search strategy, database list, inclusion/exclusion criteria, screening process, or search date is provided anywhere, so the 'systematic' label is not verifiable. More concretely, the catalog itself contains demonstrable errors. In Section V.A.1, the authors write that 'Models like ProteinLM, TAPE, and ESM-1b have been trained on Pfam' (citing refs [5], [47], [103]). The cited ESM-1b paper (Rives et al., PNAS 2021) reports pretraining on UniRef50, not Pfam. This is a factual mischaracterization of a flagship model that a reader would rely on when selecting training data. Other similar issues appear, e.g., Section VI.C cites ProteinMPNN [173] in a paragraph on RoseTTAFold applications. Because the core deliverable of this review is an accurate, comprehensive catalog, these errors undercut the central claim and the paper should not be treated as an authoritative systematic reference without correction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript claims to provide a systematic review of protein language models (PLMs) from a macro perspective. It covers model architectures (non-transformer-based and transformer-based, subdivided into encoder-only, decoder-only, and encoder-decoder), positional encoding strategies, scaling laws, pre-training and benchmark datasets, downstream applications (structure prediction, function prediction, protein design, and mutation prediction), and several widely used computational tools. The paper also presents four model tables (Tables I-IV), two dataset/benchmark tables (Tables V-VI), and a GitHub repository collecting resources related to PLMs, datasets, and tools.","tokens_in":25266,"tokens_out":5874,"duration_ms":49116,"significance":"The planned organization of the review—grouping models by architecture, position encoding, scaling behavior, and datasets—is reasonable, and the compiled tables and companion GitHub repository are practical resources for researchers entering the field. If the catalog entries were accurate, the paper would be a useful broad survey. However, the value of a review of this type rests entirely on the correctness of its factual and citation details, and the manuscript currently contains several errors that affect the reliability of its central deliverable. The absence of a documented selection methodology also weakens the 'systematic' claim. These issues are fixable, but they require substantive revision rather than copy-editing.","major_comments":[{"comment":"The sentence 'Models like ProteinLM[47], TAPE[103], and ESM-1b[5] have been trained on Pfam' is factually incorrect for ESM-1b, whose pretraining corpus is UniRef50 according to the cited paper [5]. Please correct this entry in the catalog.","section":"Section V.A.1, Pfam paragraph"},{"comment":"The text states that ProtTrans 'trains various autoencoder models (e.g., BERT, ALBERT, ELECTRA)', but BERT, ALBERT, and ELECTRA are not autoencoders; they are transformer encoders trained with masked language modeling or replaced-token detection. Please revise the architecture terminology.","section":"Section II.B.2, ProtTrans description"},{"comment":"References [122] and [124] are the same work (Zhang et al., arXiv:2203.06125), yet [122] is cited in the EC paragraph and [124] in the GO paragraph. The cited work is a structure pretraining paper, not the EC or GO benchmark source; please replace these citations with the correct benchmark references and eliminate the duplicate.","section":"Section V.B.2 and References [122]/[124]"},{"comment":"The paper labels itself a 'systematic' and 'comprehensive' review, but no review methodology is documented anywhere in the manuscript. Please add a methods subsection stating the search strategy, databases consulted, inclusion/exclusion criteria, and the date the literature search was performed, so that the 'systematic' claim is verifiable.","section":"Section I and Abstract"},{"comment":"The sentence about RoseTTAFold cites [173–175], but [173] is the ProteinMPNN paper (Dauparas et al., 2022) and is not a RoseTTAFold reference; please adjust the citation span or add a dedicated RoseTTAFold citation.","section":"Section VI.C, protein design paragraph"}],"minor_comments":[{"comment":"[92] is a duplicate of [35] (both are 'Alec Radford. Improving language understanding by generative pre-training. 2018'); please remove the duplicate and renumber.","section":"Reference list"},{"comment":"The phrase 'Joshua's findings' is informal and unclear; please replace with 'the findings of Meier et al.' or similar.","section":"Section VI.D"},{"comment":"The sentence 'GO annotations cover multiple species and play a positive role in cross-species gene function and evolutionary research' would read better as 'GO annotations cover multiple species and support cross-species gene function and evolutionary research'.","section":"Section V.B.2, GO paragraph"},{"comment":"In the sentence 'Learned positional encoding often leads to better downstream performance for protein language models, as adopted by numerous PLMs like ESM-1b[5], and others[46, 64, 69]', the comma after '[5]' is misplaced; consider removing it.","section":"Section III.A"},{"comment":"The ESM-C row lists Time '2024.12' and a checkmark under Code; please verify the source and date, and clarify whether the code is publicly available.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a useful curated overview but is not, in its current form, a systematic review in the methodological sense. The factual and citation errors noted above affect the reliability of the catalog. If the authors correct these errors and add a methodology paragraph, the paper could be acceptable; otherwise the 'comprehensive' claim is not supportable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This review is best read as a broad map of the protein language model landscape, not as a systematic synthesis. The taxonomy—architecture, position encoding, scaling, datasets, benchmarks, applications, tools—is sensible, and the model tables are a handy quick reference, especially with code links. The GitHub resource list is a genuine asset, and the coverage of recent models like ESM-3, xTrimoPGLM, and ProSST is current. People new to the field will get oriented here faster than from the primary literature alone.\n\nThat said, the central claim falls short in two ways. First, there is no methodology supporting the word “systematic.” No search strategy, inclusion/exclusion criteria, or search date is provided, so the comprehensiveness is unverifiable. Second, the catalog has concrete factual errors. The most visible is the text claiming ESM-1b was trained on Pfam (Section V.A.1), when Table II correctly lists UniRef50 and the Rives et al. paper is about UniRef50. That is exactly the kind of detail a reader relies on when choosing a model or dataset. There are also duplicate references ([122]/[124] are the same paper cited for EC and GO; [92] duplicates [35]) and a loose use of “autoencoder” for BERT/ALBERT/ELECTRA. None of these are fatal, but in a review, accuracy is the product.\n\nThe broad strokes hold up: the field does organize into these families, the scaling discussion is reasonable, and the applications survey is serviceable. The soft spots are localized rather than load-bearing. The paper could be a good entry point after a careful revision that fixes the factual errors, harmonizes the tables with the text, and either documents a systematic protocol or stops calling itself systematic.\n\nWho will get value? Graduate students and interdisciplinary researchers wanting a starting point, not experts seeking authoritative statements. I would send it to peer review because a corrected version is genuinely useful, and the problems are eminently fixable. My recommendation: major revision, with a referee asked to check every table entry against the cited primary source.","headline":"A useful but uneven PLM catalog whose 'systematic' label is not backed by methodology and whose factual errors are fixable—worth a round of major revision before it can serve as an entry point.","tokens_in":25753,"tokens_out":2735,"would_cite":false,"duration_ms":30498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper provides a systematic review of protein language models across architectures, position encoding, scaling laws, datasets, benchmarks, applications, and tools.","keywords":["protein language models","survey","protein sequence representation","Transformer architectures","position encoding","scaling laws","pretraining datasets","downstream applications"],"falsifier":"An independent audit of the tables and the companion resource collection against the cited primary papers would settle it: if a substantial fraction of entries give wrong parameter counts, wrong pretraining datasets, or dead or missing code links, or if well-known models published before the cutoff are absent, the claim of a systematic and accurate review fails.","tokens_in":24862,"feed_emoji":"🧬","tokens_out":10149,"duration_ms":87917,"temperature":0.7,"pith_summary":"Protein language models treat amino-acid sequences as text and learn representations from large unlabeled sequence databases. The paper argues that the field's growth has outpaced its overviews, with existing reviews confined to narrow subdomains, and it sets out to supply a macro-level map. The map is organized around the factors the authors say determine model behavior: architecture (from early non-transformer embeddings through encoder-only, decoder-only, and encoder-decoder transformers), position encoding, scaling laws, and pretraining data. The same framework carries through to benchmarks, downstream applications, tools, and a curated resource collection. A sympathetic reader comes away with a structured route into the field and its open problems.","feed_headline":"A survey maps protein language models from architecture to benchmarks","feed_subtitle":"It covers architectures, datasets, benchmarks, tools, and challenges in one curated resource hub.","key_machinery":"The organizing machinery is a four-way architectural taxonomy of protein language models, used as the spine of the review. Around that spine the authors place three further lenses: position encoding (absolute, relative, and rotary variants), scaling laws linking model size, data, and compute to performance, and pretraining dataset choice. The taxonomy groups the field's models so that representative examples can be compared on the same axes, and the benchmark and tool sections give the reader the instruments the field actually uses to judge them.","core_discovery":"The central claim, stated in the introduction, is that a systematic review of protein language models from a macro perspective is now possible and needed, and that this paper delivers it. The authors organize the field into non-transformer and transformer-based architectures, split the latter into encoder-only, decoder-only, and encoder-decoder families, and treat position encoding, scaling laws, and pretraining datasets as the axes along which model behavior and comparison should be understood. They further claim that evaluation is inseparable from downstream tasks, so they tie benchmarks to structure, function, mutation, and design applications. The closing argument identifies MSA-free structure prediction and multimodal sequence-structure-function models as the current mainstream directions. If the paper is right, the field now has a usable taxonomy and a curated resource repository as entry points.","pith_inferences":["The review collects the ingredients for a controlled comparison but does not run one: a testable extension would be to train or fine-tune models from the repository on matched data and compute to isolate the effect of architecture choice from dataset choice.","If the paper's claim that protein language models are more prone to underfitting than NLP models is correct, then the scaling-law section implies that modest increases in training tokens may yield larger gains than increases in parameters; the paper does not itself quantify this.","The review's focus on language models means structure-aware models are included mainly when they fuse a language model with structural tokens; a companion review of purely geometric models would be needed to see the full protein AI landscape.","The data-quality-versus-data-quantity debate noted in the discussion could be sharpened into a benchmark: compare models pretrained on redundancy-clustered versus metagenome-diverse data at fixed compute."],"forward_implications":["A newcomer can use the paper's taxonomy and resource collection to pick models, datasets, and benchmarks for a task without reconstructing the field's history from scattered papers.","The juxtaposition of scaling-law evidence and claims that protein language models are often undertrained implies that larger-scale training should continue to improve performance, and that the field's largest models are not yet at the point of diminishing returns.","The discussion of position encodings identifies rotary and relative encodings as the choices that handle long sequences best, which gives model builders a concrete design default.","The benchmark survey gives standard evaluation suites such as TAPE, PEER, ProteinGym, CASP, CAMEO, FLIP, and CAFA a single reference point, making cross-paper comparisons easier to situate.","The identified trends toward MSA-free and multimodal models point to where near-term progress is most likely to be concentrated."],"supporting_citations":[{"why":"Supplies the Transformer architecture that most protein language models build on, and defines the position-encoding problem.","marker":"[1]"},{"why":"The ESM-1b result that scaling unsupervised learning yields biological structure and function; anchors the encoder-only and scaling-law discussions.","marker":"[5]"},{"why":"ProtTrans, the large encoder-family comparison across BERT-style objectives on protein data.","marker":"[6]"},{"why":"ESM-2 and ESMFold, the basis for the MSA-free structure-prediction trend and for evidence on rotary position encoding.","marker":"[40]"},{"why":"AMPLIFY, the data-quality-versus-scale study that frames the dataset-choice discussion.","marker":"[90]"},{"why":"AlphaFold2, the MSA-based structure prediction baseline and a source of structural training data for multimodal models.","marker":"[99]"},{"why":"TAPE, one of the standard benchmark suites used to evaluate protein embeddings.","marker":"[103]"},{"why":"ProteinGym, the mutation-effect benchmark used heavily in zero-shot mutation prediction.","marker":"[134]"}],"fun_headline_variants":["Protein language models, demystified: architecture, data, and applications","One review to navigate the protein language model landscape","Protein LM review: from architectures to downstream tasks","The macro view of protein language models: tools and trends","Protein language models: a systematic look at the whole field"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole overview stands on the authors having chosen the right literature and described each model, dataset, and tool correctly; if their selection is skewed or their table entries misstate the cited papers, the review loses its usefulness.","fun_headline_variants_meta":{"raw":{"variants":["Protein language models, demystified: architecture, data, and applications","One review to navigate the protein language model landscape","Protein LM review: from architectures to downstream tasks","The macro view of protein language models: tools and trends","Protein language models: a systematic look at the whole field"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000817,"raw_usage":{"total_tokens":3519,"prompt_tokens":825,"completion_tokens":2694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":2614}},"tokens_in":441,"tokens_out":2694,"duration_ms":21940,"temperature":1.0,"reasoning_tokens":2614,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:39:11.669854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit of the tables and the companion resource collection against the cited primary papers would settle it: if a substantial fraction of entries give wrong parameter counts, wrong pretraining datasets, or dead or missing code links, or if well-known models published before the cutoff are absent, the claim of a systematic and accurate review fails.","supporting_citations":[],"review_version":1}