{"id":"bc01f917-b5ee-4ebb-a363-708b1abb7814","arxiv_id":"2506.14796","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A comprehensive benchmark of 17 protein foundation models across 38 tasks yields task correlations, a streamlined protocol, and identifies ProTrek as the strongest general performer.","lead":"PFMBench is a new benchmark that tests 17 protein foundation models on 38 tasks, from gene annotation to structure prediction. It also proposes a simplified evaluation protocol based on task correlations, and finds that ProTrek, a model using sequence and function clues, beats the widely used ESM2 on most tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ProTrek's apparent advantage may be largely label leakage from its function-annotation pretraining; the 75% winning rate is the load-bearing claim that needs a leakage-controlled recomputation.","rationale":"The reader's formal weakest assumption is the representativeness of the core task/model selection, but the rationale also flags ProTrek label leakage. I see label leakage as the more load-bearing concern because it directly threatens the headline claim that ProTrek consistently outperforms ESM2, which in turn supports the protocol and the scaling discussion. The paper itself acknowledges the leakage risk in the Table 7 caption without testing whether the conclusion survives when overlapping tasks are removed. This is not an internal inconsistency, but it is a correctness risk in the central empirical ranking. The benchmark resource and code remain valuable, and the claimed correlations and protocol may still be useful, so rejection is not warranted. The existing CONDITIONAL verdict remains appropriate: the ProTrek-vs-ESM2 ranking should not be treated as definitive until a leakage-controlled analysis is produced. My read therefore does not change the reader's verdict, hence UNCHANGED.","tokens_in":19586,"tokens_out":3239,"duration_ms":34038,"concrete_test":"Recompute Table 3 winning rates against ESM2 after removing every representative task whose labels are derived from the same annotation databases used in ProTrek pretraining (at minimum Enzyme Commission, Metal Ion Binding, and DeepLoc2 Multi; ideally all Swiss-Prot/GO-derived labels) and report ProTrek's #Win on the remaining structure, solubility, and production tasks. Additionally, run a sequence-level overlap check between ProTrek's pretraining corpus and the test splits of all 11 representative tasks; if overlap exceeds, say, 5% for any task, rerun with decontaminated splits. If ProTrek's #Win drops from 75% to at or below ESM2's, the central ranking claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical result is that ProTrek consistently outperforms ESM2 (75% winning rate on representative tasks, Section 4.1) and that contrastive pretraining on function annotations is the reason. But ProTrek is pretrained on Swiss-Prot/GO functional annotations, and PFMBench's representative tasks include Enzyme Commission, Metal Ion Binding, and DeepLoc2 Multi, whose labels are drawn from the same functional annotation sources. Table 7's own caption concedes that potential label leakage from overlapping functional annotation data remains a concern for function-aware models. If ProTrek's advantage is concentrated in annotation-overlapping tasks, then the model-ranking headline, the 'pretraining strategies matter more' scaling conclusion in Section 4.3, and the protocol's recommendation to use ProTrek as a default baseline are all built on an artifact. The paper does not provide any overlap analysis or leakage-controlled experiment; it only appends a caveat while retaining the claims.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PFMBench, a benchmark for protein foundation models spanning 38 tasks (annotation, solubility, localization, mutation, interaction, structure, production, and zero-shot ProteinGym) and 17 models, with standardized data splits, a set of PEFT methods, and a hierarchical analysis pipeline. The authors report that core-task and core-model selection lead to 11 representative tasks and 12 core models, that ProTrek achieves a 75% winning rate against ESM2 on representative tasks, that zero-shot ProteinGym performance does not correlate with supervised performance, and that scaling ESM2 only helps at 15B parameters while pretraining strategy matters more. The paper also provides code and a modular framework.","tokens_in":19835,"tokens_out":2931,"duration_ms":31015,"significance":"If the central claims hold, PFMBench would be a valuable community resource: it offers a broad task/model coverage, standardized 30% sequence-similarity splits, a reproducible Hydra/PyTorch Lightning framework, a comparison of several PEFT methods, and a concrete streamlined evaluation protocol. The analysis of task correlations and the MSA mutual-information diagnostic are useful additions beyond a simple leaderboard. However, the headline model-ranking result depends on a possible label-leakage confound that the manuscript itself flags in Appendix A.2, and the selection of core tasks/models is post hoc relative to the main conclusions. These issues need to be addressed before the benchmark's recommendations can be relied upon.","major_comments":[{"comment":"The central claim that ProTrek outperforms ESM2 with a 75% winning rate is not established because of a likely label-leakage confound. ProTrek is pretrained on Swiss-Prot/GO functional annotations, and the representative tasks include EC, GO BP, GO MF, GO CC, Metal Ion Binding, and DeepLoc2-Multi, whose labels come from the same or overlapping functional-annotation sources. The manuscript's own Table 7 caption concedes that 'potential label leakage from overlapping functional annotation data remains a concern for function-aware models,' yet no overlap analysis or leakage-controlled experiment is presented. I request a stratified analysis separating annotation-overlapping tasks from non-overlapping tasks, reporting ProTrek's win rate in each stratum, and, if feasible, a leakage-controlled evaluation where training labels overlapping ProTrek's pretraining annotations are removed.","section":"§4.1, Table 3; Appendix A.2, Table 7"},{"comment":"The core-task and core-model selection criteria are post hoc and can bias the conclusions. Core tasks are selected by requiring ESM2-Adapter run-to-run bias below 5%, and core models are selected by requiring EC F1 at least 85% of ESM2's score. Because these filters are applied before computing the task-correlation matrix, the representative tasks, and the model rankings, the reported relationships are conditional on the choice of ESM2 as the reference and on the adapter protocol. Please provide a sensitivity analysis that, for example, includes the 10 excluded tasks or relaxes the EC threshold, to show that the task clusters, the 11 representative tasks, and the ProTrek-versus-ESM2 ranking are robust to these choices.","section":"§3.2 and §3.3"},{"comment":"Almost all model-task results are single runs with no error bars or significance tests, while the paper's own task-bias filter (Table 1) shows that 10 of 38 tasks have ESM2-Adapter run-to-run bias above 5%, some as high as 114%. Many pairwise model differences in Table 3 are below 0.01 in F1 or Spearman, so the #Win rates and the 'ProTrek consistently outperforms ESM2' conclusions may reflect noise. Please report results over at least three seeds with standard errors for the representative tasks, or provide a paired significance test for the winning-rate comparisons.","section":"Tables 3, 5, 6, and 7"},{"comment":"The use of AF2DB or ESMFold predicted structures for structure-aware models, rather than experimentally determined structures, is a protocol choice that can interact with model pretraining and task difficulty, but the paper does not analyze this dependence. For example, SaProt is trained with a structure-aware vocabulary derived from predicted structures, which may give it an advantage or disadvantage on tasks evaluated with predicted structures. Please state how many tasks use predicted versus experimental structures and include a sensitivity check on a subset where experimental structures are available.","section":"§3.1"}],"minor_comments":[{"comment":"There are several typos and inconsistencies, including 'SaPort' for SaProt in Table 2 and Figure 1, 'ProtoT5' for ProtT5 in §4.2, 'foucus' in §2, 'adpot' in §2, and 'enumerious' in the Figure 3 caption.","section":"Throughout"},{"comment":"The task-correlation matrix and the model-ranking figure are difficult to read at the printed resolution; please increase the figure size and font, and consider providing a zoomable version or a table of the underlying correlations.","section":"Figure 4 and Figure 7"},{"comment":"The mutual-information difference metric depends on the aligned overlapping regions and on the masking procedure, but the text does not specify how gaps and length differences are handled in the alignment or how many MSA clusters are used for the figure; please clarify these implementation details.","section":"Appendix A.3"},{"comment":"The hyperparameter section states that the optimizer is AdamW with batch size 64 and up to 50 epochs, but it does not specify the learning-rate scheduler, warmup, weight decay, or the random seed policy across tasks; please provide these details for reproducibility.","section":"§3.4"},{"comment":"The task table lists mean performance and bias for ESM2-Adapter but does not report the number of evaluation seeds used to compute the bias; please state that explicitly, since the 5% core-task threshold depends on it.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a disclosure or discussion of the institutional relationship between PFMBench and ProTrek, given that both originate from the same research groups. This is not a claim of misconduct, but it strengthens the case for an independent or leakage-controlled verification of the central ranking result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jeremy—\n\nPFMBench is a genuinely large and useful resource: 38 tasks, 17 models, hundreds of supervised runs, plus a task-correlation analysis that reduces the set to 11 representative tasks and a PEFT comparison. That is real work and the code is out. Anyone building an evaluation stack for protein foundation models will want to borrow from it.\n\nBut the paper's central claim—that ProTrek consistently beats ESM2, with a 75% winning rate—should not be taken at face value. ProTrek was pretrained on functional annotations (Swiss-Prot/GO), and several of the representative tasks (EC, GO terms, Metal Ion Binding, DeepLoc2) draw labels from the same annotation space. The paper's own Table 7 caption concedes \"potential label leakage from overlapping functional annotation data remains a concern.\" That is not a minor caveat: the model-ranking headline, the \"pretraining strategies matter more\" scaling conclusion, and the protocol's recommendation to use ProTrek as a default baseline all depend on ProTrek's advantage being real and not an artifact of seeing the answers. The paper does not provide an overlap analysis or a leakage-controlled experiment. Until that is done, the main empirical conclusions are conditional at best.\n\nThere are softer problems too. The core task selection (bias <5% on ESM2-Adapter) and core model selection (EC >=85% of ESM2) are post hoc, so the subsequent correlation and ranking analysis risks circularity. Most results have no error bars; with 12 models on 28 tasks, the differences between the middle-ranked models are probably within noise. And the authors cite ProteinBench but never compare against it, which is odd for a benchmark paper. Also, ProTrek comes from the same research group (Westlake/BioMap), and the institutional overlap is not acknowledged.\n\nWhat is solid: the scale, the modular interface, the PEFT comparison across adapter/LoRA/DoRA/IA3, and the reminder that ProteinGym zero-shot performance does not track supervised performance. That last point is worth checking, but it may hold even after leakage control.\n\nWho is this for? Groups building practical evaluation pipelines, and anyone trying to place new protein models. It deserves a serious referee, but the referee should demand a leakage-controlled reanalysis of ProTrek, error bars, and a direct comparison with ProteinBench. I would not let the model rankings of the current version be treated as definitive.\n\nBest,\n[You]","headline":"A valuable and reusable benchmark, but its main ProTrek ranking needs a leakage-controlled redo before it can be believed.","tokens_in":20321,"tokens_out":2657,"would_cite":true,"duration_ms":25937,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PFMBench argues that ProTrek beats ESM2 on 75% of representative protein tasks, that zero-shot ProteinGym scores do not predict supervised performance, and that scaling ESM2 only pays off at 15B parameters.","keywords":["protein foundation models","benchmark","parameter-efficient fine-tuning","zero-shot evaluation","task correlation","multimodal protein models","model scaling"],"falsifier":"Run the same adapter-based protocol on a held-out set of protein tasks not used for clustering or core-task selection and compare ProTrek's winning rate against ESM2; if it drops to or below 50%, the representativeness claim fails.","tokens_in":19413,"feed_emoji":"🧬","tokens_out":11297,"duration_ms":100345,"temperature":0.7,"pith_summary":"PFMBench argues that protein foundation models can be compared fairly in a single unified suite, and that two widespread evaluation beliefs are wrong. The benchmark spans 38 tasks in eight areas and 17 models, then filters to 12 core models and 28 low-variance core tasks and clusters them into 11 representative tasks. On those tasks, the multimodal model ProTrek beats sequence-only ESM2 on 75% of tasks, while zero-shot ProteinGym rank and supervised rank do not match. The paper also reports that ESM2 scaling is flat until 15B parameters, so pretraining strategy matters more than raw scale.","feed_headline":"ProTrek beats ESM2 on 75% of protein benchmark tasks","feed_subtitle":"A new 38-task benchmark finds zero-shot ProteinGym scores do not predict supervised performance.","key_machinery":"The carrying machinery is a two-stage selection plus task-correlation protocol. ESM2-Adapter is run three times on each task, and only the 28 tasks with under 5% run-to-run bias are kept; models that reach at least 85% of ESM2's EC F1 score become the 12 core models. Spearman correlations between tasks are clustered into 11 groups, one representative task per group, and model strength is summarized by a winning rate (#Win), the share of representative tasks where a model exceeds ESM2. The same protocol is crossed with six PEFT methods and with a mutual-information-difference analysis relative to ESM2-35M to connect the rankings to pretraining behavior.","core_discovery":"The paper's discovery is a scaled, structured evaluation landscape that changes what should be reported about protein foundation models. Under a fixed adapter-tuning protocol, the 28 core tasks correlate into 11 clusters, and model rankings across those clusters show that sequence-only encoders rarely beat ESM2, decoder-only models perform worst, and multimodal models with contrastive alignment lead, with ProTrek winning 75% of representative tasks. A second finding is that ProteinGym zero-shot scores do not correlate with supervised results, so zero-shot fitness benchmarks are not a proxy for general protein understanding. A third finding is that the ESM2 scaling curve only improves at the 15B size and at disproportionate cost, whereas ProTrek-650M beats ESM2-15B on most tasks.","pith_inferences":["A direct decontamination test that removes benchmark-overlapping sequences from ProTrek's pretraining data would settle whether its annotation-task advantage is semantic alignment or memory; the paper itself flags label leakage as a concern.","The divergence between ProteinGym and supervised rankings suggests fitness prediction and general protein understanding are separate capabilities, so protein engineering pipelines may need two separate model selections.","The scaling conclusion is specific to ESM2's pretraining recipe; scaling a contrastive model like ProTrek from 650M to 3B or 15B and comparing cost-adjusted gains would test whether 'scaling is not worth it' generalizes.","The 11-task subset's representativeness can be validated externally by deriving task clusters on an independent set of protein datasets and checking whether the ProTrek-versus-ESM2 ranking reproduces."],"forward_implications":["A new protein foundation model can be evaluated meaningfully on the 11 representative tasks instead of an exhaustive 38-task sweep, cutting the cost of fair comparison.","Zero-shot ProteinGym results should be reported separately from supervised results, because they rank models differently and cannot be used interchangeably for model selection.","Multimodal, contrastively aligned models emerge as the strongest current direction, while decoder-only generative models are poor defaults for protein understanding tasks.","Scaling parameter count in the ESM2 family is not a reliable route to better downstream performance until the 15B scale, so improving pretraining data and objectives is the cheaper lever.","Adapter tuning is sufficient as a default PEFT protocol, with DoRA as a competitive alternative, so future benchmarks can standardize on a single efficient tuning method."],"supporting_citations":[{"why":"It supplies the ESM2 baseline, the sequence-only comparison point, and the ESM2 scaling series used for the 15B finding.","marker":"[35]"},{"why":"It is the ProTrek model whose 75% winning rate over ESM2 is the paper's principal positive result.","marker":"[56]"},{"why":"It defines the ProteinGym zero-shot benchmark whose rank order is shown not to correlate with supervised performance.","marker":"[47]"},{"why":"It is a large multimodal ESM3 model whose comparatively low ranking supports the function-data quality conclusion.","marker":"[18]"},{"why":"It is the structure-aware SaProt baseline that anchors the sequence-structure comparison and the stability split validation.","marker":"[55]"},{"why":"It is the prior VenusFactory benchmark and VenusPLM model source that PFMBench extends and compares against.","marker":"[59]"},{"why":"It supplies the TAPE stability and fluorescence tasks and an earlier evaluation standard for transfer learning.","marker":"[52]"},{"why":"It is the multi-task PEER benchmark that PFMBench extends in scope and model coverage.","marker":"[72]"},{"why":"It supplies the adapter tuning method used in every supervised evaluation and in the task-bias filter.","marker":"[23]"}],"fun_headline_variants":["ProTrek tops new 38-task protein benchmark, beating ESM2","Zero-shot scores mislead protein model ranking","ProTrek-650M beats ESM2-15B on new protein benchmark","New benchmark: 38 tasks reveal ProTrek leads, zero-shot fails","ProTrek wins 75% of tasks in new protein benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions depend on the assumption that the 12 core models and 28 core tasks chosen by the filtering rules represent the wider protein foundation model landscape, so that rankings and recommendations transfer to other models and tasks.","fun_headline_variants_meta":{"raw":{"variants":["ProTrek tops new 38-task protein benchmark, beating ESM2","Zero-shot scores mislead protein model ranking","ProTrek-650M beats ESM2-15B on new protein benchmark","New benchmark: 38 tasks reveal ProTrek leads, zero-shot fails","ProTrek wins 75% of tasks in new protein benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00077,"raw_usage":{"total_tokens":3375,"prompt_tokens":877,"completion_tokens":2498,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":2404}},"tokens_in":493,"tokens_out":2498,"duration_ms":17132,"temperature":1.0,"reasoning_tokens":2404,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:55:14.329330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same adapter-based protocol on a held-out set of protein tasks not used for clustering or core-task selection and compare ProTrek's winning rate against ESM2; if it drops to or below 50%, the representativeness claim fails.","supporting_citations":[{"cited_title":"Proteingym: Large- scale benchmarks for protein fitness prediction and design.Advances in Neural Information Processing Systems, 36:64331–64379, 2023","cited_arxiv_id":null,"evidence_quote":"It defines the ProteinGym zero-shot benchmark whose rank order is shown not to correlate with supervised performance."},{"cited_title":"Saprot: Protein language modeling with structure-aware vocabulary","cited_arxiv_id":null,"evidence_quote":"It is the structure-aware SaProt baseline that anchors the sequence-structure comparison and the stability split validation."},{"cited_title":"Peer: a comprehensive and multi-task benchmark for protein sequence understanding.Advances in Neural Information Processing Systems, 35:35156–35173, 2022","cited_arxiv_id":null,"evidence_quote":"It is the multi-task PEER benchmark that PFMBench extends in scope and model coverage."}],"review_version":1}