{"id":"1aad89c7-dfc8-4db6-b342-f7787a5f8448","arxiv_id":"2608.12090","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across 13 protein language models and 15 downstream tasks, the last layer is rarely the best source of embeddings; intermediate layers work better, and the best layer is set by dataset structure rather than by the task alone.","lead":"Protein language models are usually read at their last layer, but this study shows that intermediate layers often carry the most useful information for predicting protein properties. The authors map 13 models across 15 tasks and find that the best layer depends on whether the dataset is built from many mutants of one protein or from many diverse natural proteins.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'dataset, not task' claim rests on within-dataset task pairs whose labels are nested or co-derived, so the dichotomy remains confounded with label type.","rationale":"The reader's weakest assumption correctly identifies the central flaw. The paper's headline distinction between DMS and diverse-protein datasets is confounded with task/label type, and the only within-dataset evidence intended to break that confound uses label pairs that are nested or co-derived: Fluorescence GMM thresholding, nested DeepLoc and SSP labels, and correlated Meltome species/Tm. This matters because the abstract's strongest claim is precisely that the dataset, not the task, determines the layer profile; if that is unsupported, the paper's main novelty reduces to the well-supported but weaker observations (last layer rarely best, residue-level monotonic increase, artificial-protein degradation). I do not think the flaw is fatal: a dedicated experiment with independent label pairs could validate the dichotomy, and the practical guidance to use intermediate layers is robust. Hence the verdict stays CONDITIONAL. The Tsuboyama dataset being both multi-species and classified as 'centered around one protein' in §3 adds a smaller internal inconsistency that the proposed test would also clarify. Overall, the reader's conditional assessment is appropriate, and my read does not change it.","tokens_in":21679,"tokens_out":7654,"duration_ms":67972,"concrete_test":"Replace the four task pairs in Figure 3A-D with pairs whose labels are not nested or derived from the same measurement, keeping protein sets and splits fixed: e.g., DeepLoc2.0 10-class localization vs. an independent binary signal-peptide label; Meltome Tm residualized against species vs. species class; Fluorescence regression vs. a non-thresholded experimental property of the same GFP mutants; SCOPe 3-class/8-class SSP vs. solvent-accessible/buried residue labels. Recompute the mean layer-performance curve correlations across the 13 PLMs. If the ~0.85+ correlations do not reproduce for independent label pairs, the 'dataset, not task' dichotomy is confounded by label dependence; if they persist, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central dichotomy—'dataset, not task'—is supported almost entirely by the within-dataset task-pair comparisons in Figure 3A-D, but the four pairs are not independent label sources. The Fluorescence binary task is produced by fitting a GMM to the same log-fluorescence values used in the regression task and dropping the interval between the two 2-sigma boundaries (§4.3, Fig. 6); it is a thresholded coarsening of the regression label. The Meltome Tm and species tasks are acknowledged to be 'closely related' [42], and species identity is strongly associated with growth temperature. DeepLoc binary membrane vs. 10-class localization are nested labels (membrane status is one of the localization classes), and the SCOPe 3-class vs. 8-class secondary-structure tasks are nested DSSP annotations. In every pair, the two tasks share a large fraction of label information, so the high curve correlations (0.68–0.996) can be explained by label correlation rather than by dataset identity. The across-dataset comparison does not resolve this: DMS datasets come with fitness/fluorescence labels, while diverse-protein datasets come with localization/solubility/fold labels, so 'dataset' and 'task/label type' are fully confounded. The paper's own interpretation in §3—early layers capture local fine details, deeper layers general protein rules—is itself a task/label-structure explanation. The dichotomy is also blurred by the Tsuboyama stability set (64 species, 63 clusters; Table 3), which is labeled 'centered around one or several proteins' in §3 despite being a multi-protein DMS set. Unless the similarity survives genuinely independent task pairs on the same sequences, the strongest claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic probing study of 13 protein language models (PLMs) across 15 downstream tasks from 11 datasets. For each model and task, the authors train linear probes on per-layer embeddings, measure layer-wise performance, and supplement these measurements with three latent-space metrics (intrinsic dimension, variance@10, and neighborhood overlap). The main empirical findings are: (i) the deepest layer is rarely the best embedding source (reported in 17.92% of cases), with performance usually peaking in intermediate layers for protein-level tasks; (ii) residue-level tasks show a steady performance increase with depth, which the authors attribute to alignment with the masked-language-model pretraining objective; (iii) for whole-protein tasks, the paper claims that the dataset, rather than the task, determines the layer profile, with DMS-centered datasets favoring shallow layers and diverse-protein datasets favoring deeper layers; and (iv) performance on artificially generated proteins is markedly worse than on natural proteins. The paper also includes fine-tuning experiments showing that fine-tuning makes layer-wise performance monotonic toward the fine-tuning objective, and sparse-data experiments showing that 15-20% of training data often suffices to identify a near-best layer.","tokens_in":21890,"tokens_out":3151,"duration_ms":32352,"significance":"If the central findings hold, they are practically and scientifically useful: practitioners would be guided away from the default last-layer embedding and toward layer selection based on dataset structure, and the paper adds to the growing evidence that PLM representations are not uniformly organized by task relevance across depth. The study is unusually broad in model and task coverage, and the authors have made code and processed data available, which strengthens reproducibility. The paper is also careful in several secondary analyses, such as reporting three-seed stability for the sparse-probe layer selection and using DataSAIL for stratified splits. However, the headline distinction between dataset-driven and task-driven layer profiles is not yet established to journal standard, because the within-dataset task pairs used to separate these factors share heavily correlated labels, and the across-dataset comparison confounds dataset identity with label type. The central dichotomy therefore needs additional evidence or a more carefully scoped claim before the paper can be accepted.","major_comments":[{"comment":"The claim that the layer-performance profile depends on the dataset, not the task, rests on the four within-dataset task pairs in Figure 3A-D, but these pairs do not provide independent label information. In §4.3 and Figure 6, the Fluorescence binary task is created by fitting a GMM to the same log-fluorescence values used for regression and discarding the interval between the two 2-sigma boundaries, so it is a thresholded coarsening of the regression label. The Meltome species and Tm tasks are acknowledged as closely related [42]; the DeepLoc binary membrane and 10-class localization labels are nested; and the SCOPe 3-class and 8-class secondary-structure labels are nested DSSP assignments. The high curve correlations (0.681-0.996) can therefore be explained by label correlation rather than by dataset identity. The paper should either provide task pairs with orthogonal labels on the same protein set or explicitly acknowledge and statistically control for the shared label information before claiming independence from objective, split, and difficulty.","section":"§2.1 and Figure 3A-D; §3"},{"comment":"The grouping of the Tsuboyama stability dataset with the DMS-based, shallow-layer-favoring datasets is internally inconsistent. Section 3 describes these datasets as 'centered around one or several proteins and comprise many variants of them,' but Table 3 reports that the Tsuboyama set comprises 64 species, 63 CD-HIT clusters, and a mean sequence identity of 0.2192, which is far closer to the diverse-protein datasets than to Fluorescence or GB1 (which have 1 cluster each). This undermines the proposed dichotomy between 'DMS datasets' and 'diverse natural protein datasets.' The authors should either reclassify Tsuboyama, demonstrate that its multi-species DMS structure behaves like single-protein DMS for the relevant layer metrics, or otherwise revise the explanatory claim in Section 3.","section":"§3 and Table 3"},{"comment":"The central quantitative summary that the last layer is best in only 17.92% of cases is reported without error bars, confidence intervals, or per-seed variance, and Figure 1B,C show only means over models. Because the best-layer selection is a discrete claim and the performance curves are close in some regions, the 17.92% figure could be sensitive to probe initialization, data splits, or other minor variations. The authors report three seeds only for the sparse-probe experiments in Figure 3E-H; the same rigor should be applied to the main best-layer analysis, or a sensitivity analysis should be provided. This is a load-bearing summary statistic for the first main claim.","section":"Figure 1 and Table 1"}],"minor_comments":[{"comment":"The label '2-NN Intrindic Dimension' contains a typo; it should read '2-NN Intrinsic Dimension.'","section":"Supplementary Figures 2 and 3"},{"comment":"The legend states that Figure 1B and C show the mean over the 13 PLMs, but no dispersion measure is provided; adding per-layer standard deviations or model-specific faint curves would improve interpretability.","section":"Figure 1"},{"comment":"Several entries are listed as NaN for ESM-2 3B, ProGen2-large, and ProtGPT2 on residue-level tasks; the table would be clearer if a footnote explained that these models were excluded from residue-level experiments, consistent with §4.4.","section":"Table 1"},{"comment":"The statement that the fine-tuned models 'show a clear effect' in Figure 4B-G is supported by visual inspection, but the main text does not report the magnitude of the improvement or a statistical comparison; a numeric summary would strengthen the claim.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and the empirical scope is commendable, but the dataset-versus-task dichotomy is the paper's main conceptual contribution and it is currently supported by confounded comparisons. I believe the authors can address this with additional experiments or by substantially softening the claim, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nRead the PLM layer probing paper. Bottom line: the core observation is solid and useful, the central interpretation is oversold.\n\nWhat's new: systematic comparison across 13 PLMs and 15 tasks, with consistent evidence that the last layer is rarely the best (under 18% of cases), that performance peaks mid-network for whole-protein tasks, and that residue-level tasks improve monotonically. The layer-selection recipe—15-20% of data finds a near-best layer—is practically valuable and reproducible, with code/data available. The artificial-protein degradation finding is interesting, though based on two generation pipelines.\n\nThe soft spot is exactly where the stress-test lands. The 'dataset, not task' dichotomy is supported by pairs of tasks on the same dataset, but in each pair the labels share a lot of information: fluorescence classification is a thresholded coarsening of the same log-values used in regression (and the interval between GMM 2-sigma bounds is dropped); DeepLoc binary membrane is nested in the 10-class localization; SCOPe 3-class and 8-class SSP are nested DSSP; Meltome Tm and species are acknowledged as closely related. So high curve correlations within a dataset can be explained by label correlation, not by dataset identity. And across datasets, 'dataset' is confounded with label/objective type: DMS sets have fitness-style continuous labels, diverse sets have localization/fold/solubility labels. The paper's own explanation—early layers capture local detail, deeper layers general rules—is itself a label-structure story. To establish dataset-not-task you'd need independent task pairs on identical sequences, or at least pairs where the labels are not co-derived. The current evidence is compatible with a weaker claim: layer trend tracks the granularity/local vs global nature of the label.\n\nOther items are minor by comparison: Figure 1 reports means without per-seed variance, though the subsampling experiments used three seeds and report consistent layer choice; that helps but doesn't cover the main curves. The Tsuboyama set is a multi-protein DMS, which slightly blurs the DMS-vs-diverse dichotomy but isn't fatal.\n\nWho should read it: anyone using PLM embeddings as a drop-in feature source. The practical layer-selection guidance is worth having, and the empirical map is a useful reference. The conceptual claim needs revision before it's cited for the dichotomy.\n\nI'd send it to review—it's a solid empirical contribution with data and code—but I'd push the authors to either add independent task pairs or soften the dataset-not-task claim to 'label structure matters.' Deserves referee time.","headline":"A broad, well-built empirical map of layer-wise PLM embeddings, but the 'dataset, not task' claim is not yet disentangled from label type.","tokens_in":22563,"tokens_out":2305,"would_cite":true,"duration_ms":19918,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For protein language models, the most useful embeddings for a downstream task usually sit in an intermediate layer, and which layer that is depends more on the dataset than on the task itself.","keywords":["protein language models","layer probing","embeddings","downstream tasks","deep mutational scanning","latent space analysis","protein property prediction","transfer learning"],"falsifier":"Re-run the paired-task comparison with label decorrelation: for fluorescence, move the binary boundary or include the omitted middle interval and check whether classification and regression layer curves diverge; for Meltome, predict species after regressing out temperature (or vice versa). If the curves separate under either intervention, the shared dataset alone is not controlling the layer profile.","tokens_in":21414,"feed_emoji":"🧬","tokens_out":9749,"duration_ms":85962,"temperature":0.7,"pith_summary":"This paper asks where, across the layers of a protein language model, the most useful information for a downstream prediction task actually lives. Probing 13 PLMs on 15 tasks from 11 datasets, it finds that the conventional choice, the model's final layer, is the best layer only about 18% of the time; performance typically rises over early layers, peaks mid-network, and falls again near the output. The central claim is that for whole-protein tasks the dataset, not the task itself, decides where this peak sits: datasets built from deep mutational scans of one or few proteins are best served by shallow embeddings, while datasets spanning many diverse natural proteins are best served by deeper ones. For residue-level tasks, performance rises monotonically toward the last layer, matching the models' token-level pre-training objective. The authors also report a sharp drop in prediction quality when the inputs are artificial, designed proteins rather than natural sequences.","feed_headline":"Final layer is best only 18% of the time in protein models","feed_subtitle":"Layer choice hinges on the dataset: mutation scans favor early layers, diverse proteins favor deeper ones.","key_machinery":"The central object is the per-layer probe: a linear classifier or regressor (plus k-nearest-neighbor probes as a nonlinear check) trained on mean-pooled embeddings from each layer, producing a performance-versus-depth curve for every model-task pair. These curves are compared across paired tasks on the same dataset (fluorescence regression vs. binary classification, Meltome temperature vs. species, DeepLoc binary vs. 10-class, SCOPe 3-class vs. 8-class secondary structure) to isolate dataset effects from task effects. Three latent-space metrics — intrinsic dimension estimated by TwoNN, neighborhood overlap between consecutive layers, and variance explained by the first ten principal components — corroborate the probe results. Fine-tuning ESM-2 150M on six tasks and re-probing all layers, plus evaluating the original pre-training losses per layer, supplies the mechanistic link between pre-training objective and layer trends.","core_discovery":"Across 13 protein language models (five families, four architectures) and 15 downstream tasks, the paper establishes that the informativeness of embeddings varies non-monotonically with layer depth for whole-protein tasks: probe performance improves over the first layers, peaks between roughly the 10th and 90th percentile of depth, and declines in the deepest layers, so the last layer is best in only 17.92% of cases. The location of the peak is controlled by dataset composition, not by whether the task is regression or classification, binary or multi-class: variants-heavy deep mutational scanning datasets (fluorescence, GB1, parts of Rocklin, Tsuboyama) peak in shallow layers, whereas diverse multi-protein datasets (DeepLoc2.0, DeepSol, SCOPe40) peak in deeper layers. Paired tasks on the same dataset produce highly correlated layer-performance curves, which the paper reads as evidence that dataset identity, rather than task objective or difficulty, governs the layer profile. Residue-level tasks instead improve monotonically as layers deepen, which the authors attribute to alignment with the masked-language-model or next-token-prediction pre-training objectives. Fine-tuning a PLM on a downstream task makes every layer improve toward that task while degrading the pre-training objective, and performance on artificial proteins (Rosetta designs and ProGen-generated lysozymes) is markedly worse than on natural ones.","pith_inferences":["An open extension: predicting mutational effects at the residue level on the same DMS datasets would show whether the shallow-layer preference comes from local sequence context rather than from how whole-protein embeddings are pooled; the paper does not run this control.","If the artificial-protein result generalizes, PLM-based scoring functions in generative design pipelines will tend to rank natural-like designs above equally functional novel sequences; this bias can be quantified by comparing model scores against measured activity on existing Rocklin and lysozyme data.","The authors' data also imply that pretraining data composition, not architecture alone, controls where task information sits: comparing models with identical architectures but different training corpora (e.g., ESM-2 versus ESMC) on the same datasets would make that explicit.","A neighboring problem that should inherit this result is remote homology and function search with embeddings, where layer choice is usually fixed; the same 15-20% subsampling trick could identify the best layer for retrieval tasks."],"forward_implications":["In most protein-level tasks, last-layer embeddings are outperformed by embeddings from an intermediate layer; the best layer usually sits between the 10th and 90th percentile of model depth, and in some model-task pairs the relative improvement is several-fold.","For residue-level tasks such as secondary-structure and binding-site prediction, the last layer is typically the right choice because these tasks align with the token-level pre-training objective.","Dataset composition, not task type, determines the best layer for whole-protein tasks: deep-mutational-scan data favor shallow embeddings, while diverse natural protein sets favor deep embeddings.","Fine-tuning a PLM on a downstream task makes layers monotonically better for that task, so the non-monotonic profiles of frozen models are a signature of the pre-training objective rather than a fixed property of the model.","Only 15-20% of the training data is needed to identify a layer that achieves at least 95% of the best layer's performance, making layer selection feasible in practice."],"supporting_citations":[{"why":"Supplies the linear-probe method used to measure how informative each layer's embeddings are.","marker":"[6]"},{"why":"Provides the intrinsic-dimension and neighborhood-overlap metrics and the observation that the first and last layers of trained transformers resemble each other.","marker":"[10]"},{"why":"The ESM-2 family is the main set of bidirectionally trained PLMs probed across all 15 tasks.","marker":"[15]"},{"why":"The fluorescence deep mutational scan supplies the clearest shallow-layer-peak example.","marker":"[19]"},{"why":"GB1, another single-protein DMS dataset, independently shows the same shallow-layer peak.","marker":"[20]"},{"why":"The Rocklin stability dataset contributes both DMS variants and de novo designed proteins, anchoring the artificial-protein comparison.","marker":"[21]"},{"why":"Provides the ProGen-generated lysozymes whose activity becomes much harder to predict than natural lysozymes.","marker":"[23]"},{"why":"The Meltome Atlas yields the paired temperature-regression and species-classification tasks used to separate dataset from task effects.","marker":"[24]"},{"why":"DeepLoc2.0 gives paired binary and 10-class tasks on a diverse set of natural proteins.","marker":"[26]"},{"why":"SCOPe40 supplies both diverse protein-level tasks (fold, superfamily) and residue-level secondary-structure tasks.","marker":"[27]"}],"fun_headline_variants":["Best embedding layer varies by dataset, not task","Last layer only best 18% of time in protein models","Mutation scans favor early layers; diverse proteins deep","Protein models struggle on artificial proteins","Residue tasks deepen; whole-protein tasks peak midway"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset-not-task conclusion requires that the two tasks paired on each dataset are effectively independent experiments; if correlated labels — such as the GMM-derived fluorescence classes and their parent regression values, or the closely related Meltome species and temperature labels — actually drive the similar layer curves, then the dataset identity conclusion loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["Best embedding layer varies by dataset, not task","Last layer only best 18% of time in protein models","Mutation scans favor early layers; diverse proteins deep","Protein models struggle on artificial proteins","Residue tasks deepen; whole-protein tasks peak midway"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1699,"prompt_tokens":1111,"completion_tokens":588,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":727,"completion_tokens_details":{"reasoning_tokens":515}},"tokens_in":727,"tokens_out":588,"duration_ms":5516,"temperature":1.0,"reasoning_tokens":515,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:17:04.442006+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the paired-task comparison with label decorrelation: for fluorescence, move the binary boundary or include the omitted middle interval and check whether classification and regression layer curves diverge; for Meltome, predict species after regressing out temperature (or vice versa). If the curves separate under either intervention, the shared dataset alone is not controlling the layer profile.","supporting_citations":[{"cited_title":"The geometry of hidden representations of large transformer models.Ad- vances in Neural Information Processing Sys- tems, 36:51234–51252, 2023","cited_arxiv_id":null,"evidence_quote":"Provides the intrinsic-dimension and neighborhood-overlap metrics and the observation that the first and last layers of trained transformers resemble each other."},{"cited_title":"Evolutionary-scale prediction of atomic-level protein structure with a language model.Sci- ence, 379(6637):1123–1130, 2023","cited_arxiv_id":null,"evidence_quote":"The ESM-2 family is the main set of bidirectionally trained PLMs probed across all 15 tasks."},{"cited_title":"Local fitness land- scape of the green fluorescent protein.Nature, 533(7603):397–401, 2016","cited_arxiv_id":null,"evidence_quote":"The fluorescence deep mutational scan supplies the clearest shallow-layer-peak example."},{"cited_title":"A comprehensive biophysical description of pairwise epistasis throughout an entire protein domain.Current biology, 24(22):2643–2651, 2014","cited_arxiv_id":null,"evidence_quote":"GB1, another single-protein DMS dataset, independently shows the same shallow-layer peak."},{"cited_title":"Global analysis of protein folding using mas- sively parallel design, synthesis, and testing.Sci- ence, 357(6347):168–175, 2017","cited_arxiv_id":null,"evidence_quote":"The Rocklin stability dataset contributes both DMS variants and de novo designed proteins, anchoring the artificial-protein comparison."},{"cited_title":"Large language models generate functional protein se- quences across diverse families.Nature biotech- nology, 41(8):1099–1106, 2023","cited_arxiv_id":null,"evidence_quote":"Provides the ProGen-generated lysozymes whose activity becomes much harder to predict than natural lysozymes."},{"cited_title":"Meltome atlas—thermal proteome stability across the tree of life.Nature methods, 17(5):495–503, 2020","cited_arxiv_id":null,"evidence_quote":"The Meltome Atlas yields the paired temperature-regression and species-classification tasks used to separate dataset from task effects."},{"cited_title":"Deeploc 2.0: multi-label subcellular localization prediction using protein language models.Nucleic acids research, 50(W1):W228–W234, 2022","cited_arxiv_id":null,"evidence_quote":"DeepLoc2.0 gives paired binary and 10-class tasks on a diverse set of natural proteins."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SCOPe40 supplies both diverse protein-level tasks (fold, superfamily) and residue-level secondary-structure tasks."}],"review_version":1}