{"id":"1928b011-5251-4ad1-b0e0-305e5ffc5b49","arxiv_id":"2607.16087","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Scaled Gaussian Convolution, a deterministic scaling-plus-smoothing of AlphaFold2's weights, reveals a ubiquitin conformational landscape whose contact-loss order, flexibility pattern, and funnel topology match experiments, while KaiB and α-synuclein show the encoding's limits.","lead":"The study applies a deterministic perturbation — multiplying AlphaFold2's internal weights by 0.75 and lightly blurring them — and shows the perturbed model produces ordered protein shapes for ubiquitin whose contact-loss sequence, flexibility pattern, and funnel topology match experimental and simulated references. It proposes that structure-prediction networks can be read directly as records of protein conformational organization, and introduces 'neural spectroscopy' as the","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All quantitative results sit in the reduced-MSA regime (64×64); at 128×128 sensitivity collapses (atomization 3.0%→0.2%) and default 512×5,120 behavior is untested, so the readout may reflect under-constrained-input floppiness rather than encoded conformational landscapes.","rationale":"The reader's weakest_assumption identified exactly this regime premise, and the reader's rationale listed it as item (3). I agree it is the most load-bearing concern: it underpins all quantitative results, whereas other flagged issues (block-mismatch confound, OpenFold framing, code availability) affect narrower claims or are acknowledged by the author. The paper is otherwise strong: matched-power noise controls, five-model replication, external MD benchmarks, and explicit epistemic status. But the regime premise is not merely a minor limitation—it is the condition under which every headline number is produced. Since the author explicitly leaves default-MSA behavior as an open question, the conditional verdict already reflects the need for this test; raising it does not change the verdict. I therefore recommend UNCHANGED: the paper should remain CONDITIONAL pending a default-MSA control.","tokens_in":51193,"tokens_out":8334,"duration_ms":81105,"concrete_test":"Run SGC on ubiquitin (model_1_ptm) at default MSA depth 512×5,120 with (σ=0.30, λ=0.75) for depths {4,8,12,16,20,24,28,32,40,48}, 3 seeds. Compute max RMSD, min Q, the G1/G2/G3 crossing depths, and partial r|WCN at depths 4–24. If max RMSD < 1.5 Å and/or the G2-first/G1-last ordering and r|WCN ≥ 0.3 do not appear, the reduced-MSA regime is load-bearing for the headline claim. If the response is measurable, repeat with λ=0.5 and λ=0.6 to see whether the ordering and correlations recover before atomization; if they never recover, the claim that SGC reads encoded conformational landscapes from AlphaFold2 weights does not extend beyond the artificial regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AlphaFold2's weights encode a readable conformational landscape—is tested entirely in the reduced-MSA regime (64×64, Section 2.3). The author's rationale is that at reduced depth 'perturbing the Evoformer weights has a measurable effect' (Section 2.3), and sensitivity falls monotonically with depth: atomization at λ=0.75 drops from 3.0% at 64×64 to 0.2% at 128×128 (Supplementary Section S11). Behavior at the default 512×5,120 depth is untested, with the author stating 'Behaviour at default MSA depth... remains an open question' (Section 2.3; also §4.4). The risk is not merely that the effect is smaller; it is that the structured readout—contact-loss ordering (Section 3.5), flexibility correlations (Section 3.6), and landscape topology (Section 3.4)—is a product of an under-constrained transformer. With the MSA weakened, the model is generically floppy, so weight perturbation may be amplifying input ambiguity rather than reading out a landscape genuinely encoded in the weights. Because every quantitative headline result lives in this regime, the claim about AlphaFold2's weights is currently unsupported in the model's intended operating regime. The author's controls—ordering preserved at 32/64/128, noise controls, five-model replication—mitigate but do not eliminate this concern; they do not include default MSA.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Scaled Gaussian Convolution (SGC), a deterministic perturbation of AlphaFold2/OpenFold Evoformer weights, and argues that the resulting structural responses reveal a conformational landscape encoded in the learned weights. Using ubiquitin as the main test system, the author reports that native contacts break in an experimentally established stability ordering (G2 before G3, G1 last), that per-residue perturbation sensitivity correlates with microsecond MD RMSF beyond geometric baselines, that the pooled perturbed ensemble reproduces the L-shaped folding funnel topology of millisecond MD, and that matched-power noise controls do not produce such structured responses. The results are further supported by five-model replication and by application to KaiB (convergent non-recovery of the alternative fold) and α-synuclein (structured inter-model disagreement). The paper is candid about several limitations, especially the use of reduced MSA depth (64×64) in all primary analyses.","tokens_in":51550,"tokens_out":4331,"duration_ms":43052,"significance":"If the central claim holds, the paper is significant: it would show that a structure-prediction model trained only on static targets nonetheless encodes a readable, physically meaningful ordering of conformational organization, and it introduces a generally applicable perturbation-based interpretability protocol. The external validation strategy is a genuine strength: the primary comparisons are against DE Shaw MD trajectories, Went/Jackson and Sosnick folding experiments, and five separately trained OpenFold replicas, with no parameter fitted to reproduce the target orderings. The noise controls and MSA-depth comparisons are also thoughtful. However, the significance is conditional on the reduced-MSA issue described below; the paper itself labels default-depth behavior an open question, which tempers the abstract's claims.","major_comments":[{"comment":"All quantitative headline results (contact-loss ordering in §3.5, RMSF correlations in §3.6, landscape topology in §3.4) are produced at MSA depth 64×64. The paper notes that at 128×128 atomization falls from 3.0% to 0.2% and that default 512×5,120 behavior is untested. This is load-bearing: the central claim is that the Evoformer weights encode a conformational landscape, but the readout is only demonstrated in a regime where the coevolutionary input is deliberately weakened. The author's cross-depth controls (32/64/128) mitigate but do not eliminate the concern that the structured response is an artifact of an under-constrained transformer rather than a faithful readout of the weights. A control at default MSA depth—possibly with larger perturbation strength—or a demonstration that the 64×64 response is a monotone amplification of a default-depth signal is required to support the abstr","section":"§2.3, §4.4, Supplementary S11"},{"comment":"The depth axis conflates 'more blocks corrupted' with 'more early-stage computation warped.' The paper explicitly invokes perturbed/unperturbed block mismatch to explain tail recovery (Section 3.1), yet interprets the depth-resolved contact-loss ordering (Section 3.5) as differential encoding robustness. Cumulative forward perturbation changes both the number of corrupted blocks and the location of the perturbed/unperturbed boundary; the strong forward/reverse asymmetry reported in Section 3.1 shows that boundary effects are large. Without single-block or windowed perturbation controls, the G2-before-G3 ordering could reflect the specific position of the boundary rather than a stability ordering encoded in the weights. The paper defers such controls to future work, but this gap is central to the interpretation of Section 3.5.","section":"§3.1, §3.5"},{"comment":"The flexibility correlation claim is pattern-level, not magnitude-level, in the regime where it is strong. At depths 4–24 the partial r|WCN is 0.66–0.79, but Lin's CCC is about 0.08 and perturbation sensitivity is about 18-fold smaller than MD RMSF. At depths 27–31 the magnitudes converge (CCC=0.66) but the structures are partially unfolded and outside the native basin sampled by the 300 K MD. The paper states this tradeoff honestly, but the conclusion that 'Evoformer weights carry a residue-level constraint signal correlated with physical flexibility' rests on a reduced-MSA, pattern-only correlation in the shallow regime and an unfolded-ensemble magnitude agreement in the deep regime. A more direct test—e.g., comparing SGC sensitivity on native-basin structures at higher MSA depth—would strengthen the causal claim.","section":"§3.6, Table 2"},{"comment":"The α-synuclein 'shared-core' prediction is not sharply falsifiable as stated. The shared-core region in (Rg, Ree) space is defined by the overlap of the five AF2 models and the pooled MD ensemble, and the prediction is that any future experimental ensemble with sufficient sampling will occupy this region. Because the region is partly defined by the MD reference, the prediction risks circularity. To make it testable, the region should be defined from the five-model intersection alone (without MD), and the prediction should specify a quantitative occupancy threshold (e.g., a measurable fraction of experimental frames) and a sampling criterion.","section":"§3.11, §4.2"}],"minor_comments":[{"comment":"The abstract states that 'the conformational organization visible under perturbation... emerged as a byproduct' without noting that all quantitative support comes from the reduced-MSA regime and that default-MSA behavior is untested. A one-sentence qualification would align the abstract with Section 4.4.","section":"Abstract / §2.3"},{"comment":"Depth-1 is described as an anomaly in Section 3.8 but is not consistently excluded in the summary metrics; the exclusions (depths ≤5 versus depths ≤3) should be stated in each figure/table caption.","section":"§3.1 / §3.2"},{"comment":"Cross-references to 'main Fig. 13' in the supplementary gallery should be updated to the correct main-text figure numbers (the KaiB landscape appears to be Figure 14).","section":"Supplementary S8 / main text §3.10"},{"comment":"The statement that code and processed tables are 'available from the author on request' is insufficient for a study of this scale and for the reproducibility claims made in the text. At minimum, the code, YAML manifests, and processed metric tables should be deposited in a public repository.","section":"Code and data availability"},{"comment":"The sentence 'all structures at depths 2 and 4–24 maintain a fully connected polypeptide chain' is contradicted by the 18% of runs with broken bonds reported in the same section. Clarify that the 18% refers to deeper depths and the phrase 'excluding... anomalies' should be explicit here as well.","section":"§3.7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is intellectually interesting and unusually honest about its limitations, but the central claim is currently conditional on the reduced-MSA regime. The missing default-MSA control is the single most important issue; the depth-axis conflation and the α-synuclein prediction sharpening are secondary but also load-bearing. If the author can demonstrate that the structured response persists at default MSA depth (or show that the 64×64 response is a faithful amplification of it), the paper would be a strong contribution. I would not reject, but the current version is not yet ready for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It's a genuinely careful interpretability paper—the author did the controls most papers skip—but the headline claim overreaches the evidence in one specific, load-bearing way: every quantitative result lives at reduced MSA depth 64×64, chosen precisely because it makes the model more sensitive. At 128×128 the effect mostly vanishes (atomization 3.0%→0.2%), and default 512×5,120 is never tested. So the statement 'AlphaFold2's weights encode a conformational landscape' is currently supported only for an under-constrained, out-of-distribution input regime. It may be true; it is not yet shown.\n\nWhat is genuinely new: a deterministic scale-and-blur perturbation (SGC) of Evoformer weights yields graded structural responses; contact groups break in the experimentally established order across 84 informative conditions; per-residue perturbation sensitivity tracks 154.6 μs MD RMSF with partial r|WCN around 0.7; landscape topology matches 390 K MD funnel shape; five independently retrained OpenFold models converge for ubiquitin and disagree for α-synuclein; matched-power white/spectral noise produce debris or bimodal fates, not structured progressions. The external validations use published MD and folding data with no fitted parameters. That is real evidence, and the author is refreshingly explicit about what is measurement and what is interpretation.\n\nSoft spots, in order: (1) The MSA regime. The author's own Supplementary shows sensitivity falls monotonically with MSA depth, and the comparison across depths covers one protein and 7 of 192 conditions. This is the central claim's load-bearing weakness; it needs either default-depth experiments or a clear argument that the reduced-input regime is the right probe for weight-encoded structure. (2) The cumulative depth axis. 'Depth d' means first d blocks corrupted, so the perturbed/unperturbed mismatch is itself a variable; the author invokes it to explain tail recovery but treats the rest of the profile as weight-encoded. Single-block profiling would separate block-local fragility from cumulative artifacts. (3) The deterministic block-2 collapse at depth 3 is excluded from summaries rather than explained. (4) It's OpenFold, not AlphaFold2 proper, and code/data are on request only—minor but real.\n\nWho it's for: anyone doing AlphaFold interpretability, perturbation-based ensemble sampling, or using AF2 for dynamics hypotheses. It deserves a serious referee; I'd push for major revision with the MSA question and single-block control addressed. Not desk-reject.","headline":"A serious, candid interpretability study with real external validations, but its central 'encoded landscape' claim rests entirely on a reduced-MSA regime and is not yet established.","tokens_in":52132,"tokens_out":3127,"would_cite":true,"duration_ms":30852,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AlphaFold2's trained weights carry a readable record of protein conformational organization — stability orderings, flexibility patterns, funnel topology — that emerges under perturbation despite never being a training target.","keywords":["neural spectroscopy","AlphaFold2","Scaled Gaussian Convolution","weight perturbation","conformational landscapes","protein dynamics","Evoformer","interpretability"],"falsifier":"Repeat the SGC sweep at the default MSA depth (512×5,120) with scaling factors well below the tested λ range: if no perturbation strength reproduces ubiquitin's contact-breaking order (G2 first, G1 last), the RMSF correlations, and the funnel topology, then the physical signal is an artifact of the weakened-input regime rather than an encoding in the weights. The paper's own proposed proteome-scale screen offers the complementary check: structured landscapes should track training-data density if the account is memorization, or extend beyond explicit coverage if it is generalized conformational","tokens_in":50915,"feed_emoji":"🧬","tokens_out":10451,"duration_ms":86529,"temperature":0.7,"pith_summary":"This paper argues that AlphaFold2's 93 million trained parameters are not just machinery for turning sequence into structure: they encode a map of protein conformational organization that ordinary inference never expresses. The author's probe — scaled Gaussian convolution, which smooths each Evoformer weight tensor with a narrow Gaussian, scales it down, and applies this cumulatively to the first d of the 48 blocks — produces graded structural responses that track known physics. Ubiquitin's native contacts break in the experimentally established stability order (intermediate-boundary contacts first in all informative conditions, folding-nucleus contacts last), per-residue perturbation sensitivity correlates with microsecond-scale molecular dynamics flexibility beyond what atomic burial predicts, and the 912,384-structure ensemble reproduces the funnel topology of millisecond folding simulations. Matched-power random noise destroys the signal, and five independently trained models converge on the same landscape, arguing the organization is a property of the training, not of one weight realization. If the claim holds, conformational information is present, at least partially, inside a predictor trained only on static structures — a route to dynamics-like information without simulation.","feed_headline":"Perturbing AlphaFold2's weights exposes hidden protein landscapes","feed_subtitle":"Its weights trace ubiquitin's real folding order and flexibility — readable without running any simulation.","key_machinery":"Scaled Gaussian Convolution (SGC): each weight tensor W in Evoformer blocks 0 through d−1 is replaced by W′ = λ·G_σ(W), where G_σ is Gaussian smoothing with width σ and λ is a uniform scaling factor — at the primary operating point σ = 0.30, λ = 0.75 — and d (perturbation depth) controls how many of the 48 sequential blocks are modified. Scaling dominates (more than 99% of the perturbation power) and acts as a near-uniform attenuation of every dot-product in the network, shifting information routing smoothly; the sub-1% smoothing residual has a component orthogonal to uniform scaling that selectively destabilizes specific contact groups near the transition boundary. The matched noise control","core_discovery":"The central claim is that AlphaFold2's Evoformer weights encode physically structured conformational organization as a byproduct of the structure-prediction objective, and that this encoding can be read directly by deterministic weight deformation. Under SGC at the primary operating point, ubiquitin's native contacts lose coherence in the order established by folding experiments: in all 84 informative perturbation conditions the intermediate-boundary contacts break first, and in 67 of 84 the folding-nucleus contacts outlast every other group (ties in the remaining 17). Per-residue perturbation sensitivity tracks a 154.6-microsecond equilibrium simulation with partial correlation 0.66–0.79 af","pith_inferences":["My extension: if the encoding is in the weights rather than the weakened input, the same contact-order and flexibility signals should be recoverable at the default MSA depth with stronger attenuation (λ below the 0.70–0.85 grid tested here); the paper leaves behavior at default depth untested, and that experiment would localize the signal.","My extension: because perturbation sensitivity matches flexibility patterns but not magnitudes, the probe could serve as a simulation-free per-residue flexibility prior for fold classes well represented in the training corpus — worth benchmarking against crystallographic B-factors and NMR order parameters on a wider panel than the three proteins studied.","My extension: the five divergent α-synuclein landscapes generate a concrete, testable hypothesis the author leaves open — that model-specific basins correspond to real functional sub-ensembles (extended helical states resembling the membrane-bound form, compact states resembling aggregation-prone forms); mapping each model's basin to a functional context would distinguish structured extrapolation ","My extension: the block-2 anomaly (catastrophic collapse when exactly the first three blocks are perturbed, absent at depths 2 and 4) and the KaiB depth-7 dip indicate the 48 Evoformer blocks are not functionally interchangeable; per-block and per-head perturbation profiles could map a coarse-to-fine division of labor and reveal where the network compensates."],"forward_implications":["If the readout is genuine, a static-structure predictor's weights encode a differential robustness hierarchy: contacts that are evolutionarily and physically most load-bearing are encoded most robustly, so the order in which perturbation dismantles native contacts is a readout of stability ordering.","The shallow-depth flexibility signal — partial correlation 0.66–0.79 beyond burial, exceeding a geometry-only network baseline — implies the weights carry a residue-level constraint signal correlated with physical flexibility, even though training never rewarded it.","The three-protein spectrum gives a general assay: strong training signal produces convergent landscapes, absent signal produces convergent absence, ambiguous signal produces structured disagreement — a way to determine, for any protein, where a network's representation is determined by data and where it is underdetermined.","The noise-control contrast establishes that the response is not generic network damage: only weight-coherent perturbation reproduces the physical signal, so the conformational information sits in the dominant, structured directions of the learned weights rather than in fine-grained random components.","Model-independence for ubiquitin (transition onset within two blocks across five models) implies the landscape is a property of the architecture plus training data, not of a single weight configuration."],"fun_headline_variants":["AlphaFold2's weights hide a map of protein conformational states","Perturbing AlphaFold2's weights uncovers real folding orders","Reading AlphaFold2's weights exposes protein shape change landscape","AlphaFold2's encoded weights reveal hidden protein flexibility"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"All primary results come from a deliberately weakened regime — MSA depth 64×64, one-eighth of AlphaFold2's default — chosen because it makes weight perturbation measurable, and the load-bearing assumption is that the physical ordering, flexibility correlations, and funnel topology arise from what the weights encode rather than from the under-supplied coevolutionary input; behavior at default depth is untested.","fun_headline_variants_meta":{"raw":{"variants":["AlphaFold2's weights hide a map of protein conformational states","Perturbing AlphaFold2's weights uncovers real folding orders","Reading AlphaFold2's weights exposes protein shape change landscape","AlphaFold2's encoded weights reveal hidden protein flexibility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000373,"raw_usage":{"total_tokens":1836,"prompt_tokens":760,"completion_tokens":1076,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":1006}},"tokens_in":504,"tokens_out":1076,"duration_ms":10203,"temperature":1.0,"reasoning_tokens":1006,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T21:24:15.332896+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the SGC sweep at the default MSA depth (512×5,120) with scaling factors well below the tested λ range: if no perturbation strength reproduces ubiquitin's contact-breaking order (G2 first, G1 last), the RMSF correlations, and the funnel topology, then the physical signal is an artifact of the weakened-input regime rather than an encoding in the weights. The paper's own proposed proteome-scale screen offers the complementary check: structured landscapes should track training-data density if the account is memorization, or extend beyond explicit coverage if it is generalized conformational","supporting_citations":[],"review_version":1}