{"id":"3f8bf8f6-7167-420d-b355-cfbeaad43d46","arxiv_id":"2505.17237","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A pipeline from sequence alignments through a Potts model to an Ising foldon chain predicts protein folding curves, subdomains, and mutation effects, but without new experimental validation in this paper.","lead":"This paper describes a computational method that uses many related protein sequences to predict how a protein folds and unfolds. It maps evolutionary patterns into a simple physics model of cooperative folding units to estimate stability, folding temperature, and the effects of mutations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Predicted folding curves and mutation effects inherit experimentally fitted T_sel, a fixed per-residue entropy, and a user-chosen foldon partition, so 'sequence alone' remains to be demonstrated by an independent benchmark.","rationale":"The reader's weakest assumption is precisely that the quantitative mapping depends on a family-level selection temperature and an empirically fitted entropy, and that foldon partitioning could create artifacts. Reading the Methods confirms this: Step 2 requires either experimental ΔΔG data or a strong assumption of family-independent ΔΔG spread, and Step 3 fixes s to a constant fitted in earlier work. The paper is self-consistent and provides a working Colab implementation, which is real evidence for reproducibility, but it does not contain new external validation of the quantitative predictions; Figure 2 shows a DHFR example, Figure 3 summarizes 7490 sequences across 15 families, and Figure 4 gives single-mutant predictions, but the comparisons against experiment cited are prior work by the same group. I therefore do not find a reason to move from CONDITIONAL to a stronger or weaker verdict. The concrete test described above would settle whether the calibration and partition choices are benign enough for the 'sequence alone' headline to stand.","tokens_in":15226,"tokens_out":3678,"duration_ms":27857,"concrete_test":"Run an out-of-family benchmark on a protein with available experimental thermal denaturation and deep-mutational scanning data (e.g., E. coli DHFR, GB1, or lambda repressor), none of which were used to fit T_sel or the entropy. Compute T_sel both from that family's experimental ΔΔG and from the PDZ-relative procedure, and repeat the pipeline under at least three foldon partitions (secondary-structure, exon-based, neutral geometric). Then compare predicted melting temperatures, cooperativity scores, and per-mutation ΔT_f/Δcooperativity against experimental thermal melts and ΔΔG values. If predictions shift by more than ~5 K in T_f or if rank correlations of mutation effects fall below statistical significance across T_sel choices or partitions, the 'sequence information alone' claim needs to be weakened to 'sequence plus family-specific calibration and user-defined foldons.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that equilibrium folding curves, subdomain emergence, and mutation effects can be predicted for any sequence from sequence information alone (abstract; Box 1). The load-bearing link is the mapping from Potts evolutionary energies to physical folding free energies in Methods Steps 2 and 3. That mapping uses two calibrated constants: the selection temperature T_sel, which is either fit to experimental ΔΔG data from at least one family member or transferred from a PDZ reference under the assumption that the standard deviation of ΔΔG is family-independent, and the per-residue entropy s = 5 cal/mol/K/res, an empirical fit from earlier repeat-protein work [29]. The Hamiltonian's exponential Boltzmann weights make predicted T_f and cooperativity directly sensitive to both constants, yet no sensitivity analysis across families is presented. In addition, foldon assignment is user-specified and the subdomain definition uses a fixed threshold |T_fj - T_fk| < 5 K, so apparent domains may be artifacts of the chosen partition. Note 1 concedes that positions under other selection forces can locally frustrate the landscape, and Note 3 concedes that for compact beta proteins a topology-only model may suffice, which narrows the scope of the sequence-information claim. Because the quantitative examples are largely drawn from the same group's previous work [21,29], the central claim is plausible but not independently demonstrated in this manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a protocol for predicting equilibrium protein folding behavior from sequence information. The method learns a Potts model from a multiple sequence alignment, maps the resulting evolutionary energy parameters to a coarse-grained Ising chain of foldons using a selection temperature and an empirical per-residue entropy, and then runs Metropolis Monte Carlo simulations to compute thermal unfolding curves, per-foldon folding temperatures, free energy profiles, a cooperativity score, and apparent subdomain emergence. It also describes how to estimate changes in folding temperature and cooperativity upon single-point mutations, and illustrates the protocol on a DHFR sequence, a ubiquitin sequence, and an analysis of 7490 sequences from 15 families.","tokens_in":15418,"tokens_out":4218,"duration_ms":29292,"significance":"If the quantitative mapping from Potts energies to folding free energies could be shown to transfer across families without per-family experimental calibration, this would be a valuable method for converting sequence alignments into testable predictions about folding stability, cooperativity, and subdomain organization. The paper is clearly written as a step-by-step protocol, ships a reproducible Colab implementation, and builds on well-established DCA methodology; the 7490-sequence phase-space analysis is also a useful descriptive resource. However, the central claim of sequence-only prediction is not yet demonstrated, because the quantitative content rests on fitted constants and a user-specified foldon partition, and no independent experimental benchmark is presented in this manuscript.","major_comments":[{"comment":"The abstract and Box 1 claim that folding curves and mutation effects are predicted from sequence information alone, but Methods Step 2 requires experimental delta-delta-G data for at least one family member to fix T_sel, or transfers a PDZ-derived T_sel under the assumption that the standard deviation of delta-delta-G is family-independent. This fitted constant multiplicatively scales all energy differences and therefore directly sets every predicted T_f and delta-delta-G; without a sensitivity analysis across families or an independent benchmark, the 'sequence alone' claim is not supported.","section":"Methods, Step 2 (and Abstract/Box 1)"},{"comment":"The per-residue entropy s = 5 cal/mol/K/res is taken from ref [29], an empirical fit for repeat proteins, and is assumed additive and sequence-independent for all foldons. Because the Boltzmann weights in the Hamiltonian make the unfolding curves exponentially sensitive to this entropy, the transfer of this fit to globular proteins with different topologies (e.g., DHFR and ubiquitin in Figures 2 and 4) needs justification; at minimum, a sensitivity analysis over a plausible range of s should be provided.","section":"Methods, Step 3"},{"comment":"The foldon partition is user-specified, and the apparent-domain definition uses a fixed threshold |T_fj - T_fk| < 5 K. The inferred subdomains and cooperativity score therefore depend on the chosen partition and threshold. The manuscript mentions a neutral-model averaging option but does not apply it, so there is no evidence that domain emergence is robust rather than an artifact of the specific foldon assignment chosen in Figures 2 and 4.","section":"Materials, Section 3, and Methods, Step 5"},{"comment":"The central claim that the method 'predicts' folding dynamics is not tested against new experimental data in this manuscript. Figures 2 and 4 are illustrative applications, and Figure 3 is a descriptive scatter of the cooperativity score across 7490 sequences with no comparison to measured folding quantities. As a protocol paper, the manuscript should either provide an experimental validation for at least one family or explicitly frame the curves as untested predictions; the current abstract overstates what is demonstrated.","section":"Figures 2, 3, and 4"},{"comment":"The title and abstract promise 'folding dynamics,' but the Monte Carlo simulation is an equilibrium Metropolis scheme that computes thermal unfolding curves and free energy profiles, not kinetic rates or pathways. The word 'dynamics' should either be replaced by 'thermodynamics' or the model must be extended to non-equilibrium simulation; as written, the terminology overstates the scope of the method.","section":"Title and Methods, Step 4"}],"minor_comments":[{"comment":"The critical-points factor and the autocorrelation analysis are described qualitatively; please give the default values and the convergence criteria used in the reported simulations so that results are reproducible.","section":"Methods, Step 4"},{"comment":"The axis labels 'energetic heterogeneity' and 'average interaction strength' are not defined quantitatively; please provide the exact formulas used to compute these quantities.","section":"Figure 3"},{"comment":"The statement that 'using a linear fit, useful predictions can be made' should cite the specific figure or Table that supports this claim and should report the fit quality (e.g., R^2 or mean absolute error).","section":"Methods, Step 6"},{"comment":"Note 3's scope restriction, that compact beta proteins do not require a sequence-sensitive Potts model, should be reflected in the abstract, since it substantially narrows the class of proteins for which the protocol is claimed necessary.","section":"Note 3 and Abstract"},{"comment":"The equations in the Methods section as provided do not render in the manuscript text, making it impossible to verify the mapping from Potts couplings to foldon energies; please ensure all equations are fully displayed.","section":"Methods, Steps 1-3"}],"recommendation":"major_revision","confidential_remarks":"This is a protocol-style paper that leans heavily on the authors' prior work (refs [21], [29]) and does not include an independent experimental validation. The central idea is plausible and the pipeline is reproducible, but the quantitative claims in the abstract are stronger than what is demonstrated. Major revision should request either a new experimental benchmark or a clear reframing as a protocol with testable predictions, along with sensitivity analyses for T_sel, the per-residue entropy, and the foldon partition. I do not see grounds for rejection, as the method is actionable and the paper provides a useful starting point for the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a clean, well-written methods protocol that applies the authors' existing Potts-plus-Ising foldon machinery to arbitrary globular proteins. The new piece is packaging: a step-by-step Box 1, practical guidance on MSA cleaning and DCA hyperparameters, heuristics for foldon assignment, and a working Colab notebook. That is genuinely useful. The survey of 7490 sequences across 15 families in Figure 3 is a nice descriptive result. I agree with the reader's conditional verdict.\n\nWhat the paper does not do is validate the abstract's strong claim. The mapping from evolutionary energy to physical free energy rests on T_sel, which is either fit to experimental ΔΔG data or transferred from a PDZ reference, and a per-residue entropy (5 cal/mol/K/res) fit in earlier repeat-protein work. That means mutation and stability predictions inherit the calibration. It's not \"sequence information alone\"; it's sequence information plus a family-level experimental anchor. The stress-test note is right about this.\n\nThe foldon partition is user-specified and the subdomain threshold (|Tf_j - Tf_k| < 5 K) is arbitrary. No sensitivity analysis is given, so the apparent domains could be artifacts of the partition. Note 1 and Note 3 further narrow the claim: other selection forces can frustrate local stability, and for compact beta proteins a topology-only model may suffice. These are honest admissions, but they undercut the title.\n\nOn the other hand, the protocol is internally consistent, the equations work, and the code is public. For a family with a reliable T_sel, the method generates concrete, testable predictions. That's worth something. The paper is a useful recipe for practitioners, but not a demonstration of predictive power. The novelty is incremental, and a serious referee should ask for an independent benchmark, e.g., predicting ΔT_f for a family where T_sel is estimated from one protein and tested on another, or a sensitivity analysis over foldon partitions.\n\nWho is this for? Someone who wants to run this kind of simulation and needs a practical manual. The paper deserves peer review as a methods paper, but the authors should be asked to temper the claims and add robustness checks.\n\nRecommendation: engage with it, invite revision, and insist on the benchmark.","headline":"A clean, useful protocol that generalizes the authors' earlier Potts-plus-Ising work, but the 'sequence alone' claim is not validated by new experiments in this paper.","tokens_in":16030,"tokens_out":3499,"would_cite":false,"duration_ms":22186,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a method that converts a multiple sequence alignment of a protein family into quantitative predictions of how any member sequence folds and unfolds.","keywords":["protein folding dynamics","Potts model","direct coupling analysis","Ising model","foldons","protein stability","folding cooperativity","multiple sequence alignment"],"falsifier":"Apply the protocol to a protein family that already has experimental thermal denaturation curves and a deep mutational scan, then compare predicted per-foldon folding temperatures, the overall unfolding curve, and mutation-induced shifts in stability and cooperativity with the measurements. A disagreement beyond calibration noise in the rank order of stabilities across homologs, or in the location of predicted subdomain boundaries, would show that the Potts-to-free-energy mapping is not generally valid.","tokens_in":14951,"feed_emoji":"🧬","tokens_out":6453,"duration_ms":47375,"temperature":0.7,"pith_summary":"This paper presents a method for turning a multiple sequence alignment of a protein family into quantitative predictions of how any member sequence folds and unfolds. The authors claim that the observed variation in homologous sequences encodes an 'evolutionary field'—a Potts model of local fields and pairwise couplings—that can be mapped onto a coarse-grained Ising chain in which each cooperative folding unit (a foldon) is a two-state spin. With a family-level selection temperature $T_{\\mathrm{sel}}$ and an empirical per-residue entropy, Monte Carlo simulation of this chain yields equilibrium thermal unfolding curves, per-foldon folding temperatures, free-energy profiles, cooperativity scores, and the emergence of subdomains. The same machinery predicts how single-point mutations change stability and cooperativity, giving a sequence-based proxy for a deep-mutational scan. If the method is right, it fills the gap left by structure predictors, which show where a protein ends up but not how it gets there.","feed_headline":"One sequence alignment predicts how a protein unfolds","feed_subtitle":"Potts energies and an Ising foldon model yield stability, subdomains, and mutation effects with no structure input.","key_machinery":"The load-bearing object is a finite-chain Ising model of foldons. Each foldon $j$ carries a two-state spin (folded or unfolded), an internal folding free energy $\\epsilon_j^i$, and surface interactions $\\epsilon_{jk}^s$ with other folded foldons, plus an entropic penalty $s_j = L_j s$ with $s = 5\\,\\mathrm{cal\\,mol^{-1}\\,K^{-1}\\,res^{-1}}$; all energetic terms are computed by summing Potts couplings and fields inside and between foldons for the target sequence and scaling by $k_B T_{\\mathrm{sel}}$. The Potts model, learned from the MSA by DCA-type inference, provides the evolutionary energy field, and $T_{\\mathrm{sel}}$ converts evolutionary energy differences into physical free-energy differences. Metropolis Monte Carlo then samples the chain across temperatures, from which thermal unfolding curves, free-energy profiles, folding temperatures, subdomain boundaries, and cooperativity scores are read off.","core_discovery":"The central claim is that folding dynamics, not just the native structure, are accessible from sequence information alone. For any sequence belonging to a family with a sufficiently deep and clean alignment, the equilibrium folding curve can be computed: the protein is divided into foldons, each treated as a folded/unfolded spin, with internal folding free energies $\\epsilon_j^i$ and inter-foldon surface energies $\\epsilon_{jk}^s$ derived from the Potts evolutionary energy of that sequence and rescaled by $k_B T_{\\mathrm{sel}}$. The selection temperature is obtained from experimental $\\Delta\\Delta G$ data when available, or from the ratio of standard deviations of evolutionary energy changes for single mutations relative to a reference family. Subdomains are identified when neighboring foldons have folding temperatures closer than 5 K, cooperativity is measured by the fraction of intermediate folded-count states $Q$ that are never free-energy minima, and single-mutation effects are extracted from the wild-type simulation without rerunning it. The paper thus offers an end-to-end pipeline from alignment to folding dynamics.","pith_inferences":["If the relative-scale estimate of $T_{\\mathrm{sel}}$ transfers across families, the pipeline could be applied to metagenomic or poorly characterized families with no experimental stability data at all.","Because the foldon partition is a user input, averaging predictions over neutral-model partitions (as the paper suggests) could serve as a robustness check for whether predicted subdomains are real or an artifact of the chosen boundaries.","Positions conserved for reasons other than folding stability—such as binding or catalysis—may show systematic local errors, and frustration-based analysis could flag those positions ahead of time.","Substituting protein language model logits for Potts couplings (mentioned as an alternative in the paper) would make the same Ising machinery applicable to families too shallow for DCA, at the cost of trusting the language model's energy interpretation."],"forward_implications":["For any family with a deep MSA, natural sequences can be ranked by predicted folding temperature and cooperativity, which may help select thermostable variants for biotechnology.","Mutation effects on stability and cooperativity can be predicted for all single-point mutants from one wild-type simulation, providing a sequence-only proxy for deep mutational scanning.","Sequences with more favorable total evolutionary energy are predicted to be more stable by construction, so family-level energy ranking is itself a stability predictor.","Topology is predicted to matter: elongated $\\alpha$-helical families can show sequence-dependent folding mechanisms, while compact $\\beta$-containing families are more conserved in mechanism.","Synthetic sequences sampled from the Potts field can be placed in the heterogeneity–interaction phase space to estimate cooperativity before running any folding simulation."],"supporting_citations":[{"why":"introduces Direct Coupling Analysis of residue coevolution, the source of the Potts evolutionary energy field","marker":"[19]"},{"why":"provides pseudolikelihood Potts inference (plmDCA), the recommended learning method for the energy field","marker":"[32]"},{"why":"defines the selection temperature that rescales evolutionary energy differences into physical folding free energies","marker":"[34]"},{"why":"supplies the relative-scale estimate of the selection temperature when experimental $\\Delta\\Delta G$ data are unavailable","marker":"[35]"},{"why":"provides the empirically fitted per-residue entropy and the folding-model calibration against thermal denaturation scans","marker":"[29]"},{"why":"establishes the earlier observation that natural sequence variation modulates folding dynamics and supports the mutation-effect predictions","marker":"[21]"},{"why":"introduces foldons as cooperative folding units, the coarse-grained degrees of freedom of the Ising model","marker":"[17]"},{"why":"provides evidence for the exon-foldon correspondence used in foldon assignment","marker":"[18]"},{"why":"reviews inverse statistical models of protein sequences and defines the zero-sum gauge used to set the mean random-sequence energy to zero","marker":"[33]"}],"fun_headline_variants":["Folding dynamics predicted from sequence alone","No structures needed: sequences reveal folding stability","Potts energy plus Ising foldons gives folding curves","From alignments to folding curves: sequence-only route","Sequence-only pipeline maps folding dynamics and mutations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one family-level selection temperature $T_{\\mathrm{sel}}$, together with a fixed additive per-residue entropy, converts evolutionary Potts energies into real folding free energies for every sequence of the family; if those calibrations are family-dependent, or if the user-chosen foldon partition does not match the true cooperative units, the predicted curves, subdomains, and mutation effects would be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Folding dynamics predicted from sequence alone","No structures needed: sequences reveal folding stability","Potts energy plus Ising foldons gives folding curves","From alignments to folding curves: sequence-only route","Sequence-only pipeline maps folding dynamics and mutations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2736,"prompt_tokens":941,"completion_tokens":1795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":1725}},"tokens_in":557,"tokens_out":1795,"duration_ms":11696,"temperature":1.0,"reasoning_tokens":1725,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:50:50.609996+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the protocol to a protein family that already has experimental thermal denaturation curves and a deep mutational scan, then compare predicted per-foldon folding temperatures, the overall unfolding curve, and mutation-induced shifts in stability and cooperativity with the measurements. A disagreement beyond calibration noise in the rank order of stabilities across homologs, or in the location of predicted subdomain boundaries, would show that the Potts-to-free-energy mapping is not generally valid.","supporting_citations":[{"cited_title":"As the temperature increases, the folding elements are expected to transition from the folded to the unfolded state","cited_arxiv_id":null,"evidence_quote":"introduces foldons as cooperative folding units, the coarse-grained degrees of freedom of the Ising model"},{"cited_title":"Inferring protein folding mechanisms from natural sequence diversity","cited_arxiv_id":"2412.14341","evidence_quote":"provides evidence for the exon-foldon correspondence used in foldon assignment"}],"review_version":1}