Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Inferring protein folding mechanisms from natural sequence diversity

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Using only the evolutionary record in a protein family's sequences, the paper infers how individual globular proteins fold — the order of their folding elements, how cooperatively they fold, and how single mutations rewire both.

desk verdict A well-controlled extension of the foldon Ising model to globular proteins, with real validation and honest caveats; the folding-dominance assumption is the main soft spot, but the paper deserves a serious referee. read the letter →

arxiv 2412.14341 v3 pith:OV6K2VQC submitted 2024-12-18 q-bio.BM

classification q-bio.BM
keywords proteinfoldingmechanismsevolutionaryenergyIsingmodelPottsRestrictedBoltzmannMachinefoldonscooperativitytopology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to show that the evolutionary record in protein sequences carries enough information to infer the folding mechanisms of globular proteins — not just their native structures or global stabilities, but which parts fold first, which fold together, and how mutations change that choreography. The authors learn one- and two-body evolutionary energy fields from each family's sequence alignment, map them onto a coarse-grained Ising chain of folding elements called foldons, and simulate thermal unfolding for 500 sequences in each of 15 protein families. They find that native topology sets limits on how much folding cooperativity can vary within a family: beta and alpha/beta folds allow only a few mechanisms despite high sequence diversity, while alpha topologies permit diverse folding scenarios among family members. They further claim that mutation-induced changes in folding temperature and cooperativity can be computed directly from the evolutionary model, and report correlations with experimental denaturation data ($r=0.74$ between predicted cooperativity and experimental m-values). If the claims hold, protein engineers could rank natural variants by folding stability, cooperativity, or both using sequence information alone.

What carries the argument

The central object is the foldon Ising model: a finite chain of $N$ two-state elements (foldons, each folded or unfolded) whose Hamiltonian combines an internal folding free energy per element, a pairwise interaction energy between folded elements, and an entropic cost for unfolded elements. What carries the argument is the mapping of sequence information onto those energies: a Restricted Boltzmann Machine is learned on each family's multiple sequence alignment and converted into Potts-model couplings and local fields, which are summed over the residues of each foldon to yield the Ising energy terms, scaled by a family-specific selection temperature $T_{sel}$ — the apparent temperature at which nature selected the family's sequences, calibrated against experimental stability data and extrapolated via an empirical scaling rule. The foldon partition itself is supplied by Minimal Common Exons, conserved exon-intron boundaries that divide each alignment into common folding elements. From Monte Carlo simulations, the model extracts thermal unfolding curves, free-energy profiles over the number of folded elements $Q$, per-element folding temperatures $T_f$, and a cooperativity score $\rho = Q_{barrier}/(N-1)$ counting the intermediate $Q$ values that are never free-energy minima. Topology-only 'vanilla' models serve as controls that isolate the contribution of the amino-acid-level evolutionary fields.

What would settle it

Measure experimental folding-temperature shifts ($\Delta T_f$) for a set of single mutants located in and around a strongly conserved functional site — for example the active site of a DHFR or RNase H family member — and compare them with the model's per-site predictions: if the correlation between predicted and measured $\Delta T_f$ is systematically worse for functional-site mutations than for surface mutations, or collapses when the conserved functional positions are masked out of the alignment, the central assumption that folding dominates the sequence record is broken.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that the evolutionary energy fields inferred from a multiple sequence alignment can be quantitatively mapped onto a coarse-grained folding Hamiltonian. Each protein is divided into contiguous folding elements — foldons, defined by conserved exon boundaries called Minimal Common Exons — and the sequence-derived energies fix the internal stability of each foldon and the interactions between foldons. Simulating this finite Ising chain, the authors report that the folding temperature $T_f$ varies within a family in proportion to the family's selection temperature $T_{sel}$, and that the variance of the cooperativity score $\rho$ across family members is governed by native topology, quantified by the ratio of short-range to long-range contacts in the reference structure: compact $\beta$ and $\alpha$/$\beta$ proteins can realize only a few mechanisms, whereas elongated $\alpha$ proteins span the full range from all-or-none to downhill folding. For the benchmark enzyme EcDHFR, the model's per-element folding temperatures correlate at $r=0.88$ with a structure-based simulation and match the flexible regions seen in an atomistic molecular dynamics study. The same model predicts the effect of every single-point mutation on both $T_f$ and $\rho$; predicted folding-temperature changes track experimental values for the families with available data, and the sequence-based cooperativity score correlates with experimental m-values ($r = 0.74$, $p = 3.62\times 10^{-80}$).

Load-bearing premise

The load-bearing premise is that folding stability is the main evolutionary pressure recorded in a protein family's sequences, so the statistical energies learned from the alignment genuinely describe folding; if binding, catalysis, allostery, or other functional demands shape any part of the sequence record more strongly, the inferred folding energies, temperatures, and cooperativity scores would be distorted.

Editorial extensions

If this is right

  • For any protein family with a deep multiple sequence alignment, the folding temperature and cooperativity of every natural sequence can be annotated without running folding simulations, because both observables are fitted as functions of the evolutionary energy.
  • Topology sets a ceiling on mechanism diversity: families with compact beta or alpha/beta folds can realize only a few folding mechanisms even when sequence diversity is high, so the reference sequence of a family is not necessarily representative of its members' mechanisms.
  • Single-mutation effects on both stability ($\Delta T_f$) and mechanism ($\Delta \rho$) become computable for every possible amino acid substitution, providing a direct way to rank variants for protein engineering.
  • Cooperativity can be apparent rather than real: in families such as ACBP, Serpin, and Ubiquitin, all-or-none behavior can arise from similar internal stabilities of the folding elements even when inter-element interactions are removed.
  • The relationship between $T_f$ variability and $T_{sel}$ implies that the selection temperature estimated from sequences alone reports how strongly folding stability is constrained during a family's evolution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper does not run: hydrogen/deuterium exchange or NMR measurements on several members of an alpha family such as ACBP should show mechanism diversity matching the predicted spread in $\rho$, while members of a beta family such as Trypsin should not — a direct experimental check of the topology-cooperativity claim.
  • The short-to-long contact ratio rule implies a design principle for protein engineering: folding-pathway control is far easier to engineer in alpha-topology scaffolds than in beta or alpha/beta scaffolds, since the latter relax back toward a narrow set of mechanisms.
  • The weakest assumption could be stress-tested by masking strongly conserved functional positions (active sites, binding interfaces) in the alignments and re-learning the model: if the predicted $T_f$ and $\rho$ shift systematically when those positions are removed, the inferred energies are carrying functional signal rather than pure folding signal.
  • Because the model predicts per-element folding temperatures, it implicitly predicts non-native partially folded states; comparing the predicted order of element folding with experimental kinetic intermediates for a multi-domain protein would test whether the exon-defined foldons are the true cooperative units.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript presents a coarse-grained Ising model of protein folding in which proteins are partitioned into foldons (defined by Minimal Common Exons) and the internal and interfacial free energies are obtained by mapping the one- and two-body fields of a family-specific Restricted Boltzmann Machine to folding energies via Eq. (2), scaled by a family selection temperature T_sel. The authors simulate thermal unfolding for 15 PFAM families (500 sequences each), characterize each family's folding temperature and cooperativity score, and report that within-family cooperativity variability is limited for beta and alpha/beta topologies but larger for alpha proteins. They also compute the effect of single-point mutations on folding temperature and cooperativity, comparing against experimental m-values and Delta-T_f data from ProTherm. The central claim is that folding mechanisms can be inferred from sequence information alone.

Significance. The framework is original and the paper includes several useful controls: reproduction of EcDHFR foldon stabilities against a structure-based model (r=0.88), robustness to alternative foldon partitions (Fig. S4), a family-level correlation of cooperativity variance with short/long-range contact ratio, and a set of clear falsifiable predictions for point mutants. The code and data are deposited on GitHub, which supports reproducibility. However, the validity of the entire pipeline rests on the assumption that the RBM fields are dominated by folding stability rather than by functional constraints; this assumption is acknowledged but not tested. If that assumption holds, the approach could be a valuable way to connect evolutionary sequence records to folding mechanisms and to rank mutants by stability and cooperativity.

major comments (4)
  1. [Introduction; Eq. (2); Fig. S12; Table S2; Concluding remarks] The mapping from sequence to folding energetics in Eq. (2) assumes that the RBM fields are dominated by folding stability, an assumption stated in the Introduction ('we will make the approximation...') and acknowledged in the Concluding remarks as potentially affecting 'local stability and some cooperativity predictions.' This assumption is load-bearing because all downstream quantities--foldon internal energies, surface couplings, T_f, and cooperativity scores--are computed from these fields. The paper's own validations contain red flags: the m-value correlation in Fig. S12 excludes CytochromeC because of heme binding, and Table S2 reports r=0 (Kanaya 1996) and r=0.14 (Lim 1992) for RNase H and Trp syntA, which the authors attribute to active-site mutations. No control is provided that separates functional constraints from folding constraints. I request a stratified analysis: e.g., compare inferred energies against experimental Delta-Delta-G separately for functional-site and non-functional-site mutations, or recompute the model with active-site/gap columns masked, and report how the EcDHFR and m-value validations change. Without such a control, the abstract's claim that folding mechanisms are inferred from 'only sequence information' is premature.
  2. [Eq. (3); Fig. 3B] Because Eq. (2) scales all Ising energies by T_sel, the folding temperature T_f of any sequence scales linearly with T_sel for fixed dimensionless energy patterns. The near-linear relationship between the standard deviation of T_f and T_sel reported in Fig. 3B therefore holds largely by construction, and it does not independently support the evolutionary interpretation that families with low T_sel 'only permit' sequences with T_f close to the family average. The authors should report the distribution of the dimensionless ratio T_f/T_sel across families, or equivalently residual variation after removing the multiplicative T_sel factor, and should test the sensitivity of the results to the Miyazawa assumption of constant sigma(Delta-Delta-G) underlying Eq. (3). As written, the claim in the text is at risk of being a scaling artifact rather than a biological finding.
  3. [Fig. 5A; Fig. S11; Fig. S12] The predicted changes in cooperativity upon mutation are obtained from a linear fit of the cooperativity score in the heterogeneity-interaction plane (Fig. 5A) that is itself fitted to the 7500 simulated sequences; predictions from this fit are then compared to the model outputs again in Fig. S11. This is not an independent test of the model's ability to predict mutational Delta-rho. The only experimental observable linked to rho is the m-value correlation in Fig. S12, which pools many proteins and excludes CytochromeC because of the heme cofactor. I ask for an out-of-sample evaluation: cross-validate the linear surrogate, report per-family and per-mutant m-value correlations, and show the m-value comparison with CytochromeC included and without exclusion. The mutation-prediction section should clearly distinguish what is a computational shortcut from what is experimentally validated.
  4. [Fig. 4C] The central claim that topology limits cooperativity variability within a family rests on the correlation in Fig. 4C between cooperativity variance and N_short/N_long. This is a family-level scatter with only 15 points, and the manuscript does not report a correlation coefficient, p-value, or confidence interval for this relationship. The Copper-bind family is acknowledged as escaping the trend, but no explanation is offered. Please report the statistics (e.g., Spearman r, p), test the robustness of the relationship to the contact definition and to the choice of reference PDB, and discuss whether the result persists when the two or three least well-behaved families are removed.
minor comments (5)
  1. [Eq. (1)] There is a typo in Eq. (1): 'Kroeneker' should be 'Kronecker'.
  2. [Table S2] The first DHFR entry in Table S2 has 'No ID' in the PMID column; the reference should be completed or the column removed.
  3. [Methods (Data curation)] The Methods sentence 'For minimizing the phylogenetic bias within each MSA, we clustered by full sequence similarity using CD-hit at 90% cutoff and we assigned a weight to each sequence defined as being the number of sequences in the th cluster' is incomplete and the subscript is missing; it should read '...defined as 1/n_i, where n_i is the number of sequences in the i-th cluster.'
  4. [Statement of significance] The 'Statement of significance' is quite generic; it could more specifically state the discovered topology-dependence of cooperativity variability and its implications for protein engineering and for interpreting natural sequence diversity.
  5. [Fig. 4C and supplemental figures] Several correlations are described only by panels without reporting the corresponding coefficients and p-values in the main text (e.g., Fig. 4C, Fig. S8, Fig. S10); adding these statistics would make the strength of the trends easier to assess.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity: the central topology and mutation-effect results are validated against external structure-based simulations and experimental m-value/ΔT_f data; the only by-construction relation (T_f vs evolutionary energy) is explicitly acknowledged and is not load-bearing.

  1. self definitional [Results and Discussion, 'Folding mechanism variability', first paragraph]
    "For each family, T_f is correlated with the total evolutionary energy of the sequences (Fig. S5). This general relationship between folding stability and sequence probability is expected from Equation 2 and it is consistent with experimental results [36]."

    T_f is the output of the Ising simulation whose Hamiltonian (Eq. 2) is built from the same RBM one- and two-body evolutionary fields (h_a and J_ab) multiplied by T_sel. Therefore a correlation between T_f and total evolutionary energy is guaranteed by construction; it restates the model's definition rather than testing it. The paper openly labels the relationship 'expected from Equation 2', so the step is transparent and is not used as the primary evidence for the paper's main claims.

full rationale

The paper's derivation chain is: learn an RBM evolutionary energy field from an MSA, map it through Eq. 2 to coarse-grained foldon Ising energies, simulate the Ising chain, and read off T_f, free-energy profiles, and cooperativity scores. The only step where an output is forced by its input is the T_f-versus-evolutionary-energy correlation, which the paper itself says is 'expected from Equation 2'; this is a self-definitional consistency statement, not a load-bearing validation. The main claims are supported by independent external benchmarks: the EcDHFR foldon stabilities correlate with a structure-based Cα-SBM simulation (r=0.88), the cooperativity score correlates with experimental ProTherm m-values (r=0.74), and predicted ΔT_f values correlate with experimental data in Table S2 for several families, with failures in active-site mutants explicitly attributed to non-folding constraints. The exon-based foldon definition and the per-residue entropy are self-cited from prior work by the same group, but the paper tests robustness to alternative foldon partitions (Fig. S4), and the entropy scale mainly affects absolute temperatures rather than the structural correlations that anchor the conclusions. The folding-dominance assumption ('the energetics of protein folding is the main evolutionary pressure acting globally on protein sequences') is acknowledged as an approximation and could distort results for function-dominated families, but an untested assumption is a correctness risk, not a circular reduction under the stated rules. Overall the derivation has substantial independent content, so the circularity score is low.

Assumptions & free parameters 6 free parameters · 7 assumptions · 0 invented entities

The model is built on parameters and assumptions carried over from prior same-group work (entropy per residue, MCE foldon definitions, the repeat-protein Ising framework) and on the untested assumption that folding stability dominates the evolutionary record. The mutation predictions for cooperativity additionally rely on a fitted linear surface. There are no newly invented physical entities.

free parameters (6)
  • s (entropy per residue) = 5 cal mol^-1 K^-1 res^-1
    Taken from the repeat-protein model (ref 26), assumed additive and independent of amino acid identity; sets the absolute entropic cost of folding each foldon and therefore the temperature scale of all predicted folding mechanisms.
  • T_sel_PDZ = not stated in text (fitted from PDZ Delta-Delta-G vs evolutionary energy, Fig. S14)
    Obtained by linear regression of experimental Delta-Delta-G for PDZ mutants on RBM evolutionary energy differences; serves as the anchor for all family selection temperatures.
  • T_sel per family (14 families) = values 160-275 K in Table S1
    Computed from Eq. 3 as T_sel_PDZ times the ratio of standard deviations of evolutionary energy changes, relying on the assumption that experimental Delta-Delta-G spread is constant across families; each value sets the energy scale for that family's Ising model.
  • Linear fit coefficients for cooperativity prediction = not reported numerically
    The linear fit of cooperativity score on foldon energetic heterogeneity and average interaction strength (Fig. 5A) is used to estimate Delta-rho for single-site mutations without running simulations.
  • Per-family linear fit of T_f vs evolutionary energy = slopes and intercepts not reported
    Used to predict Delta-T_f for single-site mutations from evolutionary energy changes (Fig. 5B-C); fitted to 500 natural sequences per family.
  • RBM hyperparameters = 500 hidden units, 500 iterations, regularization lambda=0.25
    Chosen by grid search on DHFR and applied uniformly to all families; they affect the learned evolutionary fields and thus all downstream energies.
assumptions (7)
  • domain assumption Folding stability is the main evolutionary pressure acting globally on protein sequences
    Introduced in the Introduction and reiterated in the Concluding remarks. If false, the inferred fields contain functional signals that contaminate the folding energy terms.
  • domain assumption Evolutionary energy fields map linearly to folding free energies with a single selection temperature
    Eq. 2 converts RBM/Potts h and J into foldon internal and surface energies via T_sel; no nonlinear corrections or context-dependent terms are considered.
  • domain assumption Minimal Common Exons define the cooperative folding elements (foldons)
    The model partitions each protein into MCEs derived from exon-intron boundaries (ref 25); the control with alternative partitions (Fig. S4) mitigates but does not remove this assumption.
  • domain assumption The standard deviation of experimental Delta-Delta-G is nearly constant across protein families
    Used in Eq. 3 to scale T_sel from PDZ to other families based on Miyazawa (ref 30); if false, all non-PDZ selection temperatures are biased.
  • domain assumption Foldon entropy is additive, sequence-independent, and equal to L_j times s
    Stated in Model Definition; the per-residue entropy from repeat proteins is applied to all globular foldons without sensitivity analysis.
  • standard math The energy landscape relation 1/T_g^2 + 1/T_f^2 = 2/(T_sel T_f) holds for these families
    Taken from Pande et al. (ref 5) and used to compute T_g and the funnelness ratios in Fig. 3C.
  • standard math Monte Carlo Metropolis sampling converges to the equilibrium distribution of the Ising model
    The paper determines simulation and equilibration times by autocorrelation analysis; standard assumption for MC sampling.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Inferring protein folding mechanisms from natural sequence diversity." pith.science (2026). https://pith.science/paper/OV6K2VQC

@misc{pith2026241214341,
  author       = {Pith},
  title        = {Pith review of: Inferring protein folding mechanisms from natural sequence diversity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OV6K2VQC}},
  note         = {Machine review of arXiv:2412.14341}
}
read the original abstract

Protein sequences serve as a natural record of the evolutionary constraints that shape their functional structures. We show that it is possible to use only sequence information to go beyond predicting native structures and global stability to infer the folding mechanisms of globular proteins. The one- and two-body evolutionary energy fields at the amino-acid level are mapped to a coarse-grained description of folding, where proteins are divided into contiguous folding elements, commonly referred to as foldons. For 15 diverse protein families, we calculated the folding mechanisms of hundreds of proteins by simulating an Ising chain of foldons, with their energetics determined by the amino acid sequences. We show that protein topology imposes limits on the variability of folding cooperativity within a family. While most beta and alpha/beta structures exhibit only a few possible mechanisms despite high sequence diversity, alpha topologies allow for diverse folding scenarios among family members. We show that both the stability and cooperativity changes induced by mutations can be computed directly using sequence-based evolutionary models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting protein folding dynamics using sequence information

    q-bio.BM 2025-05 conditional novelty 4.0 of 10

    A pipeline from sequence alignments through a Potts model to an Ising foldon chain predicts protein folding curves, subdomains, and mutation effects, but without new experimental validation in this paper.

Reference graph

Works this paper leans on

43 extracted references · 35 canonical work pages · cited by 1 Pith paper

  1. [1]

    Chemical physics of protein folding,

    P. G. Wolynes, W. A. Eaton, and A. R. Fersht, “Chemical physics of protein folding,” Proc. Natl. Acad. Sci. , vol. 109, no. 44, pp. 17770–17771, Oct. 2012, doi: 10.1073/pnas.1215733109

  2. [2]

    Protein folding funnels: a kinetic approach to the sequence-structure relationship.,

    P. E. Leopold, M. Montal, and J. N. Onuchic, “Protein folding funnels: a kinetic approach to the sequence-structure relationship.,” Proc. Natl. Acad. Sci. , vol. 89, no. 18, pp. 8721–8725, Sep. 1992, doi: 10.1073/pnas.89.18.8721

  3. [3]

    Spin glasses and the statistical mechanics of protein folding,

    J. D. Bryngelson and P. G. Wolynes, “Spin glasses and the statistical mechanics of protein folding,” Proc. Natl. Acad. Sci. , vol. 84, no. 21, pp. 7524–7528, 1987

  4. [4]

    Modeling evolutionary landscapes: Mutational stability, topology, and superfunnels in sequence space,

    E. Bornberg-Bauer and H. S. Chan, “Modeling evolutionary landscapes: Mutational stability, topology, and superfunnels in sequence space,” Proc. Natl. Acad. Sci. , vol. 96, no. 19, pp. 10689–10694, Sep. 1999, doi: 10.1073/pnas.96.19.10689

  5. [5]

    Statistical mechanics of simple models of protein folding and design,

    V. S. Pande, A. Y. Grosberg, and T. Tanaka, “Statistical mechanics of simple models of protein folding and design,” Biophys. J. , vol. 73, no. 6, pp. 3192–3210, 1997, doi: 10.1016/S0006-3495(97)78345-0

  6. [6]

    Frustration in biomolecules,

    D. U. Ferreiro, E. A. Komives, and P. G. Wolynes, “Frustration in biomolecules,” Q. Rev. Biophys. , vol. 47, no. 4, pp. 285–363, Nov. 2014, doi: 10.1017/S0033583514000092

  7. [7]

    Molecular Information Theory Meets Protein Folding,

    I. E. Sánchez, E. A. Galpern, M. M. Garibaldi, and D. U. Ferreiro, “Molecular Information Theory Meets Protein Folding,” J. Phys. Chem. B , vol. 126, no. 43, pp. 8655–8668, Nov. 2022, doi: 10.1021/acs.jpcb.2c04532

  8. [9]

    Machine learning in protein structure prediction,

    M. AlQuraishi, “Machine learning in protein structure prediction,” Curr. Opin. Chem. Biol. , vol. 65, pp. 1–8, Dec. 2021, doi: 10.1016/j.cbpa.2021.04.005

Show all 43 references
  1. [10]

    Contact order, transition state placement and the refolding rates of single domain proteins 1 1Edited by P. E. Wright,

    K. W. Plaxco, K. T. Simons, and D. Baker, “Contact order, transition state placement and the refolding rates of single domain proteins 1 1Edited by P. E. Wright,” J. Mol. Biol. , vol. 277, no. 4, pp. 985–994, Apr. 1998, doi: 10.1006/jmbi.1998.1645

  2. [11]

    Coarse-grained models of protein folding: toy models or predictive tools?,

    C. Clementi, “Coarse-grained models of protein folding: toy models or predictive tools?,” Curr. Opin. Struct. Biol. , vol. 18, no. 1, pp. 10–15, Feb. 2008, doi: 10.1016/j.sbi.2007.10.005

  3. [12]

    Frustration, function and folding,

    D. U. Ferreiro, E. A. Komives, and P. G. Wolynes, “Frustration, function and folding,” Curr. Opin. Struct. Biol. , vol. 48, pp. 68–73, Feb. 2018, doi: 10.1016/j.sbi.2017.09.006

  4. [13]

    Conserved residues and the mechanism of protein folding,

    E. Shakhnovich, V. Abkevich, and O. Ptitsyn, “Conserved residues and the mechanism of protein folding,” Nature , vol. 379, no. 6560, pp. 96–98, Jan. 1996, doi: 10.1038/379096a0

  5. [14]

    Identification of direct residue contacts in protein-protein interaction by message passing,

    M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, and T. Hwa, “Identification of direct residue contacts in protein-protein interaction by message passing,” Proc. Natl. Acad. Sci. U. S. A. , vol. 106, no. 1, pp. 67–72, 2009, doi: 10.1073/pnas.0805923106

  6. [15]

    Direct-coupling analysis of residue coevolution captures native contacts across many protein families,

    F. Morcos et al. , “Direct-coupling analysis of residue coevolution captures native contacts across many protein families,” Proc. Natl. Acad. Sci. U. S. A. , vol. 108, no. 49, 2011, doi: 10.1073/pnas.1111471108

  7. [16]

    Inverse statistical physics of protein sequences: A key issues review,

    S. Cocco, C. Feinauer, M. Figliuzzi, R. Monasson, and M. Weigt, “Inverse statistical physics of protein sequences: A key issues review,” Rep. Prog. Phys. , vol. 81, no. 3, 2018, doi: 10.1088/1361-6633/aa9965

  8. [17]

    Direct Coupling Analysis for Protein Contact Prediction,

    F. Morcos, T. Hwa, J. N. Onuchic, and M. Weigt, “Direct Coupling Analysis for Protein Contact Prediction,” in Protein Structure Prediction , vol. 1137, D. Kihara, Ed., in Methods in Molecular Biology, vol. 1137. , New York, NY: Springer New York, 2014, pp. 55–70. doi: 10.1007/...

  9. [18]

    Coevolutionary Landscape Inference and the Context-Dependence of Mutations in Beta-Lactamase TEM-1,

    M. Figliuzzi, H. Jacquier, A. Schug, O. Tenaillon, and M. Weigt, “Coevolutionary Landscape Inference and the Context-Dependence of Mutations in Beta-Lactamase TEM-1,” Mol. Biol. Evol. , vol. 33, no. 1, pp. 268–280, Jan. 2016, doi: 10.1093/molbev/msv211

  10. [19]

    Inferring repeat-protein energetics from evolutionary information,

    R. Espada, R. G. Parra, T. Mora, A. M. Walczak, and D. U. Ferreiro, “Inferring repeat-protein energetics from evolutionary information,” PLoS Comput. Biol. , vol. 13, no. 6, pp. 1–16, 2017, doi: 10.1371/journal.pcbi.1005584

  11. [20]

    Size and structure of the sequence space of repeat proteins,

    J. Marchi, E. A. Galpern, R. Espada, D. U. Ferreiro, A. M. Walczak, and T. Mora, “Size and structure of the sequence space of repeat proteins,” PLoS Comput. Biol. , vol. 15, no. 8, pp. 1–23, 2019, doi: 10.1371/journal.pcbi.1007282

  12. [21]

    Epistatic contributions promote the unification of incompatible models of neutral molecular evolution,

    J. A. De La Paz, C. M. Nartey, M. Yuvaraj, and F. Morcos, “Epistatic contributions promote the unification of incompatible models of neutral molecular evolution,” Proc. Natl. Acad. Sci. , vol. 117, no. 11, pp. 5873–5882, Mar. 2020, doi: 10.1073/pnas.1913071117

  13. [22]

    Emergent time scales of epistasis in protein evolution,

    L. Di Bari, M. Bisardi, S. Cotogno, M. Weigt, and F. Zamponi, “Emergent time scales of epistasis in protein evolution,” Proc. Natl. Acad. Sci. , vol. 121, no. 40, p. e2406807121, Oct. 2024, doi: 10.1073/pnas.2406807121

  14. [23]

    Kinetic coevolutionary models predict the temporal emergence of HIV-1 resistance mutations under drug selection pressure,

    A. Biswas, I. Choudhuri, E. Arnold, D. Lyumkis, A. Haldane, and R. M. Levy, “Kinetic coevolutionary models predict the temporal emergence of HIV-1 resistance mutations under drug selection pressure,” Proc. Natl. Acad. Sci. , vol. 121, no. 15, p. e2316662121, Apr. 2024, doi: 10...

  15. [24]

    Foldons, protein structural modules, and exons.,

    A. R. Panchenko, Z. Luthey-Schulten, and P. G. Wolynes, “Foldons, protein structural modules, and exons.,” Proc. Natl. Acad. Sci. , vol. 93, no. 5, pp. 2008–2013, Mar. 1996, doi: 10.1073/pnas.93.5.2008

  16. [25]

    Reassessing the exon–foldon correspondence using frustration analysis,

    E. A. Galpern, H. Jaafari, C. Bueno, P. G. Wolynes, and D. U. Ferreiro, “Reassessing the exon–foldon correspondence using frustration analysis,” Proc. Natl. Acad. Sci. , vol. 121, no. 28, p. e2400151121, Jul. 2024, doi: 10.1073/pnas.2400151121

  17. [26]

    Evolution and folding of repeat proteins,

    E. A. Galpern, J. Marchi, T. Mora, A. M. Walczak, and D. U. Ferreiro, “Evolution and folding of repeat proteins,” Proc. Natl. Acad. Sci. , vol. 119, no. 31, p. e2204131119, 2022, doi: 10.1073/pnas.2204131119

  18. [27]

    The energy landscapes of repeat-containing proteins: Topology, cooperativity, and the folding funnels of one-dimensional architectures,

    D. U. Ferreiro, A. M. Walczak, E. A. Komives, and P. G. Wolynes, “The energy landscapes of repeat-containing proteins: Topology, cooperativity, and the folding funnels of one-dimensional architectures,” PLoS Comput. Biol. , vol. 4, no. 5, 2008, doi: 10.1371/journal.pcbi.1000070

  19. [28]

    Learning protein constitutive motifs from sequence data,

    J. Tubiana, S. Cocco, and R. Monasson, “Learning protein constitutive motifs from sequence data,” eLife , vol. 8, 2019, doi: 10.7554/eLife.39397

  20. [29]

    Coevolutionary information, protein folding landscapes, and the thermodynamics of natural selection,

    F. Morcos, N. P. Schafer, R. R. Cheng, J. N. Onuchic, and P. G. Wolynes, “Coevolutionary information, protein folding landscapes, and the thermodynamics of natural selection,” Proc. Natl. Acad. Sci. , vol. 111, no. 34, pp. 12408–12413, 2014, doi: 10.1073/pnas.1413575111

  21. [30]

    Selection originating from protein stability/foldability: Relationships between protein folding free energy, sequence ensemble, and fitness,

    S. Miyazawa, “Selection originating from protein stability/foldability: Relationships between protein folding free energy, sequence ensemble, and fitness,” J. Theor. Biol. , vol. 433, pp. 21–38, 2017, doi: 10.1016/j.jtbi.2017.08.018

  22. [31]

    Refolding of Escherichia coli dihydrofolate reductase: sequential formation of substrate binding sites.,

    C. Frieden, “Refolding of Escherichia coli dihydrofolate reductase: sequential formation of substrate binding sites.,” Proc. Natl. Acad. Sci. , vol. 87, no. 12, pp. 4413–4416, Jun. 1990, doi: 10.1073/pnas.87.12.4413

  23. [32]

    Folding of dihydrofolate reductase from Escherichia coli,

    N. A. Touchette, K. M. Perry, and C. R. Matthews, “Folding of dihydrofolate reductase from Escherichia coli,” Biochemistry , vol. 25, no. 19, pp. 5445–5452, Sep. 1986, doi: 10.1021/bi00367a015

  24. [33]

    Thermal unfolding molecular dynamics simulation of Escherichia coli dihydrofolate reductase: Thermal stability of protein domains and unfolding pathway,

    Y. Y. Sham, B. Ma, C. Tsai, and R. Nussinov, “Thermal unfolding molecular dynamics simulation of Escherichia coli dihydrofolate reductase: Thermal stability of protein domains and unfolding pathway,” Proteins Struct. Funct. Bioinforma. , vol. 46, no. 3, pp. 308–320, Feb. 2002,...

  25. [34]

    Microsecond Subdomain Folding in Dihydrofolate Reductase,

    M. Arai, M. Iwakura, C. R. Matthews, and O. Bilsel, “Microsecond Subdomain Folding in Dihydrofolate Reductase,” J. Mol. Biol. , vol. 410, no. 2, pp. 329–342, Jul. 2011, doi: 10.1016/j.jmb.2011.04.057

  26. [35]

    Structure of a partially unfolded form of E scherichia coli dihydrofolate reductase provides insight into its folding pathway,

    J. R. Kasper, P. Liu, and C. Park, “Structure of a partially unfolded form of E scherichia coli dihydrofolate reductase provides insight into its folding pathway,” Protein Sci. , vol. 23, no. 12, pp. 1728–1737, Dec. 2014, doi: 10.1002/pro.2555

  27. [36]

    Co-Evolutionary Fitness Landscapes for Sequence Design,

    P. Tian, J. M. Louis, J. L. Baber, A. Aniana, and R. B. Best, “Co-Evolutionary Fitness Landscapes for Sequence Design,” Angew. Chem. - Int. Ed. , vol. 57, no. 20, pp. 5674–5678, 2018, doi: 10.1002/anie.201713220

  28. [37]

    Solvent constraints for biopolymer folding and evolution in extraterrestrial environments,

    I. E. Sánchez, E. A. Galpern, and D. U. Ferreiro, “Solvent constraints for biopolymer folding and evolution in extraterrestrial environments,” Proc. Natl. Acad. Sci. , vol. 121, no. 21, p. e2318905121, May 2024, doi: 10.1073/pnas.2318905121

  29. [38]

    Quantitative criteria for native energetic heterogeneity influences in the prediction of protein folding kinetics,

    S. S. Cho, Y. Levy, and P. G. Wolynes, “Quantitative criteria for native energetic heterogeneity influences in the prediction of protein folding kinetics,” Proc. Natl. Acad. Sci. , vol. 106, no. 2, pp. 434–439, Jan. 2009, doi: 10.1073/pnas.0810218105

  30. [39]

    Start2Fold: A database of hydrogen/deuterium exchange data on protein folding and stability,

    R. Pancsa, M. Varadi, P. Tompa, and W. F. Vranken, “Start2Fold: A database of hydrogen/deuterium exchange data on protein folding and stability,” Nucleic Acids Res. , vol. 44, no. D1, pp. D429–D434, 2016, doi: 10.1093/nar/gkv1185

  31. [40]

    Local energetic frustration conservation in protein families and superfamilies,

    M. I. Freiberger et al. , “Local energetic frustration conservation in protein families and superfamilies,” Nat. Commun. , vol. 14, no. 1, p. 8379, Dec. 2023, doi: 10.1038/s41467-023-43801-2

  32. [41]

    The Pfam protein families database: towards a more sustainable future,

    R. D. Finn et al. , “The Pfam protein families database: towards a more sustainable future,” Nucleic Acids Res. , vol. 44, no. D1, pp. D279–D285, Jan. 2016, doi: 10.1093/nar/gkv1344

  33. [42]

    InterPro in 2022,

    T. Paysan-Lafosse et al. , “InterPro in 2022,” Nucleic Acids Res. , vol. 51, no. D1, pp. D418–D427, Jan. 2023, doi: 10.1093/nar/gkac993

  34. [43]

    Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences,

    W. Li and A. Godzik, “Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences,” Bioinformatics , vol. 22, no. 13, pp. 1658–1659, Jul. 2006, doi: 10.1093/bioinformatics/btl158

  35. [44]

    A PDZ domain recapitulates a unifying mechanism for protein folding,

    S. Gianni et al. , “A PDZ domain recapitulates a unifying mechanism for protein folding,” Proc. Natl. Acad. Sci. , vol. 104, no. 1, pp. 128–133, Jan. 2007, doi: 10.1073/pnas.0602770104. Supplemental Information for Inferring protein folding mechanisms from natural sequence div...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.