REVIEW 4 major objections 5 minor 1 cited by
Inferring protein folding mechanisms from natural sequence diversity
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Using only the evolutionary record in a protein family's sequences, the paper infers how individual globular proteins fold — the order of their folding elements, how cooperatively they fold, and how single mutations rewire both.
desk verdict A well-controlled extension of the foldon Ising model to globular proteins, with real validation and honest caveats; the folding-dominance assumption is the main soft spot, but the paper deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the foldon Ising model: a finite chain of $N$ two-state elements (foldons, each folded or unfolded) whose Hamiltonian combines an internal folding free energy per element, a pairwise interaction energy between folded elements, and an entropic cost for unfolded elements. What carries the argument is the mapping of sequence information onto those energies: a Restricted Boltzmann Machine is learned on each family's multiple sequence alignment and converted into Potts-model couplings and local fields, which are summed over the residues of each foldon to yield the Ising energy terms, scaled by a family-specific selection temperature $T_{sel}$ — the apparent temperature at which nature selected the family's sequences, calibrated against experimental stability data and extrapolated via an empirical scaling rule. The foldon partition itself is supplied by Minimal Common Exons, conserved exon-intron boundaries that divide each alignment into common folding elements. From Monte Carlo simulations, the model extracts thermal unfolding curves, free-energy profiles over the number of folded elements $Q$, per-element folding temperatures $T_f$, and a cooperativity score $\rho = Q_{barrier}/(N-1)$ counting the intermediate $Q$ values that are never free-energy minima. Topology-only 'vanilla' models serve as controls that isolate the contribution of the amino-acid-level evolutionary fields.
What would settle it
Measure experimental folding-temperature shifts ($\Delta T_f$) for a set of single mutants located in and around a strongly conserved functional site — for example the active site of a DHFR or RNase H family member — and compare them with the model's per-site predictions: if the correlation between predicted and measured $\Delta T_f$ is systematically worse for functional-site mutations than for surface mutations, or collapses when the conserved functional positions are masked out of the alignment, the central assumption that folding dominates the sequence record is broken.
Extended reading notes
Core claim
On its own terms, the paper establishes that the evolutionary energy fields inferred from a multiple sequence alignment can be quantitatively mapped onto a coarse-grained folding Hamiltonian. Each protein is divided into contiguous folding elements — foldons, defined by conserved exon boundaries called Minimal Common Exons — and the sequence-derived energies fix the internal stability of each foldon and the interactions between foldons. Simulating this finite Ising chain, the authors report that the folding temperature $T_f$ varies within a family in proportion to the family's selection temperature $T_{sel}$, and that the variance of the cooperativity score $\rho$ across family members is governed by native topology, quantified by the ratio of short-range to long-range contacts in the reference structure: compact $\beta$ and $\alpha$/$\beta$ proteins can realize only a few mechanisms, whereas elongated $\alpha$ proteins span the full range from all-or-none to downhill folding. For the benchmark enzyme EcDHFR, the model's per-element folding temperatures correlate at $r=0.88$ with a structure-based simulation and match the flexible regions seen in an atomistic molecular dynamics study. The same model predicts the effect of every single-point mutation on both $T_f$ and $\rho$; predicted folding-temperature changes track experimental values for the families with available data, and the sequence-based cooperativity score correlates with experimental m-values ($r = 0.74$, $p = 3.62\times 10^{-80}$).
Load-bearing premise
The load-bearing premise is that folding stability is the main evolutionary pressure recorded in a protein family's sequences, so the statistical energies learned from the alignment genuinely describe folding; if binding, catalysis, allostery, or other functional demands shape any part of the sequence record more strongly, the inferred folding energies, temperatures, and cooperativity scores would be distorted.
Editorial extensions
If this is right
- For any protein family with a deep multiple sequence alignment, the folding temperature and cooperativity of every natural sequence can be annotated without running folding simulations, because both observables are fitted as functions of the evolutionary energy.
- Topology sets a ceiling on mechanism diversity: families with compact beta or alpha/beta folds can realize only a few folding mechanisms even when sequence diversity is high, so the reference sequence of a family is not necessarily representative of its members' mechanisms.
- Single-mutation effects on both stability ($\Delta T_f$) and mechanism ($\Delta \rho$) become computable for every possible amino acid substitution, providing a direct way to rank variants for protein engineering.
- Cooperativity can be apparent rather than real: in families such as ACBP, Serpin, and Ubiquitin, all-or-none behavior can arise from similar internal stabilities of the folding elements even when inter-element interactions are removed.
- The relationship between $T_f$ variability and $T_{sel}$ implies that the selection temperature estimated from sequences alone reports how strongly folding stability is constrained during a family's evolution.
Reading between the lines
- A testable extension the paper does not run: hydrogen/deuterium exchange or NMR measurements on several members of an alpha family such as ACBP should show mechanism diversity matching the predicted spread in $\rho$, while members of a beta family such as Trypsin should not — a direct experimental check of the topology-cooperativity claim.
- The short-to-long contact ratio rule implies a design principle for protein engineering: folding-pathway control is far easier to engineer in alpha-topology scaffolds than in beta or alpha/beta scaffolds, since the latter relax back toward a narrow set of mechanisms.
- The weakest assumption could be stress-tested by masking strongly conserved functional positions (active sites, binding interfaces) in the alignments and re-learning the model: if the predicted $T_f$ and $\rho$ shift systematically when those positions are removed, the inferred energies are carrying functional signal rather than pure folding signal.
- Because the model predicts per-element folding temperatures, it implicitly predicts non-native partially folded states; comparing the predicted order of element folding with experimental kinetic intermediates for a multi-domain protein would test whether the exon-defined foldons are the true cooperative units.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents a coarse-grained Ising model of protein folding in which proteins are partitioned into foldons (defined by Minimal Common Exons) and the internal and interfacial free energies are obtained by mapping the one- and two-body fields of a family-specific Restricted Boltzmann Machine to folding energies via Eq. (2), scaled by a family selection temperature T_sel. The authors simulate thermal unfolding for 15 PFAM families (500 sequences each), characterize each family's folding temperature and cooperativity score, and report that within-family cooperativity variability is limited for beta and alpha/beta topologies but larger for alpha proteins. They also compute the effect of single-point mutations on folding temperature and cooperativity, comparing against experimental m-values and Delta-T_f data from ProTherm. The central claim is that folding mechanisms can be inferred from sequence information alone.
Significance. The framework is original and the paper includes several useful controls: reproduction of EcDHFR foldon stabilities against a structure-based model (r=0.88), robustness to alternative foldon partitions (Fig. S4), a family-level correlation of cooperativity variance with short/long-range contact ratio, and a set of clear falsifiable predictions for point mutants. The code and data are deposited on GitHub, which supports reproducibility. However, the validity of the entire pipeline rests on the assumption that the RBM fields are dominated by folding stability rather than by functional constraints; this assumption is acknowledged but not tested. If that assumption holds, the approach could be a valuable way to connect evolutionary sequence records to folding mechanisms and to rank mutants by stability and cooperativity.
major comments (4)
- [Introduction; Eq. (2); Fig. S12; Table S2; Concluding remarks] The mapping from sequence to folding energetics in Eq. (2) assumes that the RBM fields are dominated by folding stability, an assumption stated in the Introduction ('we will make the approximation...') and acknowledged in the Concluding remarks as potentially affecting 'local stability and some cooperativity predictions.' This assumption is load-bearing because all downstream quantities--foldon internal energies, surface couplings, T_f, and cooperativity scores--are computed from these fields. The paper's own validations contain red flags: the m-value correlation in Fig. S12 excludes CytochromeC because of heme binding, and Table S2 reports r=0 (Kanaya 1996) and r=0.14 (Lim 1992) for RNase H and Trp syntA, which the authors attribute to active-site mutations. No control is provided that separates functional constraints from folding constraints. I request a stratified analysis: e.g., compare inferred energies against experimental Delta-Delta-G separately for functional-site and non-functional-site mutations, or recompute the model with active-site/gap columns masked, and report how the EcDHFR and m-value validations change. Without such a control, the abstract's claim that folding mechanisms are inferred from 'only sequence information' is premature.
- [Eq. (3); Fig. 3B] Because Eq. (2) scales all Ising energies by T_sel, the folding temperature T_f of any sequence scales linearly with T_sel for fixed dimensionless energy patterns. The near-linear relationship between the standard deviation of T_f and T_sel reported in Fig. 3B therefore holds largely by construction, and it does not independently support the evolutionary interpretation that families with low T_sel 'only permit' sequences with T_f close to the family average. The authors should report the distribution of the dimensionless ratio T_f/T_sel across families, or equivalently residual variation after removing the multiplicative T_sel factor, and should test the sensitivity of the results to the Miyazawa assumption of constant sigma(Delta-Delta-G) underlying Eq. (3). As written, the claim in the text is at risk of being a scaling artifact rather than a biological finding.
- [Fig. 5A; Fig. S11; Fig. S12] The predicted changes in cooperativity upon mutation are obtained from a linear fit of the cooperativity score in the heterogeneity-interaction plane (Fig. 5A) that is itself fitted to the 7500 simulated sequences; predictions from this fit are then compared to the model outputs again in Fig. S11. This is not an independent test of the model's ability to predict mutational Delta-rho. The only experimental observable linked to rho is the m-value correlation in Fig. S12, which pools many proteins and excludes CytochromeC because of the heme cofactor. I ask for an out-of-sample evaluation: cross-validate the linear surrogate, report per-family and per-mutant m-value correlations, and show the m-value comparison with CytochromeC included and without exclusion. The mutation-prediction section should clearly distinguish what is a computational shortcut from what is experimentally validated.
- [Fig. 4C] The central claim that topology limits cooperativity variability within a family rests on the correlation in Fig. 4C between cooperativity variance and N_short/N_long. This is a family-level scatter with only 15 points, and the manuscript does not report a correlation coefficient, p-value, or confidence interval for this relationship. The Copper-bind family is acknowledged as escaping the trend, but no explanation is offered. Please report the statistics (e.g., Spearman r, p), test the robustness of the relationship to the contact definition and to the choice of reference PDB, and discuss whether the result persists when the two or three least well-behaved families are removed.
minor comments (5)
- [Eq. (1)] There is a typo in Eq. (1): 'Kroeneker' should be 'Kronecker'.
- [Table S2] The first DHFR entry in Table S2 has 'No ID' in the PMID column; the reference should be completed or the column removed.
- [Methods (Data curation)] The Methods sentence 'For minimizing the phylogenetic bias within each MSA, we clustered by full sequence similarity using CD-hit at 90% cutoff and we assigned a weight to each sequence defined as being the number of sequences in the th cluster' is incomplete and the subscript is missing; it should read '...defined as 1/n_i, where n_i is the number of sequences in the i-th cluster.'
- [Statement of significance] The 'Statement of significance' is quite generic; it could more specifically state the discovered topology-dependence of cooperativity variability and its implications for protein engineering and for interpreting natural sequence diversity.
- [Fig. 4C and supplemental figures] Several correlations are described only by panels without reporting the corresponding coefficients and p-values in the main text (e.g., Fig. 4C, Fig. S8, Fig. S10); adding these statistics would make the strength of the trends easier to assess.
Circularity Check
No significant circularity: the central topology and mutation-effect results are validated against external structure-based simulations and experimental m-value/ΔT_f data; the only by-construction relation (T_f vs evolutionary energy) is explicitly acknowledged and is not load-bearing.
-
self definitional
[Results and Discussion, 'Folding mechanism variability', first paragraph]
"For each family, T_f is correlated with the total evolutionary energy of the sequences (Fig. S5). This general relationship between folding stability and sequence probability is expected from Equation 2 and it is consistent with experimental results [36]."
T_f is the output of the Ising simulation whose Hamiltonian (Eq. 2) is built from the same RBM one- and two-body evolutionary fields (h_a and J_ab) multiplied by T_sel. Therefore a correlation between T_f and total evolutionary energy is guaranteed by construction; it restates the model's definition rather than testing it. The paper openly labels the relationship 'expected from Equation 2', so the step is transparent and is not used as the primary evidence for the paper's main claims.
full rationale
The paper's derivation chain is: learn an RBM evolutionary energy field from an MSA, map it through Eq. 2 to coarse-grained foldon Ising energies, simulate the Ising chain, and read off T_f, free-energy profiles, and cooperativity scores. The only step where an output is forced by its input is the T_f-versus-evolutionary-energy correlation, which the paper itself says is 'expected from Equation 2'; this is a self-definitional consistency statement, not a load-bearing validation. The main claims are supported by independent external benchmarks: the EcDHFR foldon stabilities correlate with a structure-based Cα-SBM simulation (r=0.88), the cooperativity score correlates with experimental ProTherm m-values (r=0.74), and predicted ΔT_f values correlate with experimental data in Table S2 for several families, with failures in active-site mutants explicitly attributed to non-folding constraints. The exon-based foldon definition and the per-residue entropy are self-cited from prior work by the same group, but the paper tests robustness to alternative foldon partitions (Fig. S4), and the entropy scale mainly affects absolute temperatures rather than the structural correlations that anchor the conclusions. The folding-dominance assumption ('the energetics of protein folding is the main evolutionary pressure acting globally on protein sequences') is acknowledged as an approximation and could distort results for function-dominated families, but an untested assumption is a correctness risk, not a circular reduction under the stated rules. Overall the derivation has substantial independent content, so the circularity score is low.
Assumptions & free parameters
free parameters (6)
- s (entropy per residue) =
5 cal mol^-1 K^-1 res^-1
- T_sel_PDZ =
not stated in text (fitted from PDZ Delta-Delta-G vs evolutionary energy, Fig. S14)
- T_sel per family (14 families) =
values 160-275 K in Table S1
- Linear fit coefficients for cooperativity prediction =
not reported numerically
- Per-family linear fit of T_f vs evolutionary energy =
slopes and intercepts not reported
- RBM hyperparameters =
500 hidden units, 500 iterations, regularization lambda=0.25
assumptions (7)
- domain assumption Folding stability is the main evolutionary pressure acting globally on protein sequences
- domain assumption Evolutionary energy fields map linearly to folding free energies with a single selection temperature
- domain assumption Minimal Common Exons define the cooperative folding elements (foldons)
- domain assumption The standard deviation of experimental Delta-Delta-G is nearly constant across protein families
- domain assumption Foldon entropy is additive, sequence-independent, and equal to L_j times s
- standard math The energy landscape relation 1/T_g^2 + 1/T_f^2 = 2/(T_sel T_f) holds for these families
- standard math Monte Carlo Metropolis sampling converges to the equilibrium distribution of the Ising model
Cite this review
Pith. "Pith review of Inferring protein folding mechanisms from natural sequence diversity." pith.science (2026). https://pith.science/paper/OV6K2VQC
@misc{pith2026241214341,
author = {Pith},
title = {Pith review of: Inferring protein folding mechanisms from natural sequence diversity},
year = {2026},
howpublished = {\url{https://pith.science/paper/OV6K2VQC}},
note = {Machine review of arXiv:2412.14341}
}
read the original abstract
Protein sequences serve as a natural record of the evolutionary constraints that shape their functional structures. We show that it is possible to use only sequence information to go beyond predicting native structures and global stability to infer the folding mechanisms of globular proteins. The one- and two-body evolutionary energy fields at the amino-acid level are mapped to a coarse-grained description of folding, where proteins are divided into contiguous folding elements, commonly referred to as foldons. For 15 diverse protein families, we calculated the folding mechanisms of hundreds of proteins by simulating an Ising chain of foldons, with their energetics determined by the amino acid sequences. We show that protein topology imposes limits on the variability of folding cooperativity within a family. While most beta and alpha/beta structures exhibit only a few possible mechanisms despite high sequence diversity, alpha topologies allow for diverse folding scenarios among family members. We show that both the stability and cooperativity changes induced by mutations can be computed directly using sequence-based evolutionary models.
Forward citations
Cited by 1 Pith paper
-
Predicting protein folding dynamics using sequence information
A pipeline from sequence alignments through a Potts model to an Ising foldon chain predicts protein folding curves, subdomains, and mutation effects, but without new experimental validation in this paper.
Reference graph
Works this paper leans on
-
[1]
Chemical physics of protein folding,
P. G. Wolynes, W. A. Eaton, and A. R. Fersht, “Chemical physics of protein folding,” Proc. Natl. Acad. Sci. , vol. 109, no. 44, pp. 17770–17771, Oct. 2012, doi: 10.1073/pnas.1215733109
-
[2]
Protein folding funnels: a kinetic approach to the sequence-structure relationship.,
P. E. Leopold, M. Montal, and J. N. Onuchic, “Protein folding funnels: a kinetic approach to the sequence-structure relationship.,” Proc. Natl. Acad. Sci. , vol. 89, no. 18, pp. 8721–8725, Sep. 1992, doi: 10.1073/pnas.89.18.8721
-
[3]
Spin glasses and the statistical mechanics of protein folding,
J. D. Bryngelson and P. G. Wolynes, “Spin glasses and the statistical mechanics of protein folding,” Proc. Natl. Acad. Sci. , vol. 84, no. 21, pp. 7524–7528, 1987
work page 1987
-
[4]
E. Bornberg-Bauer and H. S. Chan, “Modeling evolutionary landscapes: Mutational stability, topology, and superfunnels in sequence space,” Proc. Natl. Acad. Sci. , vol. 96, no. 19, pp. 10689–10694, Sep. 1999, doi: 10.1073/pnas.96.19.10689
-
[5]
Statistical mechanics of simple models of protein folding and design,
V. S. Pande, A. Y. Grosberg, and T. Tanaka, “Statistical mechanics of simple models of protein folding and design,” Biophys. J. , vol. 73, no. 6, pp. 3192–3210, 1997, doi: 10.1016/S0006-3495(97)78345-0
-
[6]
D. U. Ferreiro, E. A. Komives, and P. G. Wolynes, “Frustration in biomolecules,” Q. Rev. Biophys. , vol. 47, no. 4, pp. 285–363, Nov. 2014, doi: 10.1017/S0033583514000092
-
[7]
Molecular Information Theory Meets Protein Folding,
I. E. Sánchez, E. A. Galpern, M. M. Garibaldi, and D. U. Ferreiro, “Molecular Information Theory Meets Protein Folding,” J. Phys. Chem. B , vol. 126, no. 43, pp. 8655–8668, Nov. 2022, doi: 10.1021/acs.jpcb.2c04532
-
[9]
Machine learning in protein structure prediction,
M. AlQuraishi, “Machine learning in protein structure prediction,” Curr. Opin. Chem. Biol. , vol. 65, pp. 1–8, Dec. 2021, doi: 10.1016/j.cbpa.2021.04.005
Show all 43 references
-
[10]
Contact order, transition state placement and the refolding rates of single domain proteins 1 1Edited by P. E. Wright,
K. W. Plaxco, K. T. Simons, and D. Baker, “Contact order, transition state placement and the refolding rates of single domain proteins 1 1Edited by P. E. Wright,” J. Mol. Biol. , vol. 277, no. 4, pp. 985–994, Apr. 1998, doi: 10.1006/jmbi.1998.1645
1998
-
[11]
Coarse-grained models of protein folding: toy models or predictive tools?,
C. Clementi, “Coarse-grained models of protein folding: toy models or predictive tools?,” Curr. Opin. Struct. Biol. , vol. 18, no. 1, pp. 10–15, Feb. 2008, doi: 10.1016/j.sbi.2007.10.005
2008 doi
-
[12]
Frustration, function and folding,
D. U. Ferreiro, E. A. Komives, and P. G. Wolynes, “Frustration, function and folding,” Curr. Opin. Struct. Biol. , vol. 48, pp. 68–73, Feb. 2018, doi: 10.1016/j.sbi.2017.09.006
2018 doi
-
[13]
Conserved residues and the mechanism of protein folding,
E. Shakhnovich, V. Abkevich, and O. Ptitsyn, “Conserved residues and the mechanism of protein folding,” Nature , vol. 379, no. 6560, pp. 96–98, Jan. 1996, doi: 10.1038/379096a0
1996 doi
-
[14]
Identification of direct residue contacts in protein-protein interaction by message passing,
M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, and T. Hwa, “Identification of direct residue contacts in protein-protein interaction by message passing,” Proc. Natl. Acad. Sci. U. S. A. , vol. 106, no. 1, pp. 67–72, 2009, doi: 10.1073/pnas.0805923106
2009 doi
-
[15]
Direct-coupling analysis of residue coevolution captures native contacts across many protein families,
F. Morcos et al. , “Direct-coupling analysis of residue coevolution captures native contacts across many protein families,” Proc. Natl. Acad. Sci. U. S. A. , vol. 108, no. 49, 2011, doi: 10.1073/pnas.1111471108
2011 doi
-
[16]
Inverse statistical physics of protein sequences: A key issues review,
S. Cocco, C. Feinauer, M. Figliuzzi, R. Monasson, and M. Weigt, “Inverse statistical physics of protein sequences: A key issues review,” Rep. Prog. Phys. , vol. 81, no. 3, 2018, doi: 10.1088/1361-6633/aa9965
2018 doi
-
[17]
Direct Coupling Analysis for Protein Contact Prediction,
F. Morcos, T. Hwa, J. N. Onuchic, and M. Weigt, “Direct Coupling Analysis for Protein Contact Prediction,” in Protein Structure Prediction , vol. 1137, D. Kihara, Ed., in Methods in Molecular Biology, vol. 1137. , New York, NY: Springer New York, 2014, pp. 55–70. doi: 10.1007/...
2014 doi
-
[18]
Coevolutionary Landscape Inference and the Context-Dependence of Mutations in Beta-Lactamase TEM-1,
M. Figliuzzi, H. Jacquier, A. Schug, O. Tenaillon, and M. Weigt, “Coevolutionary Landscape Inference and the Context-Dependence of Mutations in Beta-Lactamase TEM-1,” Mol. Biol. Evol. , vol. 33, no. 1, pp. 268–280, Jan. 2016, doi: 10.1093/molbev/msv211
2016 doi
-
[19]
Inferring repeat-protein energetics from evolutionary information,
R. Espada, R. G. Parra, T. Mora, A. M. Walczak, and D. U. Ferreiro, “Inferring repeat-protein energetics from evolutionary information,” PLoS Comput. Biol. , vol. 13, no. 6, pp. 1–16, 2017, doi: 10.1371/journal.pcbi.1005584
2017 doi
-
[20]
Size and structure of the sequence space of repeat proteins,
J. Marchi, E. A. Galpern, R. Espada, D. U. Ferreiro, A. M. Walczak, and T. Mora, “Size and structure of the sequence space of repeat proteins,” PLoS Comput. Biol. , vol. 15, no. 8, pp. 1–23, 2019, doi: 10.1371/journal.pcbi.1007282
2019 doi
-
[21]
Epistatic contributions promote the unification of incompatible models of neutral molecular evolution,
J. A. De La Paz, C. M. Nartey, M. Yuvaraj, and F. Morcos, “Epistatic contributions promote the unification of incompatible models of neutral molecular evolution,” Proc. Natl. Acad. Sci. , vol. 117, no. 11, pp. 5873–5882, Mar. 2020, doi: 10.1073/pnas.1913071117
2020 doi
-
[22]
Emergent time scales of epistasis in protein evolution,
L. Di Bari, M. Bisardi, S. Cotogno, M. Weigt, and F. Zamponi, “Emergent time scales of epistasis in protein evolution,” Proc. Natl. Acad. Sci. , vol. 121, no. 40, p. e2406807121, Oct. 2024, doi: 10.1073/pnas.2406807121
2024 doi
-
[23]
Kinetic coevolutionary models predict the temporal emergence of HIV-1 resistance mutations under drug selection pressure,
A. Biswas, I. Choudhuri, E. Arnold, D. Lyumkis, A. Haldane, and R. M. Levy, “Kinetic coevolutionary models predict the temporal emergence of HIV-1 resistance mutations under drug selection pressure,” Proc. Natl. Acad. Sci. , vol. 121, no. 15, p. e2316662121, Apr. 2024, doi: 10...
2024 doi
-
[24]
Foldons, protein structural modules, and exons.,
A. R. Panchenko, Z. Luthey-Schulten, and P. G. Wolynes, “Foldons, protein structural modules, and exons.,” Proc. Natl. Acad. Sci. , vol. 93, no. 5, pp. 2008–2013, Mar. 1996, doi: 10.1073/pnas.93.5.2008
2008 doi
-
[25]
Reassessing the exon–foldon correspondence using frustration analysis,
E. A. Galpern, H. Jaafari, C. Bueno, P. G. Wolynes, and D. U. Ferreiro, “Reassessing the exon–foldon correspondence using frustration analysis,” Proc. Natl. Acad. Sci. , vol. 121, no. 28, p. e2400151121, Jul. 2024, doi: 10.1073/pnas.2400151121
2024 doi
-
[26]
Evolution and folding of repeat proteins,
E. A. Galpern, J. Marchi, T. Mora, A. M. Walczak, and D. U. Ferreiro, “Evolution and folding of repeat proteins,” Proc. Natl. Acad. Sci. , vol. 119, no. 31, p. e2204131119, 2022, doi: 10.1073/pnas.2204131119
2022 doi
-
[27]
The energy landscapes of repeat-containing proteins: Topology, cooperativity, and the folding funnels of one-dimensional architectures,
D. U. Ferreiro, A. M. Walczak, E. A. Komives, and P. G. Wolynes, “The energy landscapes of repeat-containing proteins: Topology, cooperativity, and the folding funnels of one-dimensional architectures,” PLoS Comput. Biol. , vol. 4, no. 5, 2008, doi: 10.1371/journal.pcbi.1000070
2008 doi
-
[28]
Learning protein constitutive motifs from sequence data,
J. Tubiana, S. Cocco, and R. Monasson, “Learning protein constitutive motifs from sequence data,” eLife , vol. 8, 2019, doi: 10.7554/eLife.39397
2019 doi
-
[29]
Coevolutionary information, protein folding landscapes, and the thermodynamics of natural selection,
F. Morcos, N. P. Schafer, R. R. Cheng, J. N. Onuchic, and P. G. Wolynes, “Coevolutionary information, protein folding landscapes, and the thermodynamics of natural selection,” Proc. Natl. Acad. Sci. , vol. 111, no. 34, pp. 12408–12413, 2014, doi: 10.1073/pnas.1413575111
2014 doi
-
[30]
Selection originating from protein stability/foldability: Relationships between protein folding free energy, sequence ensemble, and fitness,
S. Miyazawa, “Selection originating from protein stability/foldability: Relationships between protein folding free energy, sequence ensemble, and fitness,” J. Theor. Biol. , vol. 433, pp. 21–38, 2017, doi: 10.1016/j.jtbi.2017.08.018
2017 doi
-
[31]
Refolding of Escherichia coli dihydrofolate reductase: sequential formation of substrate binding sites.,
C. Frieden, “Refolding of Escherichia coli dihydrofolate reductase: sequential formation of substrate binding sites.,” Proc. Natl. Acad. Sci. , vol. 87, no. 12, pp. 4413–4416, Jun. 1990, doi: 10.1073/pnas.87.12.4413
1990 doi
-
[32]
Folding of dihydrofolate reductase from Escherichia coli,
N. A. Touchette, K. M. Perry, and C. R. Matthews, “Folding of dihydrofolate reductase from Escherichia coli,” Biochemistry , vol. 25, no. 19, pp. 5445–5452, Sep. 1986, doi: 10.1021/bi00367a015
1986 doi
-
[33]
Thermal unfolding molecular dynamics simulation of Escherichia coli dihydrofolate reductase: Thermal stability of protein domains and unfolding pathway,
Y. Y. Sham, B. Ma, C. Tsai, and R. Nussinov, “Thermal unfolding molecular dynamics simulation of Escherichia coli dihydrofolate reductase: Thermal stability of protein domains and unfolding pathway,” Proteins Struct. Funct. Bioinforma. , vol. 46, no. 3, pp. 308–320, Feb. 2002,...
2002 doi
-
[34]
Microsecond Subdomain Folding in Dihydrofolate Reductase,
M. Arai, M. Iwakura, C. R. Matthews, and O. Bilsel, “Microsecond Subdomain Folding in Dihydrofolate Reductase,” J. Mol. Biol. , vol. 410, no. 2, pp. 329–342, Jul. 2011, doi: 10.1016/j.jmb.2011.04.057
2011 doi
-
[35]
Structure of a partially unfolded form of E scherichia coli dihydrofolate reductase provides insight into its folding pathway,
J. R. Kasper, P. Liu, and C. Park, “Structure of a partially unfolded form of E scherichia coli dihydrofolate reductase provides insight into its folding pathway,” Protein Sci. , vol. 23, no. 12, pp. 1728–1737, Dec. 2014, doi: 10.1002/pro.2555
2014 doi
-
[36]
Co-Evolutionary Fitness Landscapes for Sequence Design,
P. Tian, J. M. Louis, J. L. Baber, A. Aniana, and R. B. Best, “Co-Evolutionary Fitness Landscapes for Sequence Design,” Angew. Chem. - Int. Ed. , vol. 57, no. 20, pp. 5674–5678, 2018, doi: 10.1002/anie.201713220
2018 doi
-
[37]
Solvent constraints for biopolymer folding and evolution in extraterrestrial environments,
I. E. Sánchez, E. A. Galpern, and D. U. Ferreiro, “Solvent constraints for biopolymer folding and evolution in extraterrestrial environments,” Proc. Natl. Acad. Sci. , vol. 121, no. 21, p. e2318905121, May 2024, doi: 10.1073/pnas.2318905121
2024 doi
-
[38]
Quantitative criteria for native energetic heterogeneity influences in the prediction of protein folding kinetics,
S. S. Cho, Y. Levy, and P. G. Wolynes, “Quantitative criteria for native energetic heterogeneity influences in the prediction of protein folding kinetics,” Proc. Natl. Acad. Sci. , vol. 106, no. 2, pp. 434–439, Jan. 2009, doi: 10.1073/pnas.0810218105
2009 doi
-
[39]
Start2Fold: A database of hydrogen/deuterium exchange data on protein folding and stability,
R. Pancsa, M. Varadi, P. Tompa, and W. F. Vranken, “Start2Fold: A database of hydrogen/deuterium exchange data on protein folding and stability,” Nucleic Acids Res. , vol. 44, no. D1, pp. D429–D434, 2016, doi: 10.1093/nar/gkv1185
2016 doi
-
[40]
Local energetic frustration conservation in protein families and superfamilies,
M. I. Freiberger et al. , “Local energetic frustration conservation in protein families and superfamilies,” Nat. Commun. , vol. 14, no. 1, p. 8379, Dec. 2023, doi: 10.1038/s41467-023-43801-2
2023 doi
-
[41]
The Pfam protein families database: towards a more sustainable future,
R. D. Finn et al. , “The Pfam protein families database: towards a more sustainable future,” Nucleic Acids Res. , vol. 44, no. D1, pp. D279–D285, Jan. 2016, doi: 10.1093/nar/gkv1344
2016 doi
-
[42]
InterPro in 2022,
T. Paysan-Lafosse et al. , “InterPro in 2022,” Nucleic Acids Res. , vol. 51, no. D1, pp. D418–D427, Jan. 2023, doi: 10.1093/nar/gkac993
2022 doi
-
[43]
Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences,
W. Li and A. Godzik, “Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences,” Bioinformatics , vol. 22, no. 13, pp. 1658–1659, Jul. 2006, doi: 10.1093/bioinformatics/btl158
2006 doi
-
[44]
A PDZ domain recapitulates a unifying mechanism for protein folding,
S. Gianni et al. , “A PDZ domain recapitulates a unifying mechanism for protein folding,” Proc. Natl. Acad. Sci. , vol. 104, no. 1, pp. 128–133, Jan. 2007, doi: 10.1073/pnas.0602770104. Supplemental Information for Inferring protein folding mechanisms from natural sequence div...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.