{"id":"16a42ce3-8a47-4bef-a188-098fd6c1d234","arxiv_id":"2607.20504","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Protein-folding memory friction is strongly coordinate-dependent, peaking in the folded state, and incorporating that dependence improves reduced-model folding kinetics for several fast-folding proteins.","lead":"This paper extracts a reaction-coordinate-dependent memory friction from long all-atom simulations of six fast-folding proteins and shows the friction is markedly stronger in folded states. It then simulates the resulting generalized Langevin equation and finds that including this coordinate dependence improves reduced-model folding kinetics for several proteins.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The diagonal approximation in Eq. 9 is not directly validated; the folded-state friction enhancement could be an artifact of neglected off-diagonal correlations.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the diagonal approximation in Eq. 9 is never directly tested. The paper itself promises a check ('validity of this approximation will be checked below') and instead performs an integrated comparison of τ_MFP, which cannot separate approximation error from fitting/embedding error. This is a genuine soft spot because the central quantitative claim—that RC-dependent friction is strongly enhanced in the folded state and improves kinetic predictions—rests on Γ(t,x) being extracted without bias. The qualitative direction of the result is plausible and consistent with prior internal-friction work, and the code availability is a credit. But the evidence as presented supports only a conditional acceptance: the method and qualitative RC-dependence are likely right, while the quantitative 'governs' claim needs the direct validation. Since the reader already recommended CONDITIONAL, my read does not change the verdict.","tokens_in":30546,"tokens_out":6070,"duration_ms":58844,"concrete_test":"Using the repository code, compute from the MD trajectories the double-conditional velocity correlation C_vv(s,x,x_s) and evaluate both sides of Eq. 8 with the final fitted Γ(t,x): L_full(t,x)=∫dx_s∫_0^t ds C_vv(s,x,x_s)Γ(t−s,x_s) versus L_diag(t,x)=∫_0^t ds C_vv(s,x)Γ(t−s,x). Plot the relative residual |L_full−L_diag|/|C_vv(t,x)−C_vv(0,x)| for barrier-region bins (x≈0.5–0.65) and folded bins. If the residual exceeds ~10% where γ_2(x) is steep, the Eq. 9 decoupling is invalid and the reported folded-state friction enhancement is not established. Alternatively, run the extraction pipeline on a synthetic GLE trajectory with known Γ_true(t,x) to confirm the method recovers Γ_true.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the decoupling C_vv(s,x,x_s)≈C_vv(s,x)δ(x−x_s) that converts the exact conditional Volterra equation (Eq. 8) into the per-x equation (Eq. 9). The text states 'the validity of this approximation will be checked below,' but the subsequent check is the global τ_MFP comparison in Fig. 4 after fitting Γ(t,x) and simulating the embedding. That is a joint test of the diagonal ansatz, the two-exponential fit, the Markovian embedding, and m(x) fits; it cannot localize errors from the omitted off-diagonal correlations. In barrier bins, x(t) changes rapidly, so C_vv(s,x,x_s) will have substantial weight at x_s≠x. If those terms are non-negligible, the extracted Γ(t,x) absorbs them, biasing the reported folded-state peak in γ_2(x) (Fig. 3G–L). Since the paper's quantitative claim is that RC-dependent memory friction 'governs' kinetics, this unvalidated approximation is the central weak point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an extraction method for reaction-coordinate-dependent non-Markovian friction in a generalized Langevin equation (GLE) with position-dependent mass and memory kernel. Starting from a conditional Volterra equation derived from an exact GLE, the authors approximate the double-conditional velocity autocorrelation as diagonal in x, extract Γ(t,x) from all-atom MD trajectories of six fast-folding proteins, and report that the memory friction increases strongly toward the folded state. They fit the kernels to two-exponential forms, embed them in a Markovian scheme, and compare simulated mean first-passage times with MD. The central claim is that including RC-dependent memory friction improves the low-dimensional description of protein folding kinetics and provides evidence for internal friction in compact states.","tokens_in":30880,"tokens_out":4568,"duration_ms":46013,"significance":"If the extraction is valid, this is a useful contribution: it gives a data-driven route to spatially resolved non-Markovian friction, connects to internal-friction interpretations, and supplies open-source code. The paper also derives a Fokker–Planck stationary distribution for the embedding and checks marginal distributions analytically. The multi-protein comparison is a strength, and the raw observation that the memory kernel rises in the folded basin is physically plausible. However, the quantitative claims rest on an unvalidated diagonal approximation, a two-exponential fit whose slow time constants exceed the extraction window for several proteins, and an embedding whose correction factor is not negligible. These issues need to be resolved before the headline conclusions are fully supported.","major_comments":[{"comment":"The load-bearing step is the replacement C_vv(s,x,x_s) ≈ C_vv(s,x)δ(x−x_s). The text says 'the validity of this approximation will be checked below,' but the subsequent check is the joint τ_MFP comparison in Fig. 4, which tests the diagonal ansatz together with the two-exponential fit, the mass fit, and the embedding. It cannot localize errors from omitted off-diagonal correlations. In barrier bins, x(t) changes rapidly, so C_vv(s,x,x_s) will have substantial weight at x_s≠x. If those terms contribute, the extracted Γ(t,x), especially the folded-state peak in γ2(x), is biased. Please provide a direct test of the off-diagonal terms, e.g., evaluate the omitted integral in Eq. (8) or compare the diagonal-extracted Γ with a solution of the full Volterra equation in a simplified model.","section":"Main text, Eq. (8) to Eq. (9)"},{"comment":"The slow memory times τ2 exceed the maximum extraction window for several proteins: Villin τ2=46.7 ns with t_max≤16 ns, and λ-repressor τ2=64.3 ns with t_max≤16 ns. Protein G and α3D have τ2≈14–15 ns with t_max≈16 ns, which is marginal. A two-exponential fit whose slow time constant is several times longer than the fitted time window is not identified by the data; the reported rise of γ2(x) toward the folded state may be an extrapolation artifact of the fit, not a feature of the MD data. The SI short-vs-long window comparison (Table S4) is for coordinate-independent kernels only and does not test the RC-dependent fits. Please show that the γ2(x) trends are robust to the fitting window, e.g., by fixing τ2 to values above and below the extraction window and rerunning the τ_MFP comparison.","section":"SI §III, Fig. S1 and Fig. 3G–L"},{"comment":"The Markovian embedding is claimed to reproduce the target kernel, but Eq. (S30) contains an exponential correction factor exp(∫ γ'_i/(2γ_i) v ds'), which is assumed to be ≈1 because ⟨v⟩=0. Figure S8 shows that this factor ranges between about 0.5 and 3 in the Villin simulation, and Fig. S9 shows deviations E that are non-negligible for some frames. Since the GLE simulations are used as validation of the extracted parameters, this embedding error propagates into Fig. 4 and weakens the quantitative comparison with MD. Please quantify the effect of this factor on the memory kernel and on the computed τ_MFP values, or modify the embedding to remove the approximation.","section":"SI §IV.B, Eq. (S30), Fig. S8"}],"minor_comments":[{"comment":"The color coding of the x-dependent curves is described only as 'blue unfolded and yellow folded state'; please add a color bar or explicit legend so the reader can map colors to x values.","section":"Fig. 2 caption"},{"comment":"The statement 'the GLEs in Eqs. 2 and 3 are exact' may be misinterpreted: Eq. (3) is an exact form for the projected dynamics only with the appropriate projection operator, and the extraction in this paper uses an additional approximation. Please qualify the wording to distinguish the exact GLE form from the approximate extraction scheme.","section":"Main text after Eq. (3)"},{"comment":"Some fitted mass parameters are extremely small or large (e.g., θ1=1.724e-96 for Villin), suggesting poor conditioning of the fit parameterization. A more stable basis for m(x) would improve reproducibility.","section":"SI Table S2"},{"comment":"The figure compares simulated and MD τ_MFP curves without error bars or quantitative error metrics. Given that the central claim is that RC-dependent friction improves the description, please report a quantitative measure (e.g., log-mean-squared error) and, where feasible, bootstrap uncertainties.","section":"Fig. 4"},{"comment":"Reference [42] is an arXiv preprint; if the derivation of Eqs. (2) and (3) relies on it, please ensure the reference is to a published version or peer-reviewed source, and state the extent of the reliance.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely of interest to the readership, and the raw data analysis is carefully done, but the validation strategy is circular in a specific way: the diagonal approximation is tested only through a joint simulation comparison. The slow-time-constant issue and the embedding correction factor are concrete and fixable, but they are load-bearing. I would not reject the paper, but it needs either direct tests of the diagonal approximation or a restatement of the claims to match the level of validation provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is this: the paper gives a genuinely new data-driven method to extract reaction-coordinate-dependent memory friction from molecular dynamics, applies it to six fast-folding proteins, and finds that friction is higher in the folded state than in the unfolded state. I think that qualitative result is probably right. The quantitative claim — that RC-dependent non-Markovian friction 'governs' folding dynamics — is stronger than the evidence supports.\n\nWhat's actually new and good: the conditional Volterra scheme (Eq. 10) is a natural extension of earlier RC-independent extractions, but this is the first systematic application to protein folding with RC-dependent kernels. The authors use long Anton trajectories, give code on GitHub, and check their memory kernels against prior RC-independent results. The Markovian embedding in Eq. 13 is non-trivial; they test it against the stationary distribution and show the correction factor is of order one. The mean first-passage-time comparison is a meaningful consistency check: they did not fit τ_MFP directly, and the RC-dependent GLE clearly outperforms the RC-independent one for Villin, WW domain, and Trp-cage.\n\nThe soft spots are real, and the first one is load-bearing. The diagonal approximation in Eq. 9 — replacing the double-conditional velocity autocorrelation by a delta-function in x — is asserted, with the promise 'the validity of this approximation will be checked below.' The check turns out to be the global τ_MFP agreement after fitting everything. That's a joint test of the diagonal ansatz, the two-exponential fit, the mass fit, and the embedding; it cannot localize bias from the omitted off-diagonal terms. In the barrier region, where x changes rapidly over the memory time, those terms are not obviously small. If they are not small, the extracted Γ(t,x) absorbs them, and the reported folded-state peak in γ₂ could be an artifact. This is not a fatal flaw, but it is exactly the step that separates a method paper from a strong method paper.\n\nTwo smaller issues: for Villin and λ-repressor the slow memory time τ₂ (46.7 ns and 64.3 ns) exceeds the longest extraction window (up to 40 ns), so the slow-component amplitude is partly extrapolated. And there are no error bars or uncertainty estimates anywhere in the extraction. The validation also uses the same MD trajectories for fitting and testing; it is a self-consistency check, not a prediction on held-out data.\n\nProportionately: the core idea is sound, the implementation is careful, and the authors are honest about λ-repressor, where the RC-independent model does slightly better. But the abstract's 'governs' is too strong given the unvalidated approximation and the counterexample. The paper deserves serious peer review — conditional on the authors directly testing the diagonal approximation (e.g., computing off-diagonal contributions in the barrier bins), adding uncertainty quantification, and preferably validating on trajectories not used for fitting. I would bring it to the reading group, and I'd cite it for the method even if I'm not ready to cite the quantitative claim.","headline":"Genuinely new data-driven extraction of coordinate-dependent memory friction for protein folding, with a plausible qualitative result and a real but addressable weakness in the diagonal approximation.","tokens_in":31316,"tokens_out":3596,"would_cite":true,"duration_ms":33623,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.15.hm","05.40.-a"],"model":"deepseek-v4-flash","headline":"The paper shows that the memory friction governing protein folding depends strongly on the reaction coordinate and is largest in the folded, compact state, and that including this dependence in generalized Langevin equation simulations mark","keywords":["protein folding","memory friction","generalized Langevin equation","reaction coordinate","non-Markovian dynamics","conditional Volterra equation","internal friction","mean first-passage time"],"falsifier":"Run a controlled test with a known position-dependent memory kernel: simulate long trajectories from the GLE, then apply the paper's diagonal conditional-Volterra extraction and compare the recovered kernel with the true one, especially in the barrier region where the reaction coordinate changes rapidly. A systematic mismatch there would falsify the method; alternatively, compute the full double-conditional correlation in the MD data and show that the off-diagonal terms are non-negligible.","tokens_in":30480,"feed_emoji":"🧬","tokens_out":4689,"duration_ms":41947,"temperature":0.7,"pith_summary":"The paper sets out to show that the friction memory function in a generalized Langevin equation description of protein folding is strongly coordinate-dependent, not a single function of time. Using a conditional Volterra equation applied to long all-atom molecular dynamics trajectories of six fast-folding proteins, the authors extract a coordinate-dependent memory kernel and an effective coordinate-dependent mass. They find that friction is markedly larger in the folded state than in the unfolded state, consistent with internal friction in compact conformations. When the coordinate-dependent kernel is fed into a Markovian-embedded GLE simulation, the mean first-passage times between folded and unfolded states match the molecular-dynamics reference for most proteins, and for proteins with pronounced coordinate dependence they match substantially better than the standard coordinate-independent GLE. The claim is that coordinate-dependent non-Markovian friction is a fundamental feature of protein-folding dynamics, not a minor correction.","feed_headline":"Protein folding friction peaks in the folded state","feed_subtitle":"A position-dependent memory kernel extracted from MD data beats one-size-fits-all friction in reproducing folding times.","key_machinery":"The central object is the conditional Volterra equation (Eq. 10), which relates the running integral G(t,x) of the memory kernel to single-position-conditioned velocity autocorrelation functions C_vv(t,x) and velocity-force correlations C_vF(t,x). Its practical feasibility rests on the diagonal approximation that replaces the double-conditional velocity correlation C_vv(s,x,x_s) by C_vv(s,x)δ(x−x_s), decoupling the equations per position x. The extracted kernels are fitted to factorized multi-exponentials and simulated via a Markovian embedding that samples the correct stationary distribution, allowing direct comparison of mean first-passage times with MD.","core_discovery":"The central discovery is that the memory kernel extracted from MD trajectories of six fast-folding proteins varies significantly with the reaction coordinate, rising towards the folded state, with the slow friction component dominating barrier-crossing kinetics. The paper further shows that a GLE with both coordinate-dependent mass and coordinate-dependent memory friction (Eq. 3) reproduces the MD mean first-passage times at least as well as, and for several proteins better than, the commonly used coordinate-independent GLE (Eq. 1). For one protein (λ-repressor), which has a largely coordinate-independent friction profile, the simpler model performs comparably or slightly better, indicating","pith_inferences":["Because the diagonal approximation is validated only indirectly through mean first-passage times, a direct check on the magnitude of off-diagonal correlations in the double-conditional velocity correlation could strengthen or revise the extracted friction profile; if those terms are non-negligible near the barrier, the reported folded-state enhancement could be biased upward.","The method should be portable to single-molecule force spectroscopy experiments where position-dependent diffusion coefficients have been inferred; a position-dependent memory kernel would predict refolding time distributions different from the Markovian models currently fitted to such data.","If the slow folded-state friction component is the dominant kinetic control, mutations that alter the compactness of the folded state should measurably change folding rates through the friction term, a testable prediction."],"forward_implications":["For proteins with strong coordinate dependence, coordinate-independent GLE simulations fail to reproduce MD folding and unfolding kinetics; including the coordinate-dependent kernel restores agreement.","The slow memory component, with timescales from about 10 ns to 60 ns, dominates barrier crossing and is largest in the folded state, locating the kinetic effect of internal friction in the folded basin.","The coordinate-dependent mass also rises toward the folded state, and the effective potential used in the GLE produces the correct stationary distribution in position and velocity.","The authors state that the framework should extend naturally to larger proteins and other macromolecular systems."],"fun_headline_variants":["Protein folding friction depends on reaction coordinate","Folding friction is position-dependent, peaks in folded state","Memory friction varies along folding coordinate, boosts kinetics","New GLE method: position-specific friction improves folding description","Reaction-coordinate-dependent friction key to folding dynamics"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The extracted friction profile rests on the assumption that, when the trajectory is conditioned on a starting position, velocity correlations at two different later positions are negligible; if those off-diagonal correlations matter, the position-dependent kernel and the reported folded-state friction enhancement would be biased.","fun_headline_variants_meta":{"raw":{"variants":["Protein folding friction depends on reaction coordinate","Folding friction is position-dependent, peaks in folded state","Memory friction varies along folding coordinate, boosts kinetics","New GLE method: position-specific friction improves folding description","Reaction-coordinate-dependent friction key to folding dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1247,"prompt_tokens":727,"completion_tokens":520,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":446}},"tokens_in":471,"tokens_out":520,"duration_ms":5391,"temperature":1.0,"reasoning_tokens":446,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T09:19:59.246015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled test with a known position-dependent memory kernel: simulate long trajectories from the GLE, then apply the paper's diagonal conditional-Volterra extraction and compare the recovered kernel with the true one, especially in the barrier region where the reaction coordinate changes rapidly. A systematic mismatch there would falsify the method; alternatively, compute the full double-conditional correlation in the MD data and show that the off-diagonal terms are non-negligible.","supporting_citations":[],"review_version":1}