{"id":"f2e1594c-f0a9-4ef1-a390-7ed3f587d803","arxiv_id":"2608.13219","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Prescribed-SST simulations and 30-year moving-window regressions cannot reliably recover time-varying Earth's feedback, with apparent trends under ~100 years largely statistical noise.","lead":"A new controlled experiment shows that the standard 'amip-piForcing' method for estimating Earth's feedback parameter from prescribed ocean temperatures fails to reproduce the feedback of the coupled climate model, even when ocean temperatures are identical. The paper also shows that multi-decadal trends in the feedback computed with 30-year moving windows can arise purely from statistical noise, casting doubt on recent published claims of a stabilizing pattern effect.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ~100-year noise floor is likely set by the 30-year moving-window estimator itself; the conclusion needs a window-length/estimator robustness test.","rationale":"The paper is a well-designed controlled experiment, and the demonstration that amip-piForcing fails to reproduce the coupled feedback with identical SSTs is convincing. The attribution to missing forcing is supported by the amip-histForcing improvement. My main concern is narrower: the quantitative noise floor of 60-100 years is computed with one estimator (30-year overlapping moving-window regression). Because overlapping windows impose autocorrelation at the window scale, the spectral turnover in Fig. 7 likely reflects the estimator, not the climate. This does not invalidate the paper's qualitative point that moving-window regressions create spurious low-frequency variability, but it does mean the claim 'any trends shorter than ~100 years are indistinguishable from noise' is overgeneralized. The reader's weakest assumption identified this estimator dependence; I agree partially and propose a concrete window-length/estimator test. If the noise floor scales with window length, the paper should be revised to state the bound as specific to the moving-window estimator. Thus the CONDITIONAL verdict remains appropriate; no change needed beyond the existing condition.","tokens_in":33874,"tokens_out":5466,"duration_ms":51612,"concrete_test":"Recompute the Sec. 2.3 Monte-Carlo spectra and the piControl feedback time series using moving-window regressions with window lengths of 15, 30, and 60 years, and also using non-overlapping 30-year windows and a Kalman-filter/state-space estimate. If the spectral turnover ('noise floor') shifts with window length or disappears for non-overlapping windows, the ~100-year bound is an artifact of the 30-year overlapping-window estimator, and the paper's conclusion must be restricted to that estimator. If the turnover remains at 60-100 years across all estimators, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that trends in the feedback parameter on timescales shorter than ~100 years are indistinguishable from statistical noise (Sec. 3.2.1, Conclusions). This noise floor is derived entirely from the 30-year moving-window regression of Eq. 6/7: the Monte-Carlo spectra in Fig. 7b,d are computed by applying that estimator to synthetic N and T. But a moving-window slope with overlapping windows has an autocorrelation structure imposed by the window length: estimates one year apart share 29 of 30 data points, so the estimated slope series cannot decorrelate faster than the window scale. The 60-100 year spectral turnover in Fig. 7 is therefore plausibly a property of the estimator, not of the underlying N,T processes. The paper does not vary the window length or compare with a non-overlapping or state-space estimator, so the headline bound is not established as a property of the climate system. This matters because the abstract and conclusions generalize the bound to 'any trends ... shorter than ~100 years' without the estimator caveat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a controlled set of CESM2/CAM6 experiments to test two assumptions underlying the widely used amip-piForcing method for estimating the time evolution of Earth's climate feedback parameter: (i) that prescribing observed SSTs and sea ice with preindustrial atmospheric forcing captures all relevant forcing effects, and (ii) that 30-year moving-window regressions of radiative response on temperature yield a feedback time series whose variations are physically interpretable and SST-driven. The authors find that amip-piForcing-style runs forced with SSTs from coupled historical simulations fail to reproduce the coupled feedback evolution (ensemble-mean Pearson correlation 0.23, MSE 0.90 W² m⁻⁴ K⁻²), and that prescribing historical forcing as well (amip-histForcing) substantially improves agreement (correlation 0.64) but leaves unexplained residual differences. Using piControl and piClim-control simulations together with Monte-Carlo VAR(5) and bivariate-normal surrogates, they then show that the moving-window feedback estimator itself generates low-frequency variability via the Yule-Slutzky effect, with spectral power increasing toward periods of 60–100 years.","tokens_in":34133,"tokens_out":6939,"duration_ms":64782,"significance":"If the results hold, this is an important contribution to the pattern-effect and climate-feedback literature. The strengths include a clean 'perfect model' experimental design that isolates the effect of missing atmospheric forcing from SST-pattern effects, an 11-member ensemble for the historical and GHG-only cases, three independent forcing estimates with largely consistent results, and a Monte-Carlo framework that convincingly reproduces the feedback spectrum from statistical surrogates. The paper also explicitly enumerates its limitations (single GCM, possible effects of severed atmosphere-ocean coupling, and the signal-to-noise problem), which is commendable. The demonstration that amip-piForcing feedback estimates are biased and partly statistically generated is significant for interpreting published claims of a stabilizing SST-driven feedback trend in recent decades. However, the central quantitative claim about a ~100-year noise floor rests on a specific estimator and lacks a robustness test, and the null result on SST-pattern effects is weaker than the wording suggests; these issues need to be resolved before the conclusions can be accepted in their current general form.","major_comments":[{"comment":"The ~100-year noise floor is computed using the 30-year moving-window regression of Eq. (6), and because the windows overlap by 29 out of 30 years, the estimated feedback series has autocorrelation imposed by the window length itself. The 60–100 year spectral turnover in Fig. 7b,d is therefore plausibly a property of this estimator rather than of the underlying N and T processes. The paper does not vary the window length, nor does it compare with a non-overlapping or state-space estimator, so the generalized statement in the Abstract and Conclusions that 'any trends in the feedback parameter detected on timescales shorter than ~100 years are indistinguishable from statistical noise' is not established as a property of the climate system. Please add a sensitivity analysis in which the Monte-Carlo spectra are recomputed with, e.g., 15-year and 60-year windows (or with a non-overlapping-window or state-space estimator), and revise the abstract and conclusions to state the noise floor is specific to the 30-year moving-window estimator used here.","section":"Sec. 3.2.1, Fig. 7; Abstract; Sec. 5"},{"comment":"The conclusion that there is no evidence of an unforced SST pattern effect rests on selecting periods of feedback strengthening and weakening from the moving-window feedback series. If, as the authors themselves argue, a large part of the variance in that series is statistical noise from the Yule-Slutzky effect, then the selected periods may not correspond to genuine feedback changes, which would make the composite SST-trend test insensitive by construction. The authors partly acknowledge this in Sec. 4.2, but the stronger statement in Sec. 5 ('we find no evidence of an unforced SST pattern effect in the CESM2 piControl simulation') goes beyond what the test can support. Please either soften the claim to reflect the test's limited power under a noisy feedback estimator, or add a test that explicitly accounts for the noise floor (for example, by comparing the observed composite trend distribution against the Monte-Carlo noise distribution from Sec. 3.2.1).","section":"Sec. 3.2.2, Fig. 8"}],"minor_comments":[{"comment":"The 'Code and data availability' section currently contains only the placeholder text 'TEXT.' A complete statement describing where the simulation outputs and analysis code can be obtained is required for reproducibility.","section":"Code and Data"},{"comment":"Please clarify whether the moving-window regression includes an intercept. In the Gregory framework N = F + λT, the slope of a regression with intercept is the standard estimate, but the notation ∂N/∂T and 'regression of N onto ΔT' is ambiguous; this detail affects the numerical values of λ(t).","section":"Eqs. (6)-(7), Sec. 2.2"},{"comment":"The sentence 'Any trends in the feedback parameter detected on timescales shorter than ~100 years are indistinguishable from statistical noise' would be better phrased as '...indistinguishable from statistical noise when the feedback is estimated with the 30-year moving-window regression method considered here,' consistent with the recommended robustness changes.","section":"Abstract"},{"comment":"Panels (c) and (d) are described in the text as 'Bias in feedback' and the vertical axis shows 'Feedback (Wm-2K-1),' but the plotted quantity is the difference between the two feedback curves; please label the axis 'Feedback bias' or 'Difference' for clarity.","section":"Fig. 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and addresses a question of broad interest to the climate feedback community. The main technical concern is the estimator-dependence of the headline ~100-year noise floor; this is fixable with a modest addition of sensitivity experiments. I also note the data-availability placeholder, which should be resolved before acceptance. The single-GCM limitation is acknowledged by the authors and is acceptable if the claims are appropriately scoped."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The core experiment is clean: prescribe the coupled model's own SSTs to its atmosphere-only version, so the only difference is the protocol itself. The feedback time series diverge substantially (corr 0.23), and adding historical forcing closes about half the gap (corr 0.64). The other half is statistical: 30-year moving-window regressions generate spurious multi-decadal variability even when N and T are white noise. That is the Yule-Slutzky effect, and it is an underappreciated problem for the amip-piForcing literature.\n\nWhat is genuinely new: the pairwise identical-SST design across 11 members, the high- and low-pass decomposition of the feedback discrepancy, and the Monte-Carlo demonstration that the estimator produces a 60-100 year spectral floor. The 4000-year piControl composite with proper multiple-testing correction is a useful null result: no coherent SST pattern accompanies feedback trends in an unforced run.\n\nSoft spots, in proportion. Single GCM, and the authors say so. The ~100-year bound is tied to the 30-year window; a different window length or a state-space estimator would shift it. The paper does not test this, and the abstract's \"any trends...shorter than ~100 years\" overgeneralizes. But the body is more careful, and the qualitative conclusion does not depend on the exact number. The moving-window estimator does inject noise on multi-decadal scales; that is the point. There is no code or data released, and the \"Code and data availability\" line in the preprint is a placeholder. That should be fixed. The residual differences in amip-histForcing are left unexplained, with plausible but untested speculation about severed atmosphere-ocean coupling.\n\nBottom line: this is a serious paper that deserves a real referee. The experimental design is careful, the results are potentially important for the pattern-effect and effective-climate-sensitivity literature, and the limitations are mostly acknowledged. The authors should be asked to add a window-length robustness test, release the code and data, and soften the abstract's unqualified wording. I would take it to our reading group and would cite it if I worked on feedback estimation.","headline":"A careful controlled experiment showing that amip-piForcing feedback trends are part missing-forcing bias and part moving-window noise; the ~100-year noise floor is estimator-specific, but the case against the method survives.","tokens_in":34665,"tokens_out":4128,"would_cite":true,"duration_ms":39450,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that prescribing observed sea-surface temperatures to an atmospheric model cannot recover the evolution of Earth's climate feedback, and that feedback trends on timescales shorter than about a century are statistically…","keywords":["climate feedback parameter","pattern effect","amip-piForcing","moving-window regression","Yule-Slutzky effect","sea-surface temperature patterns","climate sensitivity","CESM2"],"falsifier":"A reader could compute feedback time series from the 4000-year piControl using a 50-year or 100-year moving window, or a state-space estimator; if trends shorter than 100 years then emerge as distinguishable from the noise floor, the paper's central bound fails for that estimator.","tokens_in":33716,"feed_emoji":"🌡️","tokens_out":5778,"duration_ms":92282,"temperature":0.7,"pith_summary":"This paper tests the standard 'amip-piForcing' method for estimating how Earth's climate feedback parameter has changed over the observational record, a method that prescribes observed sea-surface temperatures to an atmospheric model with fixed pre-industrial forcing. By prescribing SSTs from fully coupled historical simulations back into the same atmospheric model, the authors show the method fails to reproduce the coupled model's feedback evolution even when the SSTs are identical. They identify two causes: omitting time-varying atmospheric forcing, especially aerosols, biases the feedback, and the 30-year moving-window regression used to estimate the feedback time series injects statistical noise via the Yule-Slutzky effect. Any feedback trends on timescales shorter than about 100 years are indistinguishable from that noise, and a 4000-year control run shows no significant link between evolving SST patterns and feedback changes. If right, recent claims that observed SST patterns caused a stabilizing trend in Earth's feedback over recent decades are not robust.","feed_headline":"Prescribed-SST runs miss Earth's true feedback trend","feed_subtitle":"Even with identical ocean temperatures, atmosphere-only runs diverge from coupled runs; short feedback trends are noise.","key_machinery":"The working object is the feedback parameter $\\lambda(t)$, computed by 30-year moving-window regression of the radiative response $R = N - F$ onto temperature change $\\Delta T$. The method's second assumption is that variations in this slope are driven by the evolving SST pattern. The paper's key mechanistic finding is the Yule-Slutzky effect: applying a moving-window regression to even pure white-noise time series of $N$ and $T$ generates low-frequency variability in the slope, with a decorrelation timescale of 60 to 100 years. This statistical machinery, rather than any physical memory, explains much of the apparent multi-decadal structure in $\\lambda(t)$, as demonstrated by million-year Monte-Carlo simulations fitted to the piClim-control and piControl data. The composite trend analysis on the 4000-year control run then tests whether the SST pattern is nonetheless detectable behind that noise.","core_discovery":"The central discovery is that the amip-piForcing setup, an atmosphere-only model forced with observed SSTs and sea ice while atmospheric forcing stays at pre-industrial levels, cannot recover the time evolution of Earth's feedback parameter even in a perfect-model setting. Forced with SSTs and sea ice from eleven fully coupled historical simulations, the atmosphere-only runs produce a feedback time series with an ensemble-mean correlation of only 0.23 against the coupled feedback, and a systematic negative bias amplifies around 1940 to 2000, the period of strongest aerosol forcing. When historical atmospheric forcing is also prescribed, agreement improves to a correlation of 0.64, but individual members still diverge substantially, showing that SSTs alone do not determine the feedback. In a 4000-year pre-industrial control, periods of feedback strengthening or weakening show no significant or spatially consistent SST trend pattern, and Monte-Carlo experiments show that 30-year moving-window regressions of two white-noise series generate multi-decadal feedback variability peaking at 60 to 100 years. The conclusion is that trends in the feedback parameter on timescales shorter than roughly a century cannot be robustly attributed to evolving SST patterns.","pith_inferences":["The 60-to-100-year noise floor is a property of the 30-year moving-window estimator; a longer window or a state-space estimator would shift the floor, so the paper's 'less than ~100 years' bound should be read as estimator-specific rather than as a fixed property of the climate system.","The same Yule-Slutzky logic applies to any moving-window regression of ratio variables in climate science, such as carbon-cycle sensitivities or transient climate response estimates, where short windows can generate spurious low-frequency structure.","A direct testable extension is to rerun the identical-SST protocol in several other coupled models with and without prescribed forcing; if some models recover the coupled feedback under amip-piForcing, the failure is model-specific rather than intrinsic to the method.","The bootstrap composite procedure could be applied to observed SST and feedback data: if observed periods of feedback change show SST pattern agreement stronger than the 4000-year control null, that would be evidence that the observed pattern effect is exceptional rather than absent."],"forward_implications":["Recent amip-piForcing-based conclusions that observed SST patterns drove a stabilizing feedback trend in recent decades are not supported once the coupled feedback is used as reference.","Any feedback trend detected over periods shorter than about 100 years could be a statistical artifact of the moving-window regression, independent of any physical SST-pattern effect.","amip-piForcing-style experiments should be replaced by amip-histForcing-style runs with time-varying historical atmospheric forcing to reduce the bias.","Even with correctly prescribed SSTs, sea ice, and forcing, atmosphere-only runs do not fully reproduce the coupled feedback, so prescribing observations adds an extra, unquantified bias.","Estimates of Earth's feedback trend from prescribed-SST simulations carry little evidential weight for constraining changes in climate sensitivity over the observational period."],"supporting_citations":[{"why":"Introduced the amip-piForcing method and the moving-window regression approach for estimating the time-varying feedback parameter.","marker":"Gregory and Andrews (2016)"},{"why":"Used amip-piForcing runs to argue that historical SST patterns raise estimates of climate sensitivity, a key claim the paper tests.","marker":"Andrews et al. (2018)"},{"why":"Explicitly compared amip-piForcing feedback evolution with coupled historical simulations and attributed the difference to SST patterns, the main comparison the paper repeats with identical SSTs.","marker":"Dong et al. (2021)"},{"why":"Provides a multi-model assessment of historical SST pattern effects on radiative feedback, representing the standard interpretation the paper challenges.","marker":"Andrews et al. (2022)"},{"why":"Showed theoretically that severing air-sea coupling in atmosphere-only simulations alters variance and surface fluxes, a physical explanation for the remaining amip-couple differences.","marker":"Barsugli and Battisti (1998)"},{"why":"Founds the Yule-Slutzky effect, the statistical mechanism the paper uses to explain spurious multi-decadal variability in moving-window regression slopes.","marker":"Slutzky (1937)"},{"why":"Showed that changing the observational SST and sea-ice dataset changes amip-piForcing feedback estimates, supporting the paper's concern about the method's reliability.","marker":"Modak and Mauritsen (2023)"}],"fun_headline_variants":["Prescribed-SST runs blind to true feedback shifts","Feedback trends under 100 years are pure noise","Atmosphere-only models miss coupled feedback evolution","SST-only forcing can't capture feedback timing","Short-term feedback trends: noise, not real change"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire noise-floor and detection argument rests on defining 'detected feedback trend' as the slope from a 30-year moving-window regression; if a different window length or estimation method is used, the noise characteristics and the roughly 100-year bound could change, and the paper does not test that robustness.","fun_headline_variants_meta":{"raw":{"variants":["Prescribed-SST runs blind to true feedback shifts","Feedback trends under 100 years are pure noise","Atmosphere-only models miss coupled feedback evolution","SST-only forcing can't capture feedback timing","Short-term feedback trends: noise, not real change"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1491,"prompt_tokens":1060,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":676,"tokens_out":431,"duration_ms":20384,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:15.938734+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could compute feedback time series from the 4000-year piControl using a 50-year or 100-year moving window, or a state-space estimator; if trends shorter than 100 years then emerge as distinguishable from the noise floor, the paper's central bound fails for that estimator.","supporting_citations":[],"review_version":1}