{"id":"47557bc9-5da5-4b49-a748-f315802fad4d","arxiv_id":"2505.08742","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An ML emulator trained on one climate model reproduces that model's SST Green's function patterns qualitatively, with significant regional biases and about a 100x wall-clock speedup.","lead":"This paper tests whether an AI emulator of a global climate model can reproduce the model's Green's function response to ocean temperature changes. It finds broadly similar responses, some regionally significant mismatches, and about a 100-fold speedup over the full climate model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform σ_N=0.2 W/m² assumption for all patches underlies the claimed significant ACE–EAMv3 discrepancies; NE Pacific patch-specific variability could push differences below 2σ.","rationale":"I agree with the reader's weakest-assumption identification. The central claim is the comparison of ACE and EAMv3 Green's functions, and the only quantitative support for 'statistically significant discrepancies' is the noise-floor calculation in Section 3.3. That calculation assumes uniform, uncorrelated interannual variability across all patches; the authors verified this on only four simulations. If the assumption fails for the NE Pacific patch—the headline discrepancy—the paper's strongest quantitative statement could become unsupported. I considered other possible concerns: checkpoint selection uses the same 1970-2020 period as the reconstruction target, and the 100x speedup depends on whether training time is included. Both are secondary to the central sensitivity-map comparison and do not undermine the qualitative similarity, which is visually and statistically (pattern correlation 0.53) supported. The concern is therefore condition-setting rather than fatal, and the reader's CONDITIONAL verdict already captures it. No verdict change is needed; a patch-specific σ_N computation would settle whether the condition is met.","tokens_in":11170,"tokens_out":7116,"duration_ms":76608,"concrete_test":"Estimate patch-specific σ_N and lag-1 autocorrelation from the 40-year ACE control and all 109 ACE patch simulations, which are already available, and from the 10-year EAMv3 patch simulations using within-patch detrended annual anomalies. Then extend the EAMv3 NE Pacific patch and two other flagged patches to 40 years to reduce sampling error. Recompute Figure 6's hatching using patch-specific σ_N in Eq. (7) instead of 0.2 W/m². If the NE Pacific patch and most hatched patches remain above 2σ_diff, the concern is resolved; if hatching flips or disappears, the claim of statistically significant discrepancies must be re-scoped.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative claim of statistically significant ACE–EAMv3 differences, especially the NE Pacific bias highlighted in the abstract and Figure 6, rests entirely on Section 3.3's use of a single interannual standard deviation σ_N = 0.2 W/m² for all 109 patches and the control, with no serial correlation. This was verified only for the control and three representative patches (tropical ascent, tropical subsidence, extratropical). The NE Pacific is a stratocumulus-rich region where low-cloud internal variability is plausibly larger than the tropical-mean value. If, for example, σ_N = 0.3 W/m² for the NE Pacific patch, the noise floor in Eq. (7) rises to roughly 6 W/m²/K, and the reported difference of about 8 W/m²/K would no longer exceed the 2σ threshold used to define significance. Similarly, positive lag-1 autocorrelation would make 10-year mean uncertainties larger than σ_N/√10. The authors themselves state they 'regard this interannual variability of N as representative of all ocean patch locations,' but that extrapolation is the load-bearing assumption for the discrepancy claim. This is not an internal inconsistency; it is an under-verified empirical assumption, and it is directly testable with the existing 40-year ACE runs plus targeted EAMv3 extensions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains the ACE2 machine-learning emulator on a 51-year (1970-2020) AMIP-style EAMv3 simulation and then applies the GFMIP Green's function protocol to compare ACE with the EAMv3 reference model. The authors run 109 warm and 109 cold SST patch simulations plus a control for both ACE and EAMv3, extending the ACE runs to 40 years while using the GFMIP-standard 10-year patches for the main comparisons. They find that the spatial patterns of TOA radiative response to individual SST patches are qualitatively similar between ACE and EAMv3, that the derived global-mean TOA radiative sensitivity maps are broadly consistent but show statistically significant discrepancies for some patches, particularly in the subtropical northeast Pacific, and that Green's function reconstructions of the historical 1970-2020 TOA radiation time series are reasonable for both models. They also report that ACE completes the GFMIP suite roughly 100 times faster in wall-clock time than EAMv3 under the configurations used.","tokens_in":11373,"tokens_out":5330,"duration_ms":53149,"significance":"If the results hold, the paper provides a valuable demonstration that a learned atmospheric emulator can approximate the patterned-SST Green's functions of a full-physics climate model at a fraction of the computational cost, which could make pattern-effect diagnostics accessible for many more experiments and models. The study is carefully designed: the GFMIP protocol is followed closely, the ACE patch simulations are genuine out-of-sample tests, the EAMv3 reference patch simulations provide a non-circular ground truth for the head-to-head comparison, and the authors report uncertainty estimates and explicitly acknowledge the limits of their training data. The open release of training/evaluation code, experiment scripts, and the ACE2-EAMv3 checkpoint is a strength that supports reproducibility. The main caveat is that the statistical significance of the ACE-EAMv3 discrepancies rests on an extrapolated interannual-variability assumption; if that assumption is too optimistic, some headline discrepancies, especially in the northeast Pacific, may not be significant.","major_comments":[{"comment":"The issue is directly testable with the available ACE data and a small number of targeted EAMv3 runs, so it is fixable within the manuscript's scope.","section":"Section 3.3, Eqs. (3)-(7)"}],"minor_comments":[{"comment":"The 'approximately 100 times faster' claim in the Abstract refers to the wall-clock time for running the GFMIP suite after training, but the training cost (50 epochs at 1.7 hours per epoch on 16 A100 GPUs) is substantial; the text should explicitly state that the training cost is excluded, or provide an end-to-end comparison including training.","section":"Abstract and Section 2.1"},{"comment":"The description of the 'average' method as 'the average of the one-sided warm and cold-patch results' is inconsistent with the formula in Eq. (5), which corresponds to the warm-minus-cold difference divided by two; please reword to avoid confusion.","section":"Section 3.3"},{"comment":"The reconstruction comparison uses 40-year ACE patch simulations but 10-year EAMv3 patch simulations; this difference in averaging length should be discussed when interpreting the reported standard deviations and RMSE values, since it affects the noise level in the reconstructed Green's functions.","section":"Section 3.4 and Figure 7"},{"comment":"The sentence 'The RMSE for Green's functions ACE with respect to EAMv3 historical target is 0.43' is ambiguous; please clarify whether this is an additional cross-target metric and also report the RMSE of the ACE reconstruction against its own target.","section":"Figure 7 caption"},{"comment":"The text refers to 'darkened regions' while the figure captions refer to 'hatches' to indicate significance; please make the terminology consistent.","section":"Figures 5 and 6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is the first controlled test of the ACE2 emulator against the same underlying model's own GFMIP simulations, and that comparison was missing from Loon et al. (2025). The paper does the protocol properly, reports disagreement as clearly as agreement, and concludes, rightly, that ACE is not yet a substitute for physics-based Green's functions. That alone makes it a useful step forward.\n\nWhat's done well: same patch protocol for both models, explicit noise-floor estimates, warm/cold differencing rather than one-sided maps, and an honest discussion of where ACE disagrees with EAMv3. The historical reconstruction is a nice self-consistency check, and the paper is appropriately scoped—no overclaiming that the emulator is ready for prime time. The 100x speedup claim is concrete and tied to their actual configuration.\n\nWhere the soft spots are, in proportion: the main one is the assumption that the interannual standard deviation of global-mean TOA net radiation is σ_N = 0.2 W/m² for all 109 patches and the control, verified on only the control and three representative patches. The stress-test arithmetic is correct: if the subtropical northeast Pacific patch—a stratocumulus region where internal variability is plausibly larger—has σ_N = 0.3 W/m², the 2σ difference-map noise floor rises from roughly 8 to about 12 W/m²/K, and some of the claimed 'significant' discrepancies would no longer clear that bar. That does not overturn the qualitative conclusion, but it does make the specific claim about which patches are biased shakier than the prose suggests. The authors flag this assumption themselves, and it is directly testable with their existing 40-year ACE runs plus targeted EAMv3 extensions; a referee should ask for that sensitivity check. Minor issues: checkpoint selection uses the same period later used as the reconstruction target, which is a mild self-consistency issue rather than circularity, and the 'remarkably well' language in Section 3.4 softens a documented under-response. Full reference data are not shipped, though code and checkpoints are.\n\nWho this is for: anyone working on ML emulators for climate, or on pattern-effect methodology, gets real value here. The central comparison is well designed and the paper deserves a serious referee. I'd send it out, with a request to address the σ_N uniformity assumption explicitly, either by bounding it or by softening the significance claims accordingly.","headline":"A careful, honest controlled comparison of ACE2 vs EAMv3 on GFMIP Green's functions; the qualitative conclusion holds, but the significance claims lean on an under-verified uniform variability assumption.","tokens_in":11985,"tokens_out":2818,"would_cite":true,"duration_ms":28877,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learning emulator reproduces a climate model's sea-surface-temperature response maps for top-of-atmosphere radiation about 100 times faster, with the largest bias in the tropical eastern Pacific.","keywords":["ACE2","climate emulator","sea surface temperature","Green's function","top-of-atmosphere radiation","pattern effect","EAMv3","radiative sensitivity"],"falsifier":"Extend the EAMv3 control and patch runs to 40 years for all 109 patch locations and estimate each patch's interannual standard deviation and autocorrelation; if any patch shows variability clearly above 0.2 W/m² or significant correlation between years, the 2-sigma thresholds marking the ACE–EAMv3 differences as significant, especially in the northeast Pacific, would need to be recomputed.","tokens_in":10927,"feed_emoji":"⚡","tokens_out":18839,"duration_ms":157321,"temperature":0.7,"pith_summary":"This paper asks whether a fast machine-learning emulator of a global atmosphere model can reproduce the model's own sea-surface temperature (SST) Green's functions—the maps that say how a warm or cool patch of ocean at any location changes global top-of-atmosphere radiation. The emulator, ACE2, is trained on 51 years of historical, SST-forced output from the physics-based EAMv3 model, then run through the standardized suite of 109 warm and 109 cool SST patch experiments, with the same experiments run in EAMv3 as ground truth. The paper finds that the emulator captures the spatial pattern of individual patch responses and most of the qualitative structure of the global radiative-sensitivity map, with statistically significant discrepancies concentrated in the low-latitude oceans, especially the tropical eastern Pacific. Both the emulator and the physics model's Green's functions reconstruct the 1970–2020 historical trend in global annual-mean net TOA radiation. The reason to care is cost: the emulator completes the same experiment suite about 100 times faster in wall-clock time, which would make pattern-effect and feedback diagnostics much cheaper if the remaining biases can be reduced.","feed_headline":"ML emulator reproduces climate model's SST response 100x faster","feed_subtitle":"It reproduces top-of-atmosphere radiation response to SST patches, with the largest bias in the tropical eastern Pacific.","key_machinery":"The load-bearing machinery is the combination of the emulator itself and the Green's-function construction. ACE2 is an autoregressive machine-learning climate emulator at 1-degree horizontal resolution and 6-hourly time steps, trained with a two-step loss, a 384-dimensional embedding, and global dry-air and moisture conservation, so it learns the radiative response implicitly rather than by predicting clouds. The Green's-function construction follows the standard protocol: 109 ocean patches, each perturbed by a smooth cosine-squared SST anomaly reaching +2 K or -2 K at its center, with the response defined as the change in global-mean net TOA radiation normalized by the area-averaged SST anomaly, $(dN/dSST)_p = \\Delta N_p / \\langle \\Delta SST_p \\rangle$, and with an uncertainty estimate based on the interannual standard deviation $\\sigma_N = 0.2$ W/m² and the 10-year patch and 20-year control lengths. The paper also uses the patch basis to reconstruct historical $\\Delta N$ by superposing patch responses weighted by the SST anomaly at each patch center, which serves as the consistency check.","core_discovery":"The paper's central claim is that an autoregressive machine-learning emulator, ACE2, trained only on a 51-year historical SST-forced simulation of the EAMv3 global atmosphere model, can serve as a surrogate for EAMv3 in the standardized SST-patch experiments used to construct Green's functions. The emulator matches the spatial pattern of top-of-atmosphere radiative response for individual +2 K SST patches, reproduces the qualitative geography of the global mean radiation sensitivity to patch location (area-weighted spatial correlation 0.53 against the reference), and reconstructs the historical global-mean net TOA radiation anomaly over 1970–2020 about as well as the physics-based model's own Green's functions. The match is not exact: for a number of low-latitude patches, especially in the subtropical northeast Pacific, the difference between emulator and model sensitivity exceeds the estimated 2-sigma significance threshold of about 8 W/m²/K. The paper attributes the residual bias to insufficient diversity in the SST anomaly patterns sampled during training, and it reports that the emulator runs the whole suite about 100 times faster in wall-clock time.","pith_inferences":["If the paper's training-diversity hypothesis is correct, the northeast Pacific bias is a data limitation rather than an architectural one; retraining on a long pre-industrial control run or on a subset of the patch simulations themselves should shrink it, and the emulator's low cost makes that test cheap.","Holding the patch protocol fixed and comparing ACE2 trained on EAMv3 with the same emulator trained on reanalysis would separate errors introduced by the reference training data from errors intrinsic to the emulator, something the paper's comparison with earlier reanalysis-trained versions only hints at.","The emulator reproduces cloud-driven radiative responses without predicting clouds, suggesting that the statistical relationship between resolved atmospheric state and radiation is learnable directly; if so, the same strategy could be applied to other subgrid processes too expensive to simulate explicitly.","Using an ensemble of ACE2 random seeds rather than a single checkpoint for the patch runs could reduce emulator-specific internal variability and sharpen the significance map at negligible additional cost."],"forward_implications":["The full 218-patch SST experiment suite runs in 2.3 wall-clock days on one GPU, versus 331 wall-clock days for EAMv3 on eight CPU nodes, a speedup of roughly 100 times.","Because the emulator is cheap, every patch simulation can be extended to 40 years, which reduces the internal-variability noise floor for the sensitivity maps by about a factor of two.","The ACE and EAMv3 radiative-sensitivity maps have an area-weighted spatial pattern correlation of 0.53, and ACE is biased relative to EAMv3 at many low-latitude patches, most strongly in the 0–20°N eastern Pacific.","Green's functions from both models reproduce the historical global annual-mean net TOA radiation anomaly, with the ACE reconstruction showing reduced interannual variability ($\\sigma = 0.32$ W/m²) relative to its 0.50 W/m² target.","The paper concludes that the current emulator is not yet a substitute for physics-model Green's functions, framing the result as a feasibility demonstration and a target for further training improvements."],"supporting_citations":[{"why":"It defines the standardized SST-patch protocol: the cosine-squared patch shape, the 109 warm and 109 cool perturbation set, the sensitivity and uncertainty formulas, and the historical reconstruction check that the paper applies to both models.","marker":"Bloch-Johnson et al. (2024)"},{"why":"It introduces the autoregressive machine-learning emulator architecture that ACE2 builds on.","marker":"Watt-Meyer et al. (2023)"},{"why":"It specifies the ACE2 training protocol, including the two-step-ahead loss, the inline-inference checkpoint selection, and the conservation upgrades used in this paper.","marker":"Watt-Meyer et al. (2024)"},{"why":"It contributes ACE2 training-protocol refinements that the paper cites for the emulator's configuration.","marker":"Clark et al. (2024)"},{"why":"It documents the earlier application of ACE to an E3SM atmosphere model, which the paper updates by training on EAMv3 instead.","marker":"Duncan et al. (2024)"},{"why":"It reports prior SST-patch experiments with earlier ACE versions without a matching physics-model ground truth, giving the comparison baseline this paper improves on.","marker":"Loon et al. (2025)"},{"why":"It establishes the Green's function approach for diagnosing how SST patterns affect cloud feedback and provides the physical interpretation used for the response maps.","marker":"Zhou et al. (2017)"},{"why":"It documents the EAMv3 atmospheric model that supplies both the training data and the reference patch simulations.","marker":"Xie et al. (2025)"}],"fun_headline_variants":["ACE2 emulator reproduces EAMv3 SST response, runs 100x faster","ML emulator reproduces climate model SST response at 100x speed","100x faster ML emulator nails SST Green's functions for EAMv3","ACE2 replicates EAMv3's SST response, 100x faster"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The significance claims assume that global-mean net TOA radiation has the same year-to-year variability (about 0.2 W/m²) and no serial correlation for every SST patch and the control run, although that value was verified for only the control and three representative patches.","fun_headline_variants_meta":{"raw":{"variants":["ACE2 emulator reproduces EAMv3 SST response, runs 100x faster","ML emulator reproduces climate model SST response at 100x speed","100x faster ML emulator nails SST Green's functions for EAMv3","ACE2 replicates EAMv3's SST response, 100x faster"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00118,"raw_usage":{"total_tokens":4926,"prompt_tokens":1050,"completion_tokens":3876,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":3790}},"tokens_in":666,"tokens_out":3876,"duration_ms":26089,"temperature":1.0,"reasoning_tokens":3790,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:47:27.583589+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Extend the EAMv3 control and patch runs to 40 years for all 109 patch locations and estimate each patch's interannual standard deviation and autocorrelation; if any patch shows variability clearly above 0.2 W/m² or significant correlation between years, the 2-sigma thresholds marking the ACE–EAMv3 differences as significant, especially in the northeast Pacific, would need to be recomputed.","supporting_citations":[{"cited_title":"Reanalysis-based Global Radiative Response to Sea Surface Temperature Patterns: Evaluating the Ai2 Climate Emulator","cited_arxiv_id":"2502.10893","evidence_quote":"It reports prior SST-patch experiments with earlier ACE versions without a matching physics-model ground truth, giving the comparison baseline this paper improves on."},{"cited_title":", Terai, C R","cited_arxiv_id":null,"evidence_quote":"It documents the EAMv3 atmospheric model that supplies both the training data and the reference patch simulations."}],"review_version":1}