Pith. sign in

REVIEW 1 major objections 5 minor 15 references

Applying the ACE2 Emulator to SST Green's Functions for the E3SMv3 Global Atmosphere Model

T0 review · 1 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A machine-learning emulator reproduces a climate model's sea-surface-temperature response maps for top-of-atmosphere radiation about 100 times faster, with the largest bias in the tropical eastern Pacific.

desk verdict A careful, honest controlled comparison of ACE2 vs EAMv3 on GFMIP Green's functions; the qualitative conclusion holds, but the significance claims lean on an under-verified uniform variability assumption. read the letter →

arxiv 2505.08742 v2 pith:MGWBXLZB submitted 2025-05-13 physics.ao-ph

classification physics.ao-ph
keywords ACE2climateemulatorseasurfacetemperatureGreen'sfunctiontop-of-atmosphereradiationpatterneffectEAMv3radiativesensitivity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a fast machine-learning emulator of a global atmosphere model can reproduce the model's own sea-surface temperature (SST) Green's functions—the maps that say how a warm or cool patch of ocean at any location changes global top-of-atmosphere radiation. The emulator, ACE2, is trained on 51 years of historical, SST-forced output from the physics-based EAMv3 model, then run through the standardized suite of 109 warm and 109 cool SST patch experiments, with the same experiments run in EAMv3 as ground truth. The paper finds that the emulator captures the spatial pattern of individual patch responses and most of the qualitative structure of the global radiative-sensitivity map, with statistically significant discrepancies concentrated in the low-latitude oceans, especially the tropical eastern Pacific. Both the emulator and the physics model's Green's functions reconstruct the 1970–2020 historical trend in global annual-mean net TOA radiation. The reason to care is cost: the emulator completes the same experiment suite about 100 times faster in wall-clock time, which would make pattern-effect and feedback diagnostics much cheaper if the remaining biases can be reduced.

What carries the argument

The load-bearing machinery is the combination of the emulator itself and the Green's-function construction. ACE2 is an autoregressive machine-learning climate emulator at 1-degree horizontal resolution and 6-hourly time steps, trained with a two-step loss, a 384-dimensional embedding, and global dry-air and moisture conservation, so it learns the radiative response implicitly rather than by predicting clouds. The Green's-function construction follows the standard protocol: 109 ocean patches, each perturbed by a smooth cosine-squared SST anomaly reaching +2 K or -2 K at its center, with the response defined as the change in global-mean net TOA radiation normalized by the area-averaged SST anomaly, $(dN/dSST)_p = \Delta N_p / \langle \Delta SST_p \rangle$, and with an uncertainty estimate based on the interannual standard deviation $\sigma_N = 0.2$ W/m² and the 10-year patch and 20-year control lengths. The paper also uses the patch basis to reconstruct historical $\Delta N$ by superposing patch responses weighted by the SST anomaly at each patch center, which serves as the consistency check.

What would settle it

Extend the EAMv3 control and patch runs to 40 years for all 109 patch locations and estimate each patch's interannual standard deviation and autocorrelation; if any patch shows variability clearly above 0.2 W/m² or significant correlation between years, the 2-sigma thresholds marking the ACE–EAMv3 differences as significant, especially in the northeast Pacific, would need to be recomputed.

Watch

Extended reading notes

Core claim

The paper's central claim is that an autoregressive machine-learning emulator, ACE2, trained only on a 51-year historical SST-forced simulation of the EAMv3 global atmosphere model, can serve as a surrogate for EAMv3 in the standardized SST-patch experiments used to construct Green's functions. The emulator matches the spatial pattern of top-of-atmosphere radiative response for individual +2 K SST patches, reproduces the qualitative geography of the global mean radiation sensitivity to patch location (area-weighted spatial correlation 0.53 against the reference), and reconstructs the historical global-mean net TOA radiation anomaly over 1970–2020 about as well as the physics-based model's own Green's functions. The match is not exact: for a number of low-latitude patches, especially in the subtropical northeast Pacific, the difference between emulator and model sensitivity exceeds the estimated 2-sigma significance threshold of about 8 W/m²/K. The paper attributes the residual bias to insufficient diversity in the SST anomaly patterns sampled during training, and it reports that the emulator runs the whole suite about 100 times faster in wall-clock time.

Load-bearing premise

The significance claims assume that global-mean net TOA radiation has the same year-to-year variability (about 0.2 W/m²) and no serial correlation for every SST patch and the control run, although that value was verified for only the control and three representative patches.

Editorial extensions

If this is right

  • The full 218-patch SST experiment suite runs in 2.3 wall-clock days on one GPU, versus 331 wall-clock days for EAMv3 on eight CPU nodes, a speedup of roughly 100 times.
  • Because the emulator is cheap, every patch simulation can be extended to 40 years, which reduces the internal-variability noise floor for the sensitivity maps by about a factor of two.
  • The ACE and EAMv3 radiative-sensitivity maps have an area-weighted spatial pattern correlation of 0.53, and ACE is biased relative to EAMv3 at many low-latitude patches, most strongly in the 0–20°N eastern Pacific.
  • Green's functions from both models reproduce the historical global annual-mean net TOA radiation anomaly, with the ACE reconstruction showing reduced interannual variability ($\sigma = 0.32$ W/m²) relative to its 0.50 W/m² target.
  • The paper concludes that the current emulator is not yet a substitute for physics-model Green's functions, framing the result as a feasibility demonstration and a target for further training improvements.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's training-diversity hypothesis is correct, the northeast Pacific bias is a data limitation rather than an architectural one; retraining on a long pre-industrial control run or on a subset of the patch simulations themselves should shrink it, and the emulator's low cost makes that test cheap.
  • Holding the patch protocol fixed and comparing ACE2 trained on EAMv3 with the same emulator trained on reanalysis would separate errors introduced by the reference training data from errors intrinsic to the emulator, something the paper's comparison with earlier reanalysis-trained versions only hints at.
  • The emulator reproduces cloud-driven radiative responses without predicting clouds, suggesting that the statistical relationship between resolved atmospheric state and radiation is learnable directly; if so, the same strategy could be applied to other subgrid processes too expensive to simulate explicitly.
  • Using an ensemble of ACE2 random seeds rather than a single checkpoint for the patch runs could reduce emulator-specific internal variability and sharpen the significance map at negligible additional cost.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper trains the ACE2 machine-learning emulator on a 51-year (1970-2020) AMIP-style EAMv3 simulation and then applies the GFMIP Green's function protocol to compare ACE with the EAMv3 reference model. The authors run 109 warm and 109 cold SST patch simulations plus a control for both ACE and EAMv3, extending the ACE runs to 40 years while using the GFMIP-standard 10-year patches for the main comparisons. They find that the spatial patterns of TOA radiative response to individual SST patches are qualitatively similar between ACE and EAMv3, that the derived global-mean TOA radiative sensitivity maps are broadly consistent but show statistically significant discrepancies for some patches, particularly in the subtropical northeast Pacific, and that Green's function reconstructions of the historical 1970-2020 TOA radiation time series are reasonable for both models. They also report that ACE completes the GFMIP suite roughly 100 times faster in wall-clock time than EAMv3 under the configurations used.

Significance. If the results hold, the paper provides a valuable demonstration that a learned atmospheric emulator can approximate the patterned-SST Green's functions of a full-physics climate model at a fraction of the computational cost, which could make pattern-effect diagnostics accessible for many more experiments and models. The study is carefully designed: the GFMIP protocol is followed closely, the ACE patch simulations are genuine out-of-sample tests, the EAMv3 reference patch simulations provide a non-circular ground truth for the head-to-head comparison, and the authors report uncertainty estimates and explicitly acknowledge the limits of their training data. The open release of training/evaluation code, experiment scripts, and the ACE2-EAMv3 checkpoint is a strength that supports reproducibility. The main caveat is that the statistical significance of the ACE-EAMv3 discrepancies rests on an extrapolated interannual-variability assumption; if that assumption is too optimistic, some headline discrepancies, especially in the northeast Pacific, may not be significant.

major comments (1)
  1. [Section 3.3, Eqs. (3)-(7)] The issue is directly testable with the available ACE data and a small number of targeted EAMv3 runs, so it is fixable within the manuscript's scope.
minor comments (5)
  1. [Abstract and Section 2.1] The 'approximately 100 times faster' claim in the Abstract refers to the wall-clock time for running the GFMIP suite after training, but the training cost (50 epochs at 1.7 hours per epoch on 16 A100 GPUs) is substantial; the text should explicitly state that the training cost is excluded, or provide an end-to-end comparison including training.
  2. [Section 3.3] The description of the 'average' method as 'the average of the one-sided warm and cold-patch results' is inconsistent with the formula in Eq. (5), which corresponds to the warm-minus-cold difference divided by two; please reword to avoid confusion.
  3. [Section 3.4 and Figure 7] The reconstruction comparison uses 40-year ACE patch simulations but 10-year EAMv3 patch simulations; this difference in averaging length should be discussed when interpreting the reported standard deviations and RMSE values, since it affects the noise level in the reconstructed Green's functions.
  4. [Figure 7 caption] The sentence 'The RMSE for Green's functions ACE with respect to EAMv3 historical target is 0.43' is ambiguous; please clarify whether this is an additional cross-target metric and also report the RMSE of the ACE reconstruction against its own target.
  5. [Figures 5 and 6] The text refers to 'darkened regions' while the figure captions refer to 'hatches' to indicate significance; please make the terminology consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central ACE-vs-EAMv3 GFMIP comparison is an out-of-sample benchmark, and the historical reconstruction is explicitly labeled a consistency check.

full rationale

The central comparison is not circular. ACE2 is trained on the EAMv3 AMIP simulation, while the GFMIP patch experiments are held out: the paper states, 'Since ACE2-EAMv3 is trained on the AMIP-style EAMv3 reference simulation, all patch simulations are out-of-sample tests of the emulator.' The EAMv3 reference patch simulations used for the sensitivity maps and for the significance test of ACE-vs-EAMv3 differences were run independently under the GFMIP protocol and were not used to fit or select ACE, so the headline discrepancy claim is a genuine benchmark rather than a fitted input. The historical reconstruction is explicitly introduced as a self-consistency check ('A consistency check that helps confirm their credibility') and is not presented as an independent prediction; it uses the same model's own target, which is a stated limitation rather than a disguised input. The selection of the checkpoint and random seed by inline-inference RMSE on the AMIP period is a model-selection step and could make the emulator's historical biases look somewhat favorable, but it does not enter the GFMIP patch-response equations (Eqs. 2, 3, 5, 7) by construction, so it is not a circular reduction. Likewise, the uniform sigma_N = 0.2 W/m^2 assumption is an empirical extrapolation and a legitimate uncertainty concern, but it is not circular: the significance thresholds are derived from an explicit statistical model applied to an independently estimated variability, not from the discrepancy being tested. Self-citations to ACE2 training and GFMIP protocol are normal methodological references and are not load-bearing substitutes for the independent EAMv3 benchmark.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No new physical entities are introduced and no new constants are fitted to the GFMIP target data. The evaluation relies on the GFMIP linear-superposition protocol, on a uniformity/independence assumption for internal variability in the significance tests, and on the representativeness of a single AMIP training simulation. The emulator itself is a trained artifact, not an ad hoc entity.

free parameters (3)
  • ACE random seed / checkpoint selection = seed with lowest 1970-2020 net TOA radiation RMSE
    Model selection on the AMIP evaluation period; mild selection for the reconstruction consistency check, but out-of-sample GF map comparisons are unaffected.
  • Patch valid-grid-point threshold = 3% of patch grid points
    Patches with fewer valid ice-free ocean points are excluded because the ocean-mean SST perturbation becomes too small for reliable sensitivity (Section 3.2).
  • Interannual standard deviation sigma_N = 0.2 W/m2
    Estimated from control and three 40-year patch runs, then assumed uniform across all 109 patches for significance thresholds (Section 3.3).
assumptions (4)
  • domain assumption GFMIP SST patch response is sufficiently linear that warm and cold patch results can be averaged and superposed (BJ24 Eqns. 2-3, this paper Eqns. 8-9).
    The paper checks linearity approximately for tropical patches and finds it less accurate in the extratropics (Section 3.2); the historical reconstruction would be distorted if nonlinearity is strong.
  • domain assumption TOA global-mean net radiation interannual variability is uncorrelated year-to-year and uniform across patch locations with sigma_N = 0.2 W/m2.
    Stated in Section 3.3 and used to draw significance thresholds in Figures 5 and 6; only verified for the control and three patches.
  • standard math The GFMIP patch grid (centers every 10 degrees latitude and 40 degrees longitude, Eq. 1) forms a basis adequate to reconstruct historical SST anomalies for TOA radiation reconstructions.
    Adopted from BJ24; the finite-volume superposition in Eq. 8 ignores patch-to-patch covariance and is an approximation inherent to the protocol.
  • domain assumption The 51-year EAMv3 AMIP simulation is a representative sample of EAMv3 behavior for the out-of-sample GFMIP SST perturbations.
    The paper's own hypothesis (Sections 3.3, 4) is that this fails for the subtropical northeast Pacific training diversity, which would explain the largest biases.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Applying the ACE2 Emulator to SST Green's Functions for the E3SMv3 Global Atmosphere Model." pith.science (2026). https://pith.science/paper/MGWBXLZB

@misc{pith2026250508742,
  author       = {Pith},
  title        = {Pith review of: Applying the ACE2 Emulator to SST Green's Functions for the E3SMv3 Global Atmosphere Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MGWBXLZB}},
  note         = {Machine review of arXiv:2505.08742}
}
read the original abstract

Green's functions are a useful technique for interpreting atmospheric state responses to changes in the spatial pattern of sea surface temperature (SST). Here we train version 2 of the Ai2 Climate Emulator (ACE2) on reference historical SST simulations of the US Department of Energy's EAMv3 global atmosphere model. We compare how well the SST Green's functions generated by ACE2 match those of EAMv3, following the protocol of the Green's Function Model Intercomparison Project (GFMIP). The spatial patterns of top-of-atmosphere (TOA) radiative response from the individual GFMIP SST patch simulations are similar for ACE and the EAMv3 reference. The derived sensitivity of global net TOA radiation sensitivity to SST patch location is qualitatively similar in ACE as in EAMv3, but there are statistically significant discrepancies for some SST patches, especially over the subtropical northeast Pacific. These discrepancies may reflect insufficient diversity in the SST patterns sampled over the course of the EAMv3 AMIP simulation used for training ACE. Both ACE and EAMv3 Green's functions reconstruct the historical record of the global annual-mean TOA radiative flux from a reference EAMv3 AMIP simulation reasonably well. Notably, under our configuration and compute resources, ACE achieves these results approximately 100 times faster in wall-clock time compared to EAMv3, highlighting its potential as a powerful and efficient tool for tackling other computationally intensive problems in climate science.

Figures

Figures reproduced from arXiv: 2505.08742 by the authors.

Figure 1
Figure 1. shows the time series of global-mean net TOA radiation N (positive down￾ward; the overline denotes a global average), TOA upward longwave (LW) radiation, and TOA upward shortwave (SW) radiation for EAMv3 and for an ACE2-EAMv3 simula￾tion with the chosen seed and checkpoint, initialized from the EAMv3 simulation at the start of 1970. ACE2-EAMv3 captures the global trend in the net TOA radiation, with a mean bias of a… view at source ↗
Figure 2
Figure 2. shows maps of 51-year mean spatially-resolved biases of ACE2-EAMv3 net, LW and SW TOA radiation vs. the EAMv3 reference AMIP simulation. The net radiation biases are everywhere less than 10 W/m2 . They are largest in low latitudes, where they are systematically positive (a ‘dim cloud’ bias) outside of stratocumulus re￾gions, especially over the tropical eastern Pacific and Atlantic Ocean [PITH_FULL_IMAGE:figures/fu… view at source ↗
Figure 3
Figure 3. Time series of annual and global mean TOA radiation for the control simulation in ACE and EAMv3, and TOA radiation change from control for tropical ascent, tropical subsi￾dence, and extratropical SST patches. Dash lines indicate the 40 yr average. EAMv3 for tropical SST patches but is less accurate for extratropical SST patches. In contrast, BJ24’s [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Map of TOA radiative flux changes from control for three patch simulations for ACE (a-c), EAMv3 (d-f), and bias (g-i). Black ellipses indicate half-amplitudes for each of the patch SST perturbation. will propagate to all the different fields and can have an undesirable…
Figure 5
Figure 5. Figure 5: Normalized derivatives of TOA radiation with respect to change in SST using Green’s function method. Warming denotes the estimate using only +2K patches, cooling uses -2K patches, and the average is the average of warming and cooling. Top row shows the estimate from AC…
Figure 6
Figure 6. Figure 6: Difference between ACE and EAMv3 Green’s function estimates of global TOA radiation response with SST, constructed by differencing 10-year warm and cold patch simula￾tions. Hatches indicates SST patches for which this difference is greater than an approximate 95% signi…
Figure 7
Figure 7. Figure 7: Historical reconstruction of TOA radiation estimated from Green’s functions cal￾culated from ACE (40 year patch simulations) and EAMv3 (10 year patch simulations) and multiplied by the sea surface temperature anomalies during 1970-2020 AMIP time period. The actual hist…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 7 canonical work pages

  1. [1]

    \ Rugenstein, M A A

    alessi_surface_2023 APACrefauthors Alessi, M J. \ Rugenstein, M A A. APACrefauthors \ 2023 . Surface Temperature Pattern Scenarios Suggest Higher Warming Rates Than Current Projections Surface Temperature Pattern Scenarios Suggest Higher Warming Rates Than Current Projections . Geophysical Research Letters 50 23 e2023GL105795 . APACrefDOI doi:10.1029/2023...

  2. [2]

    \ Sardeshmukh, P D

    barsugli_global_2002 APACrefauthors Barsugli, J J. \ Sardeshmukh, P D. APACrefauthors \ 2002 . Global Atmospheric Sensitivity to Tropical SST Anomalies throughout the Indo - Pacific Basin Global Atmospheric Sensitivity to Tropical SST Anomalies throughout the Indo - Pacific Basin . Journal of Climate

  3. [3]

    , Rugenstein, M A A

    bloch-johnson_greens_2024 APACrefauthors Bloch-Johnson, J. , Rugenstein, M A A. , Alessi, M J. , Proistosescu, C. , Zhao, M. , Zhang, B. Zhou, C. APACrefauthors \ 2024 . The Green 's Function Model Intercomparison Project ( GFMIP ) Protocol The Green 's Function Model Intercomparison Project ( GFMIP ) Protocol . Journal of Advances in Modeling Earth Syste...

  4. [4]

    APACrefauthors \ 1985

    branstator_analysis_1985 APACrefauthors Branstator, G. APACrefauthors \ 1985 . Analysis of General Circulation Model Sea - Surface Temperature Anomaly Simulations Using a Linear Model . Part I : Forced Solutions Analysis of General Circulation Model Sea - Surface Temperature Anomaly Simulations Using a Linear Model . Part I : Forced Solutions . Journal of...

  5. [5]

    , Watt-Meyer, O

    clark_ace2-som_2024 APACrefauthors Clark, S K. , Watt-Meyer, O. , Kwa, A. , McGibbon, J. , Henn, B. , Perkins, W A. Bretherton, C S. APACrefauthors \ 2024 . ACE2 - SOM : Coupling an ML atmospheric emulator to a slab ocean and learning the sensitivity of climate to changed CO2. ACE2 - SOM : Coupling an ML atmospheric emulator to a slab ocean and learning t...

  6. [6]

    , Proistosescu, C

    dong_attributing_2019 APACrefauthors Dong, Y. , Proistosescu, C. , Armour, K C. \ Battisti, D S. APACrefauthors \ 2019 . Attributing Historical and Future Evolution of Radiative Feedbacks to Regional Warming Patterns using a Green ’s Function Approach : The Preeminence of the Western Pacific Attributing Historical and Future Evolution of Radiative Feedbac...

  7. [7]

    Duncan2024 APACrefauthors Duncan, J P C. , Wu, E. , Golaz, J. , Caldwell, P M. , Watt‐Meyer, O. , Clark, S K. Bretherton, C S. APACrefauthors \ 2024 . Application of the AI2 Climate Emulator to E3SMv2’s Global Atmosphere Model, With a Focus on Precipitation Fidelity Application of the ai2 climate emulator to e3smv2’s global atmosphere model, with a focus ...

  8. [8]

    , Bony, S

    Eyring2016 APACrefauthors Eyring, V. , Bony, S. , Meehl, G A. , Senior, C A. , Stevens, B. , Stouffer, R J. \ Taylor, K E. APACrefauthors \ 2016 . Overview of the Coupled Model Intercomparison Project Phase 6 (CMIP6) experimental design and organization Overview of the coupled model intercomparison project phase 6 (cmip6) experimental design and organizat...

Show all 15 references
  1. [9]

    , Boyle, J S

    gates_overview_1999 APACrefauthors Gates, W L. , Boyle, J S. , Covey, C. , Dease, C G. , Doutriaux, C M. , Drach, R S. Williams, D N. APACrefauthors \ 1999 . An Overview of the Results of the Atmospheric Model Intercomparison Project ( AMIP I ) An Overview of the Results of th...

  2. [10]

    , Rugenstein, M

    loon_reanalysis-based_2025 APACrefauthors Loon, S V. , Rugenstein, M. \ Barnes, E A. APACrefauthors \ 2025 . Reanalysis-based Global Radiative Response to Sea Surface Temperature Patterns : Evaluating the Ai2 Climate Emulator . Reanalysis-based Global Radiative Response to Sea...

  3. [11]

    , Dresdner, G

    WattMeyer2023 APACrefauthors Watt-Meyer, O. , Dresdner, G. , McGibbon, J. , Clark, S K. , Henn, B. , Duncan, J. Bretherton, C S. APACrefauthors \ 2023 . ACE: A fast, skillful learned global atmospheric model for climate prediction. ACE: A fast, skillful learned global atmosphe...

  4. [12]

    , Henn, B

    watt-meyer_ace2_2024 APACrefauthors Watt-Meyer, O. , Henn, B. , McGibbon, J. , Clark, S K. , Kwa, A. , Perkins, W A. Bretherton, C S. APACrefauthors \ 2024 . ACE2 : Accurately learning subseasonal to decadal atmospheric variability and forced responses. ACE2 : Accurately learn...

  5. [13]

    , Terai, C R

    xie2025energy APACrefauthors Xie, S. , Terai, C R. , Wang, H. , Tang, Q. , Fan, J. , Burrows, S M. others APACrefauthors \ 2025 . The Energy Exascale Earth System Model Version 3. Part I: Overview of the Atmospheric Component The energy exascale earth system model version 3. p...

  6. [14]

    , Zhao, M

    zhang_using_2023 APACrefauthors Zhang, B. , Zhao, M. \ Tan, Z. APACrefauthors \ 2023 . Using a Green ’s Function Approach to Diagnose the Pattern Effect in GFDL AM4 and CM4 Using a Green ’s Function Approach to Diagnose the Pattern Effect in GFDL AM4 and CM4 . Journal of Clima...

  7. [15]

    , Zelinka, M D

    zhou_analyzing_2017 APACrefauthors Zhou, C. , Zelinka, M D. \ Klein, S A. APACrefauthors \ 2017 . Analyzing the dependence of global cloud feedback on the spatial pattern of sea surface temperature change with a Green 's function approach Analyzing the dependence of global clo...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.