Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Exploring the Use of Machine Learning Weather Models in Data Assimilation

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read ML weather adjoints are too noisy for data assimilation

desk verdict A useful first look at GraphCast and NeuralGCM adjoints for DA, but the central 'unphysical' verdict is confounded by comparing against a physics-free MPAS-A baseline. read the letter →

arxiv 2411.14677 v1 pith:GVBBRWAS submitted 2024-11-22 physics.ao-ph cs.LG

classification physics.ao-phcs.LG
keywords machinelearningweathermodelsdataassimilation4DVartangentlinearmodeladjointGraphCastNeuralGCMMPAS-A
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether two leading machine-learning weather models, GraphCast and NeuralGCM, can supply the tangent linear and adjoint operators that four-dimensional variational data assimilation requires. The authors build those operators by linearizing the models' learned operations and check them against the adjoint of a conventional dynamical-core model, MPAS-A. They find that, even where the formal tangent-linear and adjoint consistency checks pass, the ML adjoints are physically unreliable: responses persist at the perturbation point, vertical patterns are noisy, and specific-humidity sensitivities are exaggerated. The paper's central claim is that these unphysical properties make the current ML linearized models unsuitable for operational 4DVar and likely harmful for ensemble assimilation as well.

What carries the argument

The machinery is the tangent linear and adjoint pair derived from each model: for GraphCast, the Jacobian of the graph-neural-network update rules and its transpose; for NeuralGCM, the Jacobian of the deep neural network parameterizations combined with the linearized dynamical core, obtained by automatic differentiation. The operators are verified by a finite-difference ratio test and the exact adjoint identity before being compared with the benchmark MPAS-A adjoint on identical initial states and perturbations. The comparison itself is the instrument: physical realism is judged by whether adjoint sensitivity appears upstream in a Rossby-wave pattern, stays smooth in the vertical, and avoids implausibly large amplitudes at the perturbation location.

What would settle it

Run the same single-point jet perturbation with an MPAS-A adjoint that includes tangent-linear and adjoint versions of convection, radiation, and microphysics (or with a full-physics NWP adjoint such as an operational 4DVar system's). If the on-the-spot spike, vertical noise, and humidity magnitude ratios persist in that symmetric comparison, the paper's verdict stands; if they disappear, the unphysical properties are artifacts of comparing full-physics ML adjoints to a resolved-scale-only benchmark.

Watch

Extended reading notes

Core claim

The central claim is that the current tangent linear and adjoint models of GraphCast and NeuralGCM exhibit several unphysical properties that limit their suitability for data assimilation applications. In a series of single-point perturbation experiments on a jet-stream state, GraphCast's tangent linear model keeps a strong response at the original perturbation spot after six hours, a feature absent from MPAS-A; its adjoint shows no upstream temperature sensitivity except at that spot, noisy zonal-wind patterns, and specific-humidity sensitivities roughly an order of magnitude stronger than MPAS-A's. NeuralGCM's adjoint captures some upstream, Rossby-like patterns but adds vertical noise, likely from the linearized neural-network parameterizations. If a single observation were assimilated, the resulting 4DVar increment would be noisy and contain unphysical structures far from the observation, and ensemble methods would inherit distorted error covariances.

Load-bearing premise

The comparison assumes that test cases chosen to minimize the effect of unresolved physics make MPAS-A's missing adjoint of those processes irrelevant; if that assumption fails, the differences blamed on the ML models could be artifacts of the benchmark.

Editorial extensions

If this is right

  • A 4DVar system using GraphCast's adjoint would produce noisy increments with structures located far from the observation, worsening forecast error even if it fits the observed point.
  • NeuralGCM's adjoint, though more physical in upstream propagation, still needs smoothing and refinement before it can act as a 4DVar operator.
  • Ensemble Kalman Filter methods would inherit distorted background error covariances from these ML models' perturbation responses.
  • Both ML models show some capacity for Rossby-wave dynamics, indicating the core linearized dynamics may be salvageable with better treatment of learned physics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the persistent on-the-spot sensitivity may reflect the ML models' reliance on local input features and short-range learned correlations, producing a near-identity sensitivity at the perturbed grid cell rather than dynamically consistent wave propagation.
  • Inference: a direct testable extension would be to train or fine-tune the ML models with a physics-consistency loss penalizing adjoint noise and pointwise sensitivity, then re-run the same jet perturbation experiments.
  • Inference: the same single-point protocol could be applied to newer ML weather models; the paper establishes a cheap diagnostic (one perturbation, six-hour TLM/ADM check) for screening any candidate DA operator.
  • Inference: the exaggerated humidity sensitivities suggest the humidity channels in these ML models may be the first place to look for overfitting or data-imbalance artifacts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper develops tangent-linear (TL) and adjoint (AD) models for two machine-learning weather models, GraphCast and NeuralGCM, via automatic differentiation and linearization of their internal operations. It verifies the GraphCast TLM against finite differences of the nonlinear model and checks adjoint consistency with the inner-product identity, reporting agreement to 15 digits. It then compares six-hour TL propagations and adjoint sensitivities for a zonal-wind perturbation near a jet core against the MPAS-A model's adjoint, which includes only the resolved dynamical core and no adjoints of unresolved physical parameterizations. Based on visual inspection of horizontal and vertical sensitivity patterns, the paper concludes that the GraphCast and NeuralGCM adjoints exhibit unphysical localized sensitivities, noisy vertical structures, and exaggerated specific-humidity magnitudes, and argues that these features would degrade 4DVar and ensemble-based data assimilation.

Significance. The question addressed is timely and important: the viability of ML weather models in variational DA depends directly on the quality of their linearized operators. The paper's verification of the GraphCast TLM and adjoint is a real strength, meeting standard Taylor and inner-product consistency checks. If the central conclusion were established, the paper would provide a concrete and useful caution for DA integration of ML models. However, the conclusion is not yet established: the MPAS-A benchmark lacks adjoints of unresolved physics, the assertion that the chosen test cases minimize that gap is unquantified, the NeuralGCM adjoint is not verified, and the 'unphysical' judgment rests on unquantified, visual pattern comparison. With additional control experiments and quantitative diagnostics, the paper could become a solid assessment of ML-model linearizations for DA.

major comments (4)
  1. [§2.3, §3] The decisive comparison is between ML adjoints that include all trained processes and an MPAS-A adjoint that, as the paper states in §2.3, 'only describes the evolution of resolved scale processes.' The symptoms used to label the ML adjoints unphysical—persistent on-the-spot sensitivity, noisy vertical structures, and specific-humidity amplitudes an order of magnitude larger than MPAS-A—are also the signatures that unresolved moist convection, radiation, and microphysics would imprint on a linearized model. The paper asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes,' but no metric or experiment is provided to support that assertion. The central claim that the ML adjoints are unphysical relative to the real atmosphere therefore rests on an unverified benchmark-comparability assumption. To support the claim, the authors should either include a full-physics adjoint control (or a nonlinear full-physics MPAS-A run with finite differences) or quantitatively demonstrate that the chosen cases are insensitive to unresolved processes.
  2. [§2.2, §3, Fig. 5] Verification of the tangent-linear and adjoint models is shown only for GraphCast: the Taylor test for each variable and the inner-product consistency check are reported in §2.2. No equivalent verification is shown for NeuralGCM, even though Fig. 5 and Section 3 interpret NeuralGCM adjoint sensitivities as physically meaningful with noise. Without a Taylor test or adjoint consistency check for NeuralGCM, the reader cannot distinguish model-intrinsic unphysicality from implementation error. The paper should either provide those verification results or clearly state that the NeuralGCM TL/AD results are provisional pending verification.
  3. [§3, Figs. 2–5] The judgments of physical realism are made entirely by visual comparison of horizontal maps and vertical cross-sections. Terms such as 'slightly noisier,' 'quite noisy,' 'strong on-the-spot feature,' and 'an order of magnitude stronger' are not backed by a quantitative diagnostic. A reproducible assessment would require, for example, pattern correlation coefficients between the ML and MPAS-A sensitivities, scale decomposition or spectral filtering to characterize the noise, and amplitude ratios with stated significance levels. As the paper stands, the claim that the differences indicate unphysical behavior rather than legitimate differences in model formulation is not quantitatively supported.
  4. [§2, §3] The implementation details of the GraphCast and NeuralGCM TL/AD models are insufficient for reproducibility or for judging whether the compared fields are on equal footing. Key items are not specified: the horizontal and vertical resolution of the ERA5 initial state and of each model, whether the adjoint includes the input/output normalization layers of the neural networks, how the adjoint of the multi-step forecast operator is chained, and whether the NeuralGCM adjoint includes the dynamical core plus the neural-network parameterizations in a single consistent linearization. These details should be reported so that the comparison is interpretable and reproducible.
minor comments (5)
  1. [Abstract and Keywords] The keyword 'Data tssimilation' contains a typo; it should read 'Data assimilation.'
  2. [§2.2 and Fig. 1 caption] The test statistic is denoted Ψ(α) in the text but J(α) in the caption; please unify the notation.
  3. [Figs. 3 and 4 captions and §3] The response function is referred to as 'T = 1' in some places and 'Tad = 1' in others; please make the notation consistent.
  4. [§4] The statement that adjoint noise 'would distort error covariances' in EnKF is presented as fact; consider rewording to 'may distort' or adding a supporting demonstration.
  5. [References] The paper does not identify the automatic differentiation framework (e.g., JAX, PyTorch) used to construct the adjoints; adding this would aid reproducibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation found; the physical-realism verdict is benchmark-dependent but not circular.

full rationale

The paper's analysis is empirical and self-contained: the GraphCast and NeuralGCM tangent-linear and adjoint models are constructed by linearization or automatic differentiation, and their correctness is verified against each model's own nonlinear trajectory via the standard Taylor-ratio test (Fig. 1) and an inner-product identity that agrees to 15 digits. These checks do not assume the paper's conclusion. The central claim that the ML adjoints are 'unphysical' is based on a side-by-side comparison with the MPAS-A adjoint, which is an external benchmark rather than a quantity derived from the conclusion. Self-citations (e.g., Tian and Zou 2020 for the MPAS-A adjoint) document the benchmark's provenance and are not load-bearing for the verdict. One genuine limitation, flagged in Section 2.3, is that the MPAS-A adjoint excludes unresolved processes ('the linearized version of MPAS-A only describes the evolution of resolved scale processes'), while the ML adjoints include them, and the paper merely asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes' without demonstrating this. That is a benchmark-comparability confound for the physical-realism claim, not a circularity: the conclusion is not equivalent to an input by construction. Accordingly, no circular step is exhibited and no step is scored above zero; the single point reflects the presence of minor self-citations that do not affect the derivation's independence.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the choice of MPAS-A as the physical benchmark, correctness of the implemented TL/AD for both ML models, and qualitative visual assessment. No fitted free parameters are introduced.

assumptions (3)
  • domain assumption The MPAS-A tangent linear and adjoint models, which exclude unresolved physical processes, provide a valid benchmark for physical realism of model sensitivities.
    Section 2.3 states MPAS-A adjoint only describes resolved scale processes; the comparison treats MPAS-A as the physical reference.
  • ad hoc to paper The tangent linear and adjoint models of NeuralGCM are correctly implemented, despite no verification results being shown for NeuralGCM.
    Verification (Fig. 1 and adjoint test) is presented only for GraphCast; NeuralGCM TL/AD results are compared without demonstrating correctness.
  • domain assumption Visual inspection of sensitivity patterns is sufficient to classify model behavior as physical or unphysical.
    Conclusions rely on qualitative comparisons of horizontal and vertical structures; no quantitative skill scores or significance tests are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring the Use of Machine Learning Weather Models in Data Assimilation." pith.science (2026). https://pith.science/paper/GVBBRWAS

@misc{pith2026241114677,
  author       = {Pith},
  title        = {Pith review of: Exploring the Use of Machine Learning Weather Models in Data Assimilation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GVBBRWAS}},
  note         = {Machine review of arXiv:2411.14677}
}
read the original abstract

The use of machine learning (ML) models in meteorology has attracted significant attention for their potential to improve weather forecasting efficiency and accuracy. GraphCast and NeuralGCM, two promising ML-based weather models, are at the forefront of this innovation. However, their suitability for data assimilation (DA) systems, particularly for four-dimensional variational (4DVar) DA, remains under-explored. This study evaluates the tangent linear (TL) and adjoint (AD) models of both GraphCast and NeuralGCM to assess their viability for integration into a DA framework. We compare the TL/AD results of GraphCast and NeuralGCM with those of the Model for Prediction Across Scales - Atmosphere (MPAS-A), a well-established numerical weather prediction (NWP) model. The comparison focuses on the physical consistency and reliability of TL/AD responses to perturbations. While the adjoint results of both GraphCast and NeuralGCM show some similarity to those of MPAS-A, they also exhibit unphysical noise at various vertical levels, raising concerns about their robustness for operational DA systems. The implications of this study extend beyond 4DVar applications. Unphysical behavior and noise in ML-derived TL/AD models could lead to inaccurate error covariances and unreliable ensemble forecasts, potentially degrading the overall performance of ensemble-based DA systems, as well. Addressing these challenges is critical to ensuring that ML models, such as GraphCast and NeuralGCM, can be effectively integrated into operational DA systems, paving the way for more accurate and efficient weather predictions.

Figures

Figures reproduced from arXiv: 2411.14677 by the authors.

Figure 1
Figure 1. Variations in the function for the correctness check of the GraphCast tangent linear model for a 6-hour [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. (a) Background geopotential heights (contoured) and zonal wind (shaded) at 00 UTC on January 1, 2022. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Horizontal distribution of adjoint sensitivity 6 hours prior to the response function of [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Vertical cross-sections of adjoint sensitivity 6 hours prior to response function of [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: The horizontal distribution (top row) and vertical cross sections (bottom row) of NeuralGCM adjoint sensitivity [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Long-time accuracy of ensemble Kalman filters for chaotic and machine-learned dynamical systems

    math.DS 2024-12 conditional novelty 7.0 of 10

    Under a squeezing condition on the dynamics, a discrete-time square-root ensemble Kalman filter (and its surrogate-model variant) achieves long-time mean state estimation error of order ε, the observation noise level,...

  2. GraphDOP: Towards skilful data-driven medium-range weather forecasts learnt and initialised directly from observations

    physics.ao-ph 2024-12 conditional novelty 6.0 of 10

    A graph-neural-network weather model trained only on raw observations produces skillful global forecasts out to five days, with tropical 2-meter temperature forecasts competitive with the operational IFS.

  3. Jacobian-Enforced Neural Networks (JENN) for Improved Data Assimilation Consistency in Dynamical Models

    cs.LG 2024-12 reject novelty 5.0 of 10

    A two-phase training scheme that adds tangent-linear and adjoint loss terms to a neural network emulator of Lorenz 96 improves Jacobian consistency while preserving forecast accuracy.

Reference graph

Works this paper leans on

20 extracted references · 16 canonical work pages · cited by 3 Pith papers

  1. [1]

    Brotzge, Don Berchoff, DaNa L

    Jerald A. Brotzge, Don Berchoff, DaNa L. Carlis, Frederick H. Carr, Rachel Hogan Carr, Jordan J. Gerth, Brian D. Gross, Thomas M. Hamill, Sue Ellen Haupt, and Neil Jacobs. Challenges and opportunities in numerical weather prediction. Bull. Amer. Meteor. Soc., 104 0 (3): 0 E698--E705, 2023. ISSN 0003-0007

  2. [2]

    Learning skillful medium-range global weather forecasting

    Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, and Weihua Hu. Learning skillful medium-range global weather forecasting. Science, 382 0 (6677): 0 1416--1421, 2023. ISSN 0036-8075

  3. [3]

    Accurate medium-range global weather forecasting with 3d neural networks

    Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619 0 (7970): 0 533--538, 2023. ISSN 0028-0836

  4. [4]

    Neural general circulation models for weather and climate

    Dmitrii Kochkov, Janni Yuval, Ian Langmore, Peter Norgaard, Jamie Smith, Griffin Mooers, Milan Klöwer, James Lottes, Stephan Rasp, and Peter Düben. Neural general circulation models for weather and climate. Nature, pages 1--7, 2024. ISSN 0028-0836

  5. [5]

    Skamarock, Joseph B

    William C. Skamarock, Joseph B. Klemp, Michael G. Duda, Laura D. Fowler, Sang-Hun Park, and Todd D. Ringler. A multiscale nonhydrostatic atmospheric model using centroidal voronoi tesselations and c-grid staggering. Mon. Wea. Rev., 140 0 (9): 0 3090--3105, 2012. ISSN 0027-0644. doi:10.1175/MWR-D-11-00215.1. URL https://doi.org/10.1175/MWR-D-11-00215.1

  6. [6]

    Michaelis, Gary M

    Allison C. Michaelis, Gary M. Lackmann, and Walter A. Robinson. Evaluation of a unique approach to high-resolution climate modeling using the model for prediction across scales–atmosphere (mpas-a) version 5.1. Geosci. Model Dev., 12 0 (8): 0 3725--3743, 2019. ISSN 1991-959X

  7. [7]

    The quiet revolution of numerical weather prediction

    Peter Bauer, Alan Thorpe, and Gilbert Brunet. The quiet revolution of numerical weather prediction. Nature, 525 0 (7567): 0 47--55, 2015. ISSN 0028-0836

  8. [8]

    Markus Reichstein, Gustau Camps-Valls, Bjorn Stevens, Martin Jung, Joachim Denzler, Nuno Carvalhais, and F. Prabhat. Deep learning and process understanding for data-driven earth system science. Nature, 566 0 (7743): 0 195--204, 2019. ISSN 0028-0836

Show all 20 references
  1. [9]

    Järvinen, E

    Florence Rabier, H. Järvinen, E. Klinker, J‐F Mahfouf, and A. Simmons. The ecmwf operational implementation of four‐dimensional variational assimilation. i: Experimental results with simplified physics. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1143--1170, 2000. ISSN 0035-9009

  2. [10]

    The ecmwf operational implementation of four‐dimensional variational assimilation

    J‐F Mahfouf and Florence Rabier. The ecmwf operational implementation of four‐dimensional variational assimilation. ii: Experimental results with improved physics. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1171--1190, 2000. ISSN 0035-9009

  3. [11]

    Rabier, G

    Ernst Klinker, F. Rabier, G. Kelly, and J‐F Mahfouf. The ecmwf operational implementation of four‐dimensional variational assimilation. iii: Experimental results and diagnostics with operational configuration. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1191--1215, 2000. ISSN 0035-9009

  4. [12]

    A strategy for operational implementation of 4d‐var, using an incremental approach

    Philippe Courtier, J‐N Thépaut, and Anthony Hollingsworth. A strategy for operational implementation of 4d‐var, using an incremental approach. Q. J. Royal Meteorol. Soc., 120 0 (519): 0 1367--1387, 1994. ISSN 0035-9009

  5. [13]

    Dueben, Sebastian Scher, Jonathan A

    Stephan Rasp, Peter D. Dueben, Sebastian Scher, Jonathan A. Weyn, Soukayna Mouatadid, and Nils Thuerey. Weatherbench: a benchmark data set for data‐driven weather forecasting. J. Adv. Model Earth Syst., 12 0 (11): 0 e2020MS002203, 2020. ISSN 1942-2466

  6. [14]

    A neural-network based mpas-shallow water model and its 4d-var data assimilation system

    Xiaoxu Tian, Luke Conibear, and Jeffrey Steward. A neural-network based mpas-shallow water model and its 4d-var data assimilation system. Atmosphere, 14 0 (1): 0 157, 2023. ISSN 2073-4433. URL https://www.mdpi.com/2073-4433/14/1/157

  7. [15]

    Watkins, Jordan S

    Arka Daw, Anuj Karpatne, William D. Watkins, Jordan S. Read, and Vipin Kumar. Physics-guided neural networks (pgnn): An application in lake temperature modeling, pages 353--372. Chapman and Hall/CRC, 2022

  8. [16]

    Alan J. Geer. Learning earth system models from observations: machine learning or data assimilation? Philos. Trans. R. Soc. A, 379 0 (2194): 0 20200089, 2021. ISSN 1364-503X

  9. [17]

    Development of the tangent linear and adjoint models of the mpas-atmosphere dynamic core and applications in adjoint relative sensitivity studies

    Xiaoxu Tian and Xiaolei Zou. Development of the tangent linear and adjoint models of the mpas-atmosphere dynamic core and applications in adjoint relative sensitivity studies. Tellus A, 72 0 (1): 0 1--17, 2020. ISSN 1600-0870

  10. [18]

    Validation of a prototype global 4d-var data assimilation system for the mpas-atmosphere model

    Xiaoxu Tian and Xiaolei Zou. Validation of a prototype global 4d-var data assimilation system for the mpas-atmosphere model. Mon. Wea. Rev., 149 0 (8): 0 2803--2817, 2021. ISSN 0027-0644

  11. [19]

    Evolutions of errors in the global multiresolution model for prediction across scales - shallow water (mpas-sw)

    Xiaoxu Tian. Evolutions of errors in the global multiresolution model for prediction across scales - shallow water (mpas-sw). Q. J. Royal Meteorol. Soc., 147 0 (734): 0 382--391, 2020. ISSN 0035-9009. doi:10.1002/qj.3923. URL https://doi.org/10.1002/qj.3923

  12. [20]

    Forecasting global weather with graph neural networks

    Ryan Keisler. Forecasting global weather with graph neural networks. arXiv preprint arXiv:2202.07575, 2022

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.