REVIEW 4 major objections 5 minor 3 cited by
Exploring the Use of Machine Learning Weather Models in Data Assimilation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read ML weather adjoints are too noisy for data assimilation
desk verdict A useful first look at GraphCast and NeuralGCM adjoints for DA, but the central 'unphysical' verdict is confounded by comparing against a physics-free MPAS-A baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the tangent linear and adjoint pair derived from each model: for GraphCast, the Jacobian of the graph-neural-network update rules and its transpose; for NeuralGCM, the Jacobian of the deep neural network parameterizations combined with the linearized dynamical core, obtained by automatic differentiation. The operators are verified by a finite-difference ratio test and the exact adjoint identity before being compared with the benchmark MPAS-A adjoint on identical initial states and perturbations. The comparison itself is the instrument: physical realism is judged by whether adjoint sensitivity appears upstream in a Rossby-wave pattern, stays smooth in the vertical, and avoids implausibly large amplitudes at the perturbation location.
What would settle it
Run the same single-point jet perturbation with an MPAS-A adjoint that includes tangent-linear and adjoint versions of convection, radiation, and microphysics (or with a full-physics NWP adjoint such as an operational 4DVar system's). If the on-the-spot spike, vertical noise, and humidity magnitude ratios persist in that symmetric comparison, the paper's verdict stands; if they disappear, the unphysical properties are artifacts of comparing full-physics ML adjoints to a resolved-scale-only benchmark.
Extended reading notes
Core claim
The central claim is that the current tangent linear and adjoint models of GraphCast and NeuralGCM exhibit several unphysical properties that limit their suitability for data assimilation applications. In a series of single-point perturbation experiments on a jet-stream state, GraphCast's tangent linear model keeps a strong response at the original perturbation spot after six hours, a feature absent from MPAS-A; its adjoint shows no upstream temperature sensitivity except at that spot, noisy zonal-wind patterns, and specific-humidity sensitivities roughly an order of magnitude stronger than MPAS-A's. NeuralGCM's adjoint captures some upstream, Rossby-like patterns but adds vertical noise, likely from the linearized neural-network parameterizations. If a single observation were assimilated, the resulting 4DVar increment would be noisy and contain unphysical structures far from the observation, and ensemble methods would inherit distorted error covariances.
Load-bearing premise
The comparison assumes that test cases chosen to minimize the effect of unresolved physics make MPAS-A's missing adjoint of those processes irrelevant; if that assumption fails, the differences blamed on the ML models could be artifacts of the benchmark.
Editorial extensions
If this is right
- A 4DVar system using GraphCast's adjoint would produce noisy increments with structures located far from the observation, worsening forecast error even if it fits the observed point.
- NeuralGCM's adjoint, though more physical in upstream propagation, still needs smoothing and refinement before it can act as a 4DVar operator.
- Ensemble Kalman Filter methods would inherit distorted background error covariances from these ML models' perturbation responses.
- Both ML models show some capacity for Rossby-wave dynamics, indicating the core linearized dynamics may be salvageable with better treatment of learned physics.
Reading between the lines
- Inference: the persistent on-the-spot sensitivity may reflect the ML models' reliance on local input features and short-range learned correlations, producing a near-identity sensitivity at the perturbed grid cell rather than dynamically consistent wave propagation.
- Inference: a direct testable extension would be to train or fine-tune the ML models with a physics-consistency loss penalizing adjoint noise and pointwise sensitivity, then re-run the same jet perturbation experiments.
- Inference: the same single-point protocol could be applied to newer ML weather models; the paper establishes a cheap diagnostic (one perturbation, six-hour TLM/ADM check) for screening any candidate DA operator.
- Inference: the exaggerated humidity sensitivities suggest the humidity channels in these ML models may be the first place to look for overfitting or data-imbalance artifacts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops tangent-linear (TL) and adjoint (AD) models for two machine-learning weather models, GraphCast and NeuralGCM, via automatic differentiation and linearization of their internal operations. It verifies the GraphCast TLM against finite differences of the nonlinear model and checks adjoint consistency with the inner-product identity, reporting agreement to 15 digits. It then compares six-hour TL propagations and adjoint sensitivities for a zonal-wind perturbation near a jet core against the MPAS-A model's adjoint, which includes only the resolved dynamical core and no adjoints of unresolved physical parameterizations. Based on visual inspection of horizontal and vertical sensitivity patterns, the paper concludes that the GraphCast and NeuralGCM adjoints exhibit unphysical localized sensitivities, noisy vertical structures, and exaggerated specific-humidity magnitudes, and argues that these features would degrade 4DVar and ensemble-based data assimilation.
Significance. The question addressed is timely and important: the viability of ML weather models in variational DA depends directly on the quality of their linearized operators. The paper's verification of the GraphCast TLM and adjoint is a real strength, meeting standard Taylor and inner-product consistency checks. If the central conclusion were established, the paper would provide a concrete and useful caution for DA integration of ML models. However, the conclusion is not yet established: the MPAS-A benchmark lacks adjoints of unresolved physics, the assertion that the chosen test cases minimize that gap is unquantified, the NeuralGCM adjoint is not verified, and the 'unphysical' judgment rests on unquantified, visual pattern comparison. With additional control experiments and quantitative diagnostics, the paper could become a solid assessment of ML-model linearizations for DA.
major comments (4)
- [§2.3, §3] The decisive comparison is between ML adjoints that include all trained processes and an MPAS-A adjoint that, as the paper states in §2.3, 'only describes the evolution of resolved scale processes.' The symptoms used to label the ML adjoints unphysical—persistent on-the-spot sensitivity, noisy vertical structures, and specific-humidity amplitudes an order of magnitude larger than MPAS-A—are also the signatures that unresolved moist convection, radiation, and microphysics would imprint on a linearized model. The paper asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes,' but no metric or experiment is provided to support that assertion. The central claim that the ML adjoints are unphysical relative to the real atmosphere therefore rests on an unverified benchmark-comparability assumption. To support the claim, the authors should either include a full-physics adjoint control (or a nonlinear full-physics MPAS-A run with finite differences) or quantitatively demonstrate that the chosen cases are insensitive to unresolved processes.
- [§2.2, §3, Fig. 5] Verification of the tangent-linear and adjoint models is shown only for GraphCast: the Taylor test for each variable and the inner-product consistency check are reported in §2.2. No equivalent verification is shown for NeuralGCM, even though Fig. 5 and Section 3 interpret NeuralGCM adjoint sensitivities as physically meaningful with noise. Without a Taylor test or adjoint consistency check for NeuralGCM, the reader cannot distinguish model-intrinsic unphysicality from implementation error. The paper should either provide those verification results or clearly state that the NeuralGCM TL/AD results are provisional pending verification.
- [§3, Figs. 2–5] The judgments of physical realism are made entirely by visual comparison of horizontal maps and vertical cross-sections. Terms such as 'slightly noisier,' 'quite noisy,' 'strong on-the-spot feature,' and 'an order of magnitude stronger' are not backed by a quantitative diagnostic. A reproducible assessment would require, for example, pattern correlation coefficients between the ML and MPAS-A sensitivities, scale decomposition or spectral filtering to characterize the noise, and amplitude ratios with stated significance levels. As the paper stands, the claim that the differences indicate unphysical behavior rather than legitimate differences in model formulation is not quantitatively supported.
- [§2, §3] The implementation details of the GraphCast and NeuralGCM TL/AD models are insufficient for reproducibility or for judging whether the compared fields are on equal footing. Key items are not specified: the horizontal and vertical resolution of the ERA5 initial state and of each model, whether the adjoint includes the input/output normalization layers of the neural networks, how the adjoint of the multi-step forecast operator is chained, and whether the NeuralGCM adjoint includes the dynamical core plus the neural-network parameterizations in a single consistent linearization. These details should be reported so that the comparison is interpretable and reproducible.
minor comments (5)
- [Abstract and Keywords] The keyword 'Data tssimilation' contains a typo; it should read 'Data assimilation.'
- [§2.2 and Fig. 1 caption] The test statistic is denoted Ψ(α) in the text but J(α) in the caption; please unify the notation.
- [Figs. 3 and 4 captions and §3] The response function is referred to as 'T = 1' in some places and 'Tad = 1' in others; please make the notation consistent.
- [§4] The statement that adjoint noise 'would distort error covariances' in EnKF is presented as fact; consider rewording to 'may distort' or adding a supporting demonstration.
- [References] The paper does not identify the automatic differentiation framework (e.g., JAX, PyTorch) used to construct the adjoints; adding this would aid reproducibility.
Circularity Check
No circular derivation found; the physical-realism verdict is benchmark-dependent but not circular.
full rationale
The paper's analysis is empirical and self-contained: the GraphCast and NeuralGCM tangent-linear and adjoint models are constructed by linearization or automatic differentiation, and their correctness is verified against each model's own nonlinear trajectory via the standard Taylor-ratio test (Fig. 1) and an inner-product identity that agrees to 15 digits. These checks do not assume the paper's conclusion. The central claim that the ML adjoints are 'unphysical' is based on a side-by-side comparison with the MPAS-A adjoint, which is an external benchmark rather than a quantity derived from the conclusion. Self-citations (e.g., Tian and Zou 2020 for the MPAS-A adjoint) document the benchmark's provenance and are not load-bearing for the verdict. One genuine limitation, flagged in Section 2.3, is that the MPAS-A adjoint excludes unresolved processes ('the linearized version of MPAS-A only describes the evolution of resolved scale processes'), while the ML adjoints include them, and the paper merely asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes' without demonstrating this. That is a benchmark-comparability confound for the physical-realism claim, not a circularity: the conclusion is not equivalent to an input by construction. Accordingly, no circular step is exhibited and no step is scored above zero; the single point reflects the presence of minor self-citations that do not affect the derivation's independence.
Assumptions & free parameters
assumptions (3)
- domain assumption The MPAS-A tangent linear and adjoint models, which exclude unresolved physical processes, provide a valid benchmark for physical realism of model sensitivities.
- ad hoc to paper The tangent linear and adjoint models of NeuralGCM are correctly implemented, despite no verification results being shown for NeuralGCM.
- domain assumption Visual inspection of sensitivity patterns is sufficient to classify model behavior as physical or unphysical.
Cite this review
Pith. "Pith review of Exploring the Use of Machine Learning Weather Models in Data Assimilation." pith.science (2026). https://pith.science/paper/GVBBRWAS
@misc{pith2026241114677,
author = {Pith},
title = {Pith review of: Exploring the Use of Machine Learning Weather Models in Data Assimilation},
year = {2026},
howpublished = {\url{https://pith.science/paper/GVBBRWAS}},
note = {Machine review of arXiv:2411.14677}
}
read the original abstract
The use of machine learning (ML) models in meteorology has attracted significant attention for their potential to improve weather forecasting efficiency and accuracy. GraphCast and NeuralGCM, two promising ML-based weather models, are at the forefront of this innovation. However, their suitability for data assimilation (DA) systems, particularly for four-dimensional variational (4DVar) DA, remains under-explored. This study evaluates the tangent linear (TL) and adjoint (AD) models of both GraphCast and NeuralGCM to assess their viability for integration into a DA framework. We compare the TL/AD results of GraphCast and NeuralGCM with those of the Model for Prediction Across Scales - Atmosphere (MPAS-A), a well-established numerical weather prediction (NWP) model. The comparison focuses on the physical consistency and reliability of TL/AD responses to perturbations. While the adjoint results of both GraphCast and NeuralGCM show some similarity to those of MPAS-A, they also exhibit unphysical noise at various vertical levels, raising concerns about their robustness for operational DA systems. The implications of this study extend beyond 4DVar applications. Unphysical behavior and noise in ML-derived TL/AD models could lead to inaccurate error covariances and unreliable ensemble forecasts, potentially degrading the overall performance of ensemble-based DA systems, as well. Addressing these challenges is critical to ensuring that ML models, such as GraphCast and NeuralGCM, can be effectively integrated into operational DA systems, paving the way for more accurate and efficient weather predictions.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 3 Pith papers
-
Long-time accuracy of ensemble Kalman filters for chaotic and machine-learned dynamical systems
Under a squeezing condition on the dynamics, a discrete-time square-root ensemble Kalman filter (and its surrogate-model variant) achieves long-time mean state estimation error of order ε, the observation noise level,...
-
GraphDOP: Towards skilful data-driven medium-range weather forecasts learnt and initialised directly from observations
A graph-neural-network weather model trained only on raw observations produces skillful global forecasts out to five days, with tropical 2-meter temperature forecasts competitive with the operational IFS.
-
Jacobian-Enforced Neural Networks (JENN) for Improved Data Assimilation Consistency in Dynamical Models
A two-phase training scheme that adds tangent-linear and adjoint loss terms to a neural network emulator of Lorenz 96 improves Jacobian consistency while preserving forecast accuracy.
Reference graph
Works this paper leans on
-
[1]
Jerald A. Brotzge, Don Berchoff, DaNa L. Carlis, Frederick H. Carr, Rachel Hogan Carr, Jordan J. Gerth, Brian D. Gross, Thomas M. Hamill, Sue Ellen Haupt, and Neil Jacobs. Challenges and opportunities in numerical weather prediction. Bull. Amer. Meteor. Soc., 104 0 (3): 0 E698--E705, 2023. ISSN 0003-0007
work page 2023
-
[2]
Learning skillful medium-range global weather forecasting
Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, and Weihua Hu. Learning skillful medium-range global weather forecasting. Science, 382 0 (6677): 0 1416--1421, 2023. ISSN 0036-8075
work page 2023
-
[3]
Accurate medium-range global weather forecasting with 3d neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619 0 (7970): 0 533--538, 2023. ISSN 0028-0836
work page 2023
-
[4]
Neural general circulation models for weather and climate
Dmitrii Kochkov, Janni Yuval, Ian Langmore, Peter Norgaard, Jamie Smith, Griffin Mooers, Milan Klöwer, James Lottes, Stephan Rasp, and Peter Düben. Neural general circulation models for weather and climate. Nature, pages 1--7, 2024. ISSN 0028-0836
2024
-
[5]
William C. Skamarock, Joseph B. Klemp, Michael G. Duda, Laura D. Fowler, Sang-Hun Park, and Todd D. Ringler. A multiscale nonhydrostatic atmospheric model using centroidal voronoi tesselations and c-grid staggering. Mon. Wea. Rev., 140 0 (9): 0 3090--3105, 2012. ISSN 0027-0644. doi:10.1175/MWR-D-11-00215.1. URL https://doi.org/10.1175/MWR-D-11-00215.1
-
[6]
Allison C. Michaelis, Gary M. Lackmann, and Walter A. Robinson. Evaluation of a unique approach to high-resolution climate modeling using the model for prediction across scales–atmosphere (mpas-a) version 5.1. Geosci. Model Dev., 12 0 (8): 0 3725--3743, 2019. ISSN 1991-959X
work page 2019
-
[7]
The quiet revolution of numerical weather prediction
Peter Bauer, Alan Thorpe, and Gilbert Brunet. The quiet revolution of numerical weather prediction. Nature, 525 0 (7567): 0 47--55, 2015. ISSN 0028-0836
work page 2015
-
[8]
Markus Reichstein, Gustau Camps-Valls, Bjorn Stevens, Martin Jung, Joachim Denzler, Nuno Carvalhais, and F. Prabhat. Deep learning and process understanding for data-driven earth system science. Nature, 566 0 (7743): 0 195--204, 2019. ISSN 0028-0836
work page 2019
Show all 20 references
-
[9]
Järvinen, E
Florence Rabier, H. Järvinen, E. Klinker, J‐F Mahfouf, and A. Simmons. The ecmwf operational implementation of four‐dimensional variational assimilation. i: Experimental results with simplified physics. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1143--1170, 2000. ISSN 0035-9009
2000
-
[10]
The ecmwf operational implementation of four‐dimensional variational assimilation
J‐F Mahfouf and Florence Rabier. The ecmwf operational implementation of four‐dimensional variational assimilation. ii: Experimental results with improved physics. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1171--1190, 2000. ISSN 0035-9009
2000
-
[11]
Rabier, G
Ernst Klinker, F. Rabier, G. Kelly, and J‐F Mahfouf. The ecmwf operational implementation of four‐dimensional variational assimilation. iii: Experimental results and diagnostics with operational configuration. Q. J. Royal Meteorol. Soc., 126 0 (564): 0 1191--1215, 2000. ISSN 0035-9009
2000
-
[12]
A strategy for operational implementation of 4d‐var, using an incremental approach
Philippe Courtier, J‐N Thépaut, and Anthony Hollingsworth. A strategy for operational implementation of 4d‐var, using an incremental approach. Q. J. Royal Meteorol. Soc., 120 0 (519): 0 1367--1387, 1994. ISSN 0035-9009
1994
-
[13]
Dueben, Sebastian Scher, Jonathan A
Stephan Rasp, Peter D. Dueben, Sebastian Scher, Jonathan A. Weyn, Soukayna Mouatadid, and Nils Thuerey. Weatherbench: a benchmark data set for data‐driven weather forecasting. J. Adv. Model Earth Syst., 12 0 (11): 0 e2020MS002203, 2020. ISSN 1942-2466
2020
-
[14]
A neural-network based mpas-shallow water model and its 4d-var data assimilation system
Xiaoxu Tian, Luke Conibear, and Jeffrey Steward. A neural-network based mpas-shallow water model and its 4d-var data assimilation system. Atmosphere, 14 0 (1): 0 157, 2023. ISSN 2073-4433. URL https://www.mdpi.com/2073-4433/14/1/157
2023
-
[15]
Watkins, Jordan S
Arka Daw, Anuj Karpatne, William D. Watkins, Jordan S. Read, and Vipin Kumar. Physics-guided neural networks (pgnn): An application in lake temperature modeling, pages 353--372. Chapman and Hall/CRC, 2022
2022
-
[16]
Alan J. Geer. Learning earth system models from observations: machine learning or data assimilation? Philos. Trans. R. Soc. A, 379 0 (2194): 0 20200089, 2021. ISSN 1364-503X
2021
-
[17]
Development of the tangent linear and adjoint models of the mpas-atmosphere dynamic core and applications in adjoint relative sensitivity studies
Xiaoxu Tian and Xiaolei Zou. Development of the tangent linear and adjoint models of the mpas-atmosphere dynamic core and applications in adjoint relative sensitivity studies. Tellus A, 72 0 (1): 0 1--17, 2020. ISSN 1600-0870
2020
-
[18]
Validation of a prototype global 4d-var data assimilation system for the mpas-atmosphere model
Xiaoxu Tian and Xiaolei Zou. Validation of a prototype global 4d-var data assimilation system for the mpas-atmosphere model. Mon. Wea. Rev., 149 0 (8): 0 2803--2817, 2021. ISSN 0027-0644
2021
-
[19]
Evolutions of errors in the global multiresolution model for prediction across scales - shallow water (mpas-sw)
Xiaoxu Tian. Evolutions of errors in the global multiresolution model for prediction across scales - shallow water (mpas-sw). Q. J. Royal Meteorol. Soc., 147 0 (734): 0 382--391, 2020. ISSN 0035-9009. doi:10.1002/qj.3923. URL https://doi.org/10.1002/qj.3923
2020 doi
-
[20]
Forecasting global weather with graph neural networks
Ryan Keisler. Forecasting global weather with graph neural networks. arXiv preprint arXiv:2202.07575, 2022
2022 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.