{"id":"350e8249-de22-4dfb-8fb4-2cd0c12e2774","arxiv_id":"2411.14677","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The tangent linear and adjoint versions of GraphCast and NeuralGCM produce noisier and more localized sensitivity patterns than those of MPAS-A, calling into question their readiness for data assimilation.","lead":"This study builds and tests the tangent linear and adjoint versions of two machine learning weather models, GraphCast and NeuralGCM, for use in data assimilation systems. The results show noisy and localized sensitivity patterns compared with a traditional model, suggesting these ML models are not yet ready for operational data assimilation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'unphysical' verdict is confounded by the benchmark: the MPAS-A adjoint in §2.3 omits all unresolved physics, while the ML adjoints include them, and the paper never demonstrates its claim that the test cases minimize this gap.","rationale":"The paper is attempting to evaluate whether ML weather model adjoints are physically trustworthy enough for data assimilation. For that central claim to hold, the observed differences must be attributable to properties of the ML models rather than to an unfair or incomplete reference. The weakest condition is therefore the benchmark itself. The paper's own §2.3 flags that the MPAS-A adjoint excludes unresolved physics while the ML adjoints include them, and the assertion that the test cases minimize this impact is not supported by any experiment. The GraphCast TLM/adjoint consistency checks are real evidence that the linearizations are correct for GraphCast, but correctness with respect to the ML model does not establish physical realism; it only rules out implementation bugs. NeuralGCM additionally lacks such verification, but the benchmark asymmetry is the more fundamental concern because it affects the interpretation of both models. A full-physics reference would settle whether the reported on-the-spot sensitivities, vertical noise, and specific-humidity amplification are genuinely unphysical ML artifacts or simply unresolved physics missing from the MPAS-A adjoint. Since the reader already identified this as the weakest assumption and issued a conditional verdict, my stress-test does not change that verdict.","tokens_in":8391,"tokens_out":5348,"duration_ms":83302,"concrete_test":"Build a full-physics reference for the same test cases: run the nonlinear MPAS-A model with all parameterizations enabled and approximate the tangent-linear response by finite differences for the same zonal-wind perturbation and the same T=1 response function, validating the perturbation amplitude against the no-physics MPAS-A TLM. If the on-the-spot sensitivity, the vertical noise, and the specific-humidity magnitudes persist relative to this full-physics reference, the central 'unphysical' claim is supported; if they largely match the no-physics benchmark, the missing unresolved physics in the MPAS-A adjoint is the source of the discrepancy and the conclusion needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 states that the MPAS-A adjoint 'only describes the evolution of resolved scale processes,' while the GraphCast and NeuralGCM adjoints 'inherently include' unresolved processes. The paper then asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes,' but no test or metric is given to support that assertion. The central conclusion that the ML adjoints are unphysical rests on a comparison against this physics-free reference. The reported symptoms — persistent sensitivity at the perturbation location, noisy vertical structures, and specific-humidity amplitudes an order of magnitude larger than MPAS-A — are exactly the kind of signatures that missing moist convection, radiation, and microphysics would produce in a linearized model. The GraphCast TLM is correctly verified against its own nonlinear model, so the patterns are not simple coding errors; the issue is whether they are 'unphysical' relative to the full atmospheric model or merely relative to a truncated dry-dynamics benchmark. Because the paper lacks a full-physics control and any quantitative similarity metric, the claim that these properties are ML-specific and would degrade DA is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops tangent-linear (TL) and adjoint (AD) models for two machine-learning weather models, GraphCast and NeuralGCM, via automatic differentiation and linearization of their internal operations. It verifies the GraphCast TLM against finite differences of the nonlinear model and checks adjoint consistency with the inner-product identity, reporting agreement to 15 digits. It then compares six-hour TL propagations and adjoint sensitivities for a zonal-wind perturbation near a jet core against the MPAS-A model's adjoint, which includes only the resolved dynamical core and no adjoints of unresolved physical parameterizations. Based on visual inspection of horizontal and vertical sensitivity patterns, the paper concludes that the GraphCast and NeuralGCM adjoints exhibit unphysical localized sensitivities, noisy vertical structures, and exaggerated specific-humidity magnitudes, and argues that these features would degrade 4DVar and ensemble-based data assimilation.","tokens_in":8652,"tokens_out":6003,"duration_ms":58298,"significance":"The question addressed is timely and important: the viability of ML weather models in variational DA depends directly on the quality of their linearized operators. The paper's verification of the GraphCast TLM and adjoint is a real strength, meeting standard Taylor and inner-product consistency checks. If the central conclusion were established, the paper would provide a concrete and useful caution for DA integration of ML models. However, the conclusion is not yet established: the MPAS-A benchmark lacks adjoints of unresolved physics, the assertion that the chosen test cases minimize that gap is unquantified, the NeuralGCM adjoint is not verified, and the 'unphysical' judgment rests on unquantified, visual pattern comparison. With additional control experiments and quantitative diagnostics, the paper could become a solid assessment of ML-model linearizations for DA.","major_comments":[{"comment":"The decisive comparison is between ML adjoints that include all trained processes and an MPAS-A adjoint that, as the paper states in §2.3, 'only describes the evolution of resolved scale processes.' The symptoms used to label the ML adjoints unphysical—persistent on-the-spot sensitivity, noisy vertical structures, and specific-humidity amplitudes an order of magnitude larger than MPAS-A—are also the signatures that unresolved moist convection, radiation, and microphysics would imprint on a linearized model. The paper asserts that 'some of the test cases considered are chosen so as to minimize the impact of the unresolved processes,' but no metric or experiment is provided to support that assertion. The central claim that the ML adjoints are unphysical relative to the real atmosphere therefore rests on an unverified benchmark-comparability assumption. To support the claim, the authors should either include a full-physics adjoint control (or a nonlinear full-physics MPAS-A run with finite differences) or quantitatively demonstrate that the chosen cases are insensitive to unresolved processes.","section":"§2.3, §3"},{"comment":"Verification of the tangent-linear and adjoint models is shown only for GraphCast: the Taylor test for each variable and the inner-product consistency check are reported in §2.2. No equivalent verification is shown for NeuralGCM, even though Fig. 5 and Section 3 interpret NeuralGCM adjoint sensitivities as physically meaningful with noise. Without a Taylor test or adjoint consistency check for NeuralGCM, the reader cannot distinguish model-intrinsic unphysicality from implementation error. The paper should either provide those verification results or clearly state that the NeuralGCM TL/AD results are provisional pending verification.","section":"§2.2, §3, Fig. 5"},{"comment":"The judgments of physical realism are made entirely by visual comparison of horizontal maps and vertical cross-sections. Terms such as 'slightly noisier,' 'quite noisy,' 'strong on-the-spot feature,' and 'an order of magnitude stronger' are not backed by a quantitative diagnostic. A reproducible assessment would require, for example, pattern correlation coefficients between the ML and MPAS-A sensitivities, scale decomposition or spectral filtering to characterize the noise, and amplitude ratios with stated significance levels. As the paper stands, the claim that the differences indicate unphysical behavior rather than legitimate differences in model formulation is not quantitatively supported.","section":"§3, Figs. 2–5"},{"comment":"The implementation details of the GraphCast and NeuralGCM TL/AD models are insufficient for reproducibility or for judging whether the compared fields are on equal footing. Key items are not specified: the horizontal and vertical resolution of the ERA5 initial state and of each model, whether the adjoint includes the input/output normalization layers of the neural networks, how the adjoint of the multi-step forecast operator is chained, and whether the NeuralGCM adjoint includes the dynamical core plus the neural-network parameterizations in a single consistent linearization. These details should be reported so that the comparison is interpretable and reproducible.","section":"§2, §3"}],"minor_comments":[{"comment":"The keyword 'Data tssimilation' contains a typo; it should read 'Data assimilation.'","section":"Abstract and Keywords"},{"comment":"The test statistic is denoted Ψ(α) in the text but J(α) in the caption; please unify the notation.","section":"§2.2 and Fig. 1 caption"},{"comment":"The response function is referred to as 'T = 1' in some places and 'Tad = 1' in others; please make the notation consistent.","section":"Figs. 3 and 4 captions and §3"},{"comment":"The statement that adjoint noise 'would distort error covariances' in EnKF is presented as fact; consider rewording to 'may distort' or adding a supporting demonstration.","section":"§4"},{"comment":"The paper does not identify the automatic differentiation framework (e.g., JAX, PyTorch) used to construct the adjoints; adding this would aid reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an early preprint with a solid GraphCast verification component, but the benchmark asymmetry between the full-physics ML adjoints and the dry-dynamics MPAS-A adjoint is central to the paper's conclusion. I recommend major revision: the authors should add a full-physics control or quantitative evidence that omitted unresolved processes do not drive the comparison, verify the NeuralGCM adjoint, and introduce quantitative pattern metrics. The EnKF implications in the discussion are speculative and should be toned down. No concerns about citation practice or novelty disclosure beyond the unverified NeuralGCM results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is the first to build and verify tangent linear/adjoint versions of GraphCast and NeuralGCM for data assimilation, and that work is internally solid. The GraphCast Taylor test and adjoint inner-product check agree to 15 digits, which is real evidence the codes are correct. The visual comparisons are also worth having: persistent on-the-spot sensitivity, noisy vertical structures, and inflated specific-humidity amplitudes are exactly the features a DA developer would want flagged.\n\nThe trouble is the interpretation. The MPAS-A adjoint is a dry-dynamics core with no adjoint for convection, radiation, or microphysics. The ML adjoints include those processes. The symptoms the paper calls unphysical — localized persistence, vertical noise, strong humidity sensitivity — are what you would expect from a linearized model with active physics set against a physics-free baseline. The paper acknowledges this in Section 2.3 and asserts that the test cases minimize the impact of unresolved processes, but it never shows that. No metric, no sensitivity experiment, no full-physics control. So the central claim that these ML adjoints are unphysical is not established. A more accurate statement would be that they differ from a truncated benchmark, not that they are unphysical.\n\nMinor issues: NeuralGCM gets no Taylor or adjoint consistency check of its own; the EnKF implications are speculative extrapolation; and there are no quantitative similarity metrics — just visual inspection. These are fixable, and the paper's value as a cautionary signal survives even if the verdict is overdrawn.\n\nWho is this for? Anyone thinking about putting ML weather models into 4DVar or ensemble DA. It is a preliminary evaluation, not a definitive one, but it is an honest and technically competent start. I would send it to peer review, with the expectation that the conclusions get toned down and the benchmark-dependence issue gets addressed.","headline":"A useful first look at GraphCast and NeuralGCM adjoints for DA, but the central 'unphysical' verdict is confounded by comparing against a physics-free MPAS-A baseline.","tokens_in":9092,"tokens_out":1930,"would_cite":false,"duration_ms":21872,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ML weather adjoints are too noisy for data assimilation","keywords":["machine learning weather models","data assimilation","4DVar","tangent linear model","adjoint model","GraphCast","NeuralGCM","MPAS-A"],"falsifier":"Run the same single-point jet perturbation with an MPAS-A adjoint that includes tangent-linear and adjoint versions of convection, radiation, and microphysics (or with a full-physics NWP adjoint such as an operational 4DVar system's). If the on-the-spot spike, vertical noise, and humidity magnitude ratios persist in that symmetric comparison, the paper's verdict stands; if they disappear, the unphysical properties are artifacts of comparing full-physics ML adjoints to a resolved-scale-only benchmark.","tokens_in":8222,"feed_emoji":"🌦️","tokens_out":8046,"duration_ms":66045,"temperature":0.7,"pith_summary":"The paper asks whether two leading machine-learning weather models, GraphCast and NeuralGCM, can supply the tangent linear and adjoint operators that four-dimensional variational data assimilation requires. The authors build those operators by linearizing the models' learned operations and check them against the adjoint of a conventional dynamical-core model, MPAS-A. They find that, even where the formal tangent-linear and adjoint consistency checks pass, the ML adjoints are physically unreliable: responses persist at the perturbation point, vertical patterns are noisy, and specific-humidity sensitivities are exaggerated. The paper's central claim is that these unphysical properties make the current ML linearized models unsuitable for operational 4DVar and likely harmful for ensemble assimilation as well.","feed_headline":"ML weather adjoints are too noisy for data assimilation","feed_subtitle":"Their adjoints show unphysical spikes at the perturbation site, so 4DVar increments would be noisy.","key_machinery":"The machinery is the tangent linear and adjoint pair derived from each model: for GraphCast, the Jacobian of the graph-neural-network update rules and its transpose; for NeuralGCM, the Jacobian of the deep neural network parameterizations combined with the linearized dynamical core, obtained by automatic differentiation. The operators are verified by a finite-difference ratio test and the exact adjoint identity before being compared with the benchmark MPAS-A adjoint on identical initial states and perturbations. The comparison itself is the instrument: physical realism is judged by whether adjoint sensitivity appears upstream in a Rossby-wave pattern, stays smooth in the vertical, and avoids implausibly large amplitudes at the perturbation location.","core_discovery":"The central claim is that the current tangent linear and adjoint models of GraphCast and NeuralGCM exhibit several unphysical properties that limit their suitability for data assimilation applications. In a series of single-point perturbation experiments on a jet-stream state, GraphCast's tangent linear model keeps a strong response at the original perturbation spot after six hours, a feature absent from MPAS-A; its adjoint shows no upstream temperature sensitivity except at that spot, noisy zonal-wind patterns, and specific-humidity sensitivities roughly an order of magnitude stronger than MPAS-A's. NeuralGCM's adjoint captures some upstream, Rossby-like patterns but adds vertical noise, likely from the linearized neural-network parameterizations. If a single observation were assimilated, the resulting 4DVar increment would be noisy and contain unphysical structures far from the observation, and ensemble methods would inherit distorted error covariances.","pith_inferences":["Inference: the persistent on-the-spot sensitivity may reflect the ML models' reliance on local input features and short-range learned correlations, producing a near-identity sensitivity at the perturbed grid cell rather than dynamically consistent wave propagation.","Inference: a direct testable extension would be to train or fine-tune the ML models with a physics-consistency loss penalizing adjoint noise and pointwise sensitivity, then re-run the same jet perturbation experiments.","Inference: the same single-point protocol could be applied to newer ML weather models; the paper establishes a cheap diagnostic (one perturbation, six-hour TLM/ADM check) for screening any candidate DA operator.","Inference: the exaggerated humidity sensitivities suggest the humidity channels in these ML models may be the first place to look for overfitting or data-imbalance artifacts."],"forward_implications":["A 4DVar system using GraphCast's adjoint would produce noisy increments with structures located far from the observation, worsening forecast error even if it fits the observed point.","NeuralGCM's adjoint, though more physical in upstream propagation, still needs smoothing and refinement before it can act as a 4DVar operator.","Ensemble Kalman Filter methods would inherit distorted background error covariances from these ML models' perturbation responses.","Both ML models show some capacity for Rossby-wave dynamics, indicating the core linearized dynamics may be salvageable with better treatment of learned physics."],"supporting_citations":[{"why":"introduces GraphCast, the graph-neural-network weather model whose tangent linear and adjoint operators this paper evaluates.","marker":"[Lam et al., 2023]"},{"why":"introduces NeuralGCM, the hybrid model whose neural-network parameterizations and linearized operators are tested here.","marker":"[Kochkov et al., 2024]"},{"why":"supplies the MPAS-A tangent linear and adjoint model used as the physical benchmark.","marker":"[Tian and Zou, 2020]"},{"why":"describes MPAS-A, the base forecast model from which the benchmark adjoint is built.","marker":"[Skamarock et al., 2012]"},{"why":"provides the incremental 4DVar strategy that motivates the need for physically consistent TL/AD operators.","marker":"[Courtier et al., 1994]"},{"why":"frames the relationship between machine learning and data assimilation, supporting the claim that distorted covariances harm ensemble DA.","marker":"[Geer, 2021]"}],"fun_headline_variants":["ML adjoints too noisy for operational data assimilation","GraphCast and NeuralGCM adjoints show unphysical spikes","4DVar gets noisy with ML weather adjoints","Neural weather adjoints fail the data assimilation test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that test cases chosen to minimize the effect of unresolved physics make MPAS-A's missing adjoint of those processes irrelevant; if that assumption fails, the differences blamed on the ML models could be artifacts of the benchmark.","fun_headline_variants_meta":{"raw":{"variants":["ML adjoints too noisy for operational data assimilation","GraphCast and NeuralGCM adjoints show unphysical spikes","4DVar gets noisy with ML weather adjoints","Neural weather adjoints fail the data assimilation test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000269,"raw_usage":{"total_tokens":1647,"prompt_tokens":999,"completion_tokens":648,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":584}},"tokens_in":615,"tokens_out":648,"duration_ms":7030,"temperature":1.0,"reasoning_tokens":584,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:00:28.446397+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same single-point jet perturbation with an MPAS-A adjoint that includes tangent-linear and adjoint versions of convection, radiation, and microphysics (or with a full-physics NWP adjoint such as an operational 4DVar system's). If the on-the-spot spike, vertical noise, and humidity magnitude ratios persist in that symmetric comparison, the paper's verdict stands; if they disappear, the unphysical properties are artifacts of comparing full-physics ML adjoints to a resolved-scale-only benchmark.","supporting_citations":[{"cited_title":"Learning skillful medium-range global weather forecasting","cited_arxiv_id":null,"evidence_quote":"introduces GraphCast, the graph-neural-network weather model whose tangent linear and adjoint operators this paper evaluates."},{"cited_title":"A strategy for operational implementation of 4d‐var, using an incremental approach","cited_arxiv_id":null,"evidence_quote":"provides the incremental 4DVar strategy that motivates the need for physically consistent TL/AD operators."}],"review_version":1}