{"id":"f2599d82-eba5-45e8-ab27-abfe2059de9c","arxiv_id":"2412.15687","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A graph-neural-network weather model trained only on raw observations produces skillful global forecasts out to five days, with tropical 2-meter temperature forecasts competitive with the operational IFS.","lead":"ECMWF researchers built GraphDOP, a weather forecasting model trained and started entirely from satellite and ground observations, with no use of the gridded reanalysis fields that power today's AI weather models. It produces useful global forecasts up to five days ahead, and for some variables like tropical 2-meter temperature it matches or beats ECMWF's physics-based system.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline IFS comparison is confounded: GraphDOP uses a full 12-hour observation window while IFS early-delivery uses ~5 hours less data, so the day-5 tropical t2m 'advantage' may be an initialization artifact rather than forecast skill.","rationale":"The reader identified a genuine limitation—the ERA5-based QC filter introduces partial reanalysis dependence into the training data, so the phrase 'trained exclusively from observations' is not literally true. That concern is acknowledged in the paper and may be minor in practice, but it is not the single most load-bearing issue for the central claim. The central quantitative claim is the day-5 tropical t2m comparison with IFS, and that comparison is directly confounded by the stated ~5-hour initialization advantage: GraphDOP sees five additional hours of observations, so its day-5 forecast is effectively shorter-lead than the IFS early-delivery forecast it is scored against. The paper reports no analysis of how much this advantage contributes to the headline result. The grid-space t2m degradation reported in Section 5.2 is additional evidence that the observation-space skill does not straightforwardly generalize, but the initialization mismatch alone is sufficient to make the headline claim unverified. I therefore recommend keeping the conditional verdict, with the requirement that the authors either correct the comparison or quantify the 5-hour effect.","tokens_in":19370,"tokens_out":12390,"duration_ms":114230,"concrete_test":"Recompute Figure 7 using IFS forecasts initialized from the operational 12-hour-window analysis instead of the early-delivery analysis, keeping all other verification choices fixed. If the day-5 tropical t2m normalized RMS advantage of GraphDOP shrinks or reverses, the headline comparison is an artifact of the ~5-hour initialization mismatch.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest quantitative claim—that GraphDOP has smaller t2m forecast departures than IFS over the Tropics at day 5—rests on Section 5.1 and Figures 7–9. That comparison is not initialized on equal footing. Section 5.1 states: 'For technical reasons, the IFS operational forecasts initialised from the early-delivery analysis had to be used in this study which means that the GraphDOP forecasts have an approximately 5-hour advantage.' GraphDOP ingests a full 12-hour observation window (09z–21z or 21z–09z), whereas the early-delivery IFS analysis is produced from shorter windows (09z–16z or 21z–04z). At a fixed valid time, GraphDOP therefore has 5 hours of more recent observations, i.e., a shorter effective lead time. For t2m, where the diurnal cycle and surface boundary layer have short memory, 5 hours can substantially affect error statistics. The paper provides no sensitivity test or correction for this mismatch. Consequently, the abstract and Section 1 claim that GraphDOP t2m is competitive with IFS at day 5 is not established by the evidence as presented. The Section 5.2 grid-space IFS comparisons share the same confound, and the additional unexplained gap between observation-space and grid-space t2m skill further suggests the headline result is not robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"GraphDOP is an end-to-end graph-neural-network weather forecast system that is trained and initialised from conventional and satellite observations rather than from gridded reanalysis fields. The paper describes the observation dataset and quality-control choices, the encoder-processor-decoder architecture on a latent O96 mesh, the weighted MSE objective, and qualitative case studies covering IASI radiance evolution, an Arctic sea-ice freezing event, and Hurricane Ian. Quantitative verification is performed in observation space against matched conventional and satellite observations, with operational IFS as a benchmark, and in grid space against ERA5 with persistence and climatology baselines. The headline claims are that the model produces skilful forecasts up to five days and that its two-metre temperature forecasts are competitive with, and over the Tropics at day five better than, the operational IFS.","tokens_in":19639,"tokens_out":5269,"duration_ms":47124,"significance":"If correct, this would be a substantial step toward a new forecasting paradigm: a data-driven system that learns atmospheric dynamics and observation operators jointly from observational time series, avoiding the cost and complexity of 4D-Var while still exploiting the full observing system. The paper is commendably explicit about several limitations, including the ERA5-based quality control of training data (§2), the early-delivery IFS comparison (§5.1), and the gap between grid-space and observation-space t2m skill (§5.2). The matched-observation verification and the inclusion of persistence and climatology baselines are strengths. However, the two central claims — 'exclusively from observations' and 'competitive with IFS at day five' — are currently weakened by a known but unquantified 5-hour observation-window advantage in the IFS comparison and by a training-data quality-control dependence on ERA5. These concerns are addressable with additional experiments and more cautious claims, so the paper merits a major revision rather than rejection.","major_comments":[{"comment":"The abstract claims that GraphDOP is 'trained and initialised exclusively from Earth System observations, with no physics-based (re)analysis inputs or feedbacks', but §2 states that conventional observations are filtered with 'a conservative ERA5 departure (observation minus forecast) QC check' and that this 'introduces a partial dependence of the dataset generation pipeline on the reanalysis system'. Because the conventional observations are also the targets for t2m and other headline verification variables, this is not a cosmetic caveat: if the filter preferentially removes observations that disagree with ERA5, the strict observation-only claim is weakened and the ERA5-based grid-space verification in §5.2 becomes partially circular. Please either remove or substantially soften the exclusivity claim, quantify how many conventional observations are removed by the QC check and show that the retained t2m distribution is not materially altered, or retrain without the ERA5-based filter.","section":"Abstract and §2"},{"comment":"The day-5 tropical t2m and AMSU-A channel-5 advantages over IFS are confounded by the stated ~5-hour advantage of GraphDOP relative to the IFS early-delivery analysis. Section 5.1 explains that GraphDOP uses a full 12-hour window (09z-21z or 21z-09z) while the early-delivery IFS forecasts are initialised from shorter windows (09z-16z or 21z-04z), so at a fixed valid time the two systems are not compared at equal effective lead times. For a surface variable with short memory such as t2m, an extra five hours of observations can plausibly explain part or all of the reported day-1 15% improvement and the tropical day-5 advantage. The paper provides no sensitivity test or correction for this mismatch; please add an experiment in which GraphDOP is initialised from the same truncated window as IFS early delivery, or in which the comparison is made against IFS forecasts from the late-delivery analysis, before the 'competitive with IFS' claim is made.","section":"§5.1 and Figures 7–9"},{"comment":"The grid-space verification shows a rapid growth of t2m RMSE from day 3 onward, yet the paper's headline t2m claim is based on the observation-space departures of §5.1, which are 'much lower and in line with results presented in Section 5.1'. The text says this discrepancy is 'currently being investigated', but the t2m result is the flagship quantitative claim of the paper. Please provide a quantitative explanation of the discrepancy (for example, station-elevation representativeness, diurnal sampling, or the difference between point observations and grid-cell averages), or explicitly state that the claimed IFS competitiveness applies only to station-location verification and not to gridded forecasts.","section":"§5.2 and Figure 11"},{"comment":"All quantitative verification is computed over one boreal winter (December 10, 2022 to February 28, 2023) for the observation-space comparisons and one month (January 2023) for the grid-space comparisons, with no confidence intervals or seasonal breakdown. The tropical day-5 improvements shown in Figures 7 and 9 could be specific to this season or to the particular observing-system configuration of the period. Please add uncertainty estimates, a second independent evaluation period, or a clear statement of the seasonal/regime dependence of the reported skill differences.","section":"§5.1–5.2 and Figures 7–11"}],"minor_comments":[{"comment":"The title appears as 'GRAPH DOP: T OWARDS SKILFUL...' in the preprint text; please correct the spacing and capitalization.","section":"Title page"},{"comment":"The citation to Lessig (2025) as 'Manuscript in preparation' should be updated to a citable preprint if one becomes available, or removed from the reference list.","section":"References"},{"comment":"Please clarify the summation convention in Equation (1): the normalization by |I| × |C| × |O| is written as a product after the sums, and it is not immediately clear how the sets I, C, and O are nested; a worked example for one instrument would help.","section":"§3, Eq. (1)"},{"comment":"Figure 15 is introduced as a sensitivity result from a network trained with a 3-hour decoder output interval, but this alternative model is not described elsewhere; please provide its training setup or clearly label it as an illustrative preliminary result.","section":"Appendix, Figure 15"},{"comment":"The training section states that 70,000 steps were run on 64 H100 GPUs, but it does not give the total compute cost or wall-clock time; adding this information would aid reproducibility and resource planning.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"This is an internally consistent and candid paper from a major operational centre, and I see no grounds for rejection. The main revision needs are: (1) qualifying or defending the 'exclusively from observations' claim given the ERA5-based QC; (2) removing or controlling the 5-hour window confound in the IFS comparison; and (3) addressing the t2m grid-space versus observation-space gap. The paper also leans heavily on ECMWF-internal reporting conventions; an independent reader would benefit from a clearer separation of established results from ongoing work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nGraphDOP is the first global medium-range forecast system I know of that is trained and initialized from raw observations only, with no reanalysis input to the model itself. That is a real milestone, and the paper is worth reading. The architecture (GNN encoder-processor-decoder over a latent O96 mesh) is sensible, the observation-space verification against IFS is a good idea, and the qualitative results—sea-ice radiance, Hurricane Ian—show the system has learned meaningful dynamics.\n\nThe authors are also unusually candid about limitations, which makes me trust the reporting more. They disclose the ERA5-based QC filter on conventional observations, the 5-hour advantage over IFS early-delivery forecasts, the short evaluation window, and the absence of code/data.\n\nNow the soft spots, in proportion. The abstract claims 'exclusively from observations, with no physics-based (re)analysis inputs.' That is not literally true: the training pipeline uses a conservative ERA5 departure QC to filter gross outliers. The authors acknowledge this, and I agree it is a minor leak rather than a fatal one—the QC is coarse and could be removed in a future version—but the abstract should say 'training targets' rather than 'exclusively from observations.'\n\nThe bigger concern is the IFS comparison. The stress-test note is right: GraphDOP ingests a full 12-hour window, while the IFS early-delivery analysis uses ~7-hour windows, so GraphDOP has roughly five hours of more recent data at a fixed valid time. For t2m, where the diurnal cycle and boundary layer have short memory, that can materially affect error statistics. The paper admits this but does not quantify the effect. Consequently, the claim that GraphDOP is 'competitive with' or 'better than' IFS over the Tropics at day 5 is not established by the evidence. It is a suggestive result, not a solid one. The observation-space t2m versus grid-space t2m discrepancy (RMSE much lower at observation points than on the grid) adds to the unease—the model may be good at hitting stations but not as good on the full field.\n\nThe skill against persistence and climatology out to day 5, though, is not affected by the IFS confound, and that is the core proof-of-concept. The comparison with IFS should be de-emphasized or properly corrected (for example, run IFS from a longer window or quote both systems with matched lead time). The evaluation period (one winter) is thin, and no error bars are given, so the skill numbers are suggestive rather than definitive.\n\nWho is this for? Anyone working on data-driven weather forecasting, especially the 'learn from observations without reanalysis' line. It deserves a serious referee: the idea is important, the failure modes are largely acknowledged, and the remaining issues are about precise claims, not about the basic feasibility.\n\nRecommendation: send to peer review, but insist the authors (1) either remove or qualify the ERA5-QC dependence in the abstract, (2) re-run the IFS comparison with matched observation windows or present the result as an initialization advantage rather than forecast skill, and (3) state whether code and data can be shared.\n\nIn short: a solid proof-of-concept with an overstated headline. Engage with it.","headline":"A genuine first for observation-only data-driven forecasting, but the abstract overclaims 'exclusively from observations' given ERA5-based QC, and the headline IFS comparison is confounded by a five-hour initialization advantage the paper itself admits.","tokens_in":20234,"tokens_out":3577,"would_cite":true,"duration_ms":29622,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a graph neural network exclusively on Earth System observations, without reanalysis inputs, yields medium-range forecasts that stay skilful to day five and beat the operational physics-based system on tropical two-metre…","keywords":["graph neural network","data-driven weather forecasting","observation-space learning","medium-range forecast","satellite brightness temperatures","direct observation prediction","two-metre temperature","latent state representation"],"falsifier":"Retrain GraphDOP with the reanalysis-based quality-control filter disabled and evaluate its day-five two-metre temperature departures over the Tropics against independent in-situ observations that never entered any filter; if the advantage over the physics-based system disappears or reverses, the reanalysis dependence is carrying the skill.","tokens_in":19139,"feed_emoji":"🌡️","tokens_out":8348,"duration_ms":66091,"temperature":0.7,"pith_summary":"This paper introduces GraphDOP, a forecast system that is trained and initialised exclusively from Earth System observations, with no gridded reanalysis fields used as inputs or training targets. The authors claim that a graph neural network can learn the correlations between satellite radiances and conventional in-situ measurements well enough to produce skilful global forecasts of surface and upper-air weather out to five days. In their evaluation, GraphDOP's two-metre temperature forecasts are competitive with those of the operational physics-based system, with smaller departures from verifying observations over the Tropics at day five. A sympathetic reading of the paper is that it establishes observation-only learning as a viable path toward medium-range weather prediction, even if full parity with established systems remains an open goal.","feed_headline":"Observation-only weather AI stays skilful through day five","feed_subtitle":"Learned from satellite and in-situ data alone, it matches a physics-based model on tropical 2m temperature at day five.","key_machinery":"The central machinery is an encoder–processor–decoder graph neural network whose encoder and decoder are built on the fly from the latitudes and longitudes of observations, projecting irregular observations onto and off a latent state of 40,320 nodes on a reduced Gaussian grid of roughly one-degree spacing. The processor is a transformer with windowed attention that advances the latent state through the forecast window, and autoregressive rollout extends the forecast to longer lead times. Because the decoder only needs observation metadata to produce a forecast, predictions can be issued at arbitrary locations and times. The training objective is a weighted mean squared error over all observation targets, $\\mathcal{L}_{\\mathrm{DOP}} = \\tfrac{1}{T |I| |C| |O|} \\sum_{t,i,c,o} w_i w_{c,i} (y_{tico}-\\hat{y}_{tico})^2$, which balances per-instrument and per-channel weights so that the network learns to predict every observed quantity jointly. A quarter of satellite observations and half of the conventional observations are dropped during training to prevent overfitting.","core_discovery":"The paper's central claim is that an end-to-end graph neural network, trained and initialised directly from conventional observations and level-1 satellite brightness temperatures with no physics-based reanalysis inputs, can form a coherent latent representation of the Earth system and produce skilful medium-range forecasts. Concretely, GraphDOP forecasts surface and upper-air parameters up to five days ahead; its two-metre temperature forecasts are competitive with those of the operational physics-based forecast system, and it has smaller departures from verifying observations over the Tropics at day five. The network reproduces synoptic-scale features in brightness-temperature space, including moving frontal cloud bands, jet-stream structures, sea-ice growth during a rapid Arctic freezing event, and the trajectory and intensification of Hurricane Ian. The authors acknowledge that a reanalysis-based quality-control filter is applied to some conventional observations during dataset generation, and they stress that reanalysis fields are used only for verification, never for training or initialisation.","pith_inferences":["The strongest caveat is the paper's own: a reanalysis-based quality-control filter removes gross outliers from conventional observations before training. If that filter is what teaches the model to agree with the reanalysis, the reported advantage might partly reflect the verification target rather than independent skill; retraining without the reanalysis-dependent QC is a direct test.","The model's ability to forecast at arbitrary locations suggests it could be repurposed as a learned observation operator or an observation-space prior inside a hybrid data-assimilation cycle, feeding conventional analysis systems with radiance-to-state relationships.","The deterministic WMSE objective visibly smooths forecasts at long lead times; switching to a probabilistic or diffusion-based objective would likely sharpen features and may change the day-five comparison.","The headline comparison covers one winter season (December 2022 to February 2023); seasonal and interannual robustness of the tropical skill is untested, and the authors state that more work is needed on diurnal, seasonal, and regional error dependence."],"forward_implications":["If the claim holds, producing a global medium-range forecast no longer requires running a large data-assimilation system: the forecast is initialised directly from the latest observation windows and can be issued within minutes of data arrival.","A purely observation-driven forecast model can be verified and used on any grid or location, including regions with sparse or no conventional observations, because the decoder is not tied to fixed analysis fields.","The reported day-five tropical two-metre temperature skill suggests that data-driven systems may challenge physics-based systems in data-sparse regions first, not in well-observed mid-latitudes.","Because the model ingests cloudy and surface-sensitive radiances directly, it opens a path to exploiting observation types that traditional variational assimilation must reject or approximate with complex operators.","The fully differentiable model can be used to compute adjoint sensitivities, giving a new tool for estimating the information content of individual observation types for forecast skill."],"supporting_citations":[{"why":"Introduces the AI-DOP concept of learning forecasts directly from observations, which GraphDOP operationalises.","marker":"[McNally et al., 2024a]"},{"why":"Presents Aardvark, the prior hybrid end-to-end system pre-trained on reanalysis and fine-tuned on observations, which GraphDOP contrasts with and extends by removing reanalysis training.","marker":"[Vaughan et al., 2024]"},{"why":"Supplies the encoder–processor–decoder architecture with windowed attention and sequence parallelism that GraphDOP is built around.","marker":"[Lang et al., 2024a]"},{"why":"Documents the reanalysis used for the quality-control filter on conventional observations and for grid-space verification.","marker":"[Hersbach et al., 2020]"},{"why":"Basis for the approximately five-hour advantage GraphDOP holds over the physics-based system in the observation-space comparison, which must be accounted for when interpreting the skill scores.","marker":"[Lean et al., 2021]"},{"why":"Provides the method for computing observation equivalents from physics-based forecasts, used to calculate forecast departures in observation space.","marker":"[Dahoui et al., 2016]"},{"why":"Supplies the persistence and climatology baselines used in grid-space verification.","marker":"[Rasp et al., 2024]"},{"why":"Establishes the autoregressive rollout technique used to extend GraphDOP forecasts to longer lead times.","marker":"[Keisler, 2022]"}],"fun_headline_variants":["No-reanalysis AI weather forecasts skilful through day five","Weather AI that never touches reanalysis: skilful to day five","AI weather model trained on obs alone matches physics at day five","From raw satellite data to 5-day forecasts: GraphDOP","End-to-end weather AI from observations: skilful day-five forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that filtering conventional observations by how well they agree with a reanalysis dataset does not materially shape what the model learns; if that filter is doing the work, the system is not truly trained from observations alone and its reanalysis-verified skill could be partly circular.","fun_headline_variants_meta":{"raw":{"variants":["No-reanalysis AI weather forecasts skilful through day five","Weather AI that never touches reanalysis: skilful to day five","AI weather model trained on obs alone matches physics at day five","From raw satellite data to 5-day forecasts: GraphDOP","End-to-end weather AI from observations: skilful day-five forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00071,"raw_usage":{"total_tokens":3150,"prompt_tokens":851,"completion_tokens":2299,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":2208}},"tokens_in":467,"tokens_out":2299,"duration_ms":14558,"temperature":1.0,"reasoning_tokens":2208,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:10:30.510290+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain GraphDOP with the reanalysis-based quality-control filter disabled and evaluate its day-five two-metre temperature departures over the Tropics against independent in-situ observations that never entered any filter; if the advantage over the physics-based system disappears or reverses, the reanalysis dependence is carrying the skill.","supporting_citations":[{"cited_title":"Continuous data assimilation for global numerical weather prediction","cited_arxiv_id":null,"evidence_quote":"Basis for the approximately five-hour advantage GraphDOP holds over the physics-based system in the observation-space comparison, which must be accounted for when interpreting the skill scores."},{"cited_title":"Use of forecast departures in verification against observations","cited_arxiv_id":null,"evidence_quote":"Provides the method for computing observation equivalents from physics-based forecasts, used to calculate forecast departures in observation space."}],"review_version":1}