{"id":"53e09776-ef54-4439-ab94-f18ef8bf6dd9","arxiv_id":"2509.15942","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ArchesClimate autoregressively generates 10-year monthly climate states with a flow-matching emulator trained on IPSL-CM6A-LR decadal hindcasts, producing ensembles that match the climate model for some variables.","lead":"ArchesClimate is an AI model that mimics a large climate model, generating 10 years of monthly ocean and atmosphere states at low cost. It is trained on the IPSL-CM6A-LR model's own decadal forecasts and aims to produce many alternative climate histories, called ensembles, cheaply.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Interchangeability evidence may be dominated by the seasonal cycle; the model's underpowered decadal variability undermines the central probabilistic-ensemble claim.","rationale":"The reader's verdict is CONDITIONAL and already flags the variance shortfall as a required condition, so our concern does not change the verdict. We partially agree with the reader's weakest assumption: the missing initial-condition information may contribute to low variance, but the more directly load-bearing issue is that the evaluation method (rank histograms on raw values) can conceal the variance deficit, and the paper's own spectral analysis shows the deficit exists at exactly the decadal timescales that the model claims to emulate. We diverge from the reader in that we do not think the core problem is the absence of initial-condition perturbations per se; an emulator can validly generate new members without reproducing specific initial perturbations, provided the ensemble distribution matches the target. The unresolved issue is that the distribution does not match: variance is roughly halved, and the attempts to fix it degrade accuracy. This does not falsify the paper, but it means the 'interchangeable' claim is only supported for the seasonal-cycle-dominated component, not for decadal variability. The proposed test would settle whether the rank-histogram evidence is an artifact of the seasonal cycle. The paper is transparent about its limitations, and the condition the reader places on reporting non-overlapping-split metrics and addressing the variance shortfall remains appropriate.","tokens_in":17440,"tokens_out":6018,"duration_ms":55906,"concrete_test":"Recompute the rank histograms of Section 3.6 on anomalies (raw values minus the monthly climatology over the training period, as in Section 3.1) and on band-pass filtered anomalies (periods 2-10 years) for the variables in Figure 6. If flatness seen in the raw-value rank histograms disappears for anomalies or low-frequency-filtered fields, the 'interchangeable' claim is an artifact of the seasonal cycle. As a second check, compute ensemble variance as a function of lead time for the first 1-5 years; if ArchesClimate spread does not grow comparably to IPSL-DCPP, the model is not capturing initial-condition uncertainty and decadal predictability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ArchesClimate generates ensembles 'interchangeable' with IPSL-DCPP rests on rank histograms computed on raw pixel values (Section 3.6). Raw values are dominated by the seasonal cycle and spatial mean, so flat rank histograms can coexist with severely deficient interannual-to-decadal variability. The paper's own Temporal Power Spectra (Section 3.7, Figure 6) show ArchesClimate is underpowered at all periods beyond the monthly, seasonal, and annual peaks for every variable examined. Table 2 further shows ArchesClimate variance is roughly half of IPSL-DCPP for all variables across all three periods (e.g., tos 0.26 vs 0.51, psl 30529 vs 109017, ta 0.68 vs 1.58). Section 3.10 demonstrates that attempts to restore variance (noise scaling, per-variable scaling) have little effect, and the alternate loss that improves variance degrades CRPS, leaving the variance shortfall unresolved. Because the stated purpose is to emulate decadal-scale internal variability, an under-dispersed ensemble that fails to capture low-frequency variance is not truly interchangeable, even if marginal rank statistics on raw fields appear flat. The 'physically consistent' claim is also only qualitative; no conservation or dynamical consistency test is reported. The most load-bearing concern is that the evaluation masks a genuine decadal-variance deficit, so the headline claim of probabilistic decadal ensemble generation is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ArchesClimate, a flow-matching based emulator of the IPSL-CM6A-LR climate model, trained on decadal hindcast outputs from the DCPP project. The model autoregressively generates monthly states from the two preceding months and is designed to produce probabilistic ensemble members at monthly-to-decadal timescales. The authors evaluate the model using CRPS, ensemble variance, rank histograms, temporal power spectra, Pearson correlation of decadal trends, and seasonal spatial anomaly maps, and they compare several training and variance-enhancement strategies. The central claims are that ArchesClimate generates stable, physically consistent 10-year sequences and that for several climate variables its ensemble members are interchangeable with IPSL-DCPP members.","tokens_in":17762,"tokens_out":4307,"duration_ms":38770,"significance":"If the central claims were fully established, ArchesClimate would be a valuable contribution: it is a relatively low-cost, open-source method for generating probabilistic ensemble members from a coupled climate model, and it extends the ArchesWeatherGen approach to a coupled ocean-atmosphere setting with external forcings. The paper is commendable for its transparent experimental design, the comparison of multiple training schemes, the inclusion of CRPS and rank-histogram diagnostics, the availability of code, and the explicit discussion of limitations in Appendices B and C. However, the headline claims of 'interchangeable' and 'physically consistent' are not yet supported by the evidence presented: the variance of generated ensembles is consistently lower than that of IPSL-DCPP, the temporal power spectra show deficits at periods beyond the annual cycle, and the physical-consistency assessment is qualitative. The significance is therefore conditional on strengthening these evaluations.","major_comments":[{"comment":"The interchangeability claim rests on rank histograms computed from raw pixel values, which are dominated by the seasonal cycle and the spatial mean. The paper's own Temporal Power Spectra (Section 3.7, Figure 6) show that ArchesClimate is underpowered at all periods beyond the monthly, seasonal, and annual peaks for every variable examined, and Table 2 shows that ArchesClimate variance is roughly half of IPSL-DCPP for all variables (e.g., psl 30529 vs 109017 Pa^2, ta 0.68 vs 1.58 K^2, tos 0.26 vs 0.51 K^2). A flat rank histogram on raw fields does not demonstrate that the generated ensembles reproduce decadal internal variability. The authors should either delimit the interchangeability claim to sub-annual-to-annual timescales or provide complementary diagnostics (e.g., rank histograms on anomalies, spectral variance ratios, or low-frequency variance scores) that directly assess decadal-scale calibration.","section":"Section 3.6 / Section 3.7 / Table 2"},{"comment":"The variance shortfall, which is acknowledged in the manuscript, is shown to be unresolved by the proposed remedies: noise scaling by 1.1, per-variable noise scaling, and the alternative spectral/gradient loss. The alternative loss improves variance for tos and ta but increases CRPS, and the authors state that balancing this tradeoff is left for future work. Since an under-dispersed ensemble undermines the core purpose of probabilistic decadal ensemble generation, the paper should provide a more definitive treatment, such as a calibrated variance-inflation procedure evaluated with proper scoring rules, or a demonstration that the deficit is confined to variables or timescales not central to the intended application.","section":"Section 3.10"},{"comment":"The test set includes time periods that overlap the training distribution (ensembles initialized in 1969, 1979, and 2010–2015, with other ensembles overlapping those decades), and the paper's own alternative non-overlapping split shows degraded skill, particularly for tos and ta toward the end of the decade. This means the headline interchangeability result may partly reflect interpolation within the training distribution rather than a general emulation capability. The manuscript should either adopt the stricter split for the central claims or explicitly state that the claims are limited to interpolation within the observed period, and quantify the degradation in the abstract and conclusions.","section":"Appendix C / Section 2.3"},{"comment":"The abstract's claim that the generations are 'physically consistent' is supported only by qualitative visual inspection of Figure 4; no quantitative conservation, dynamical-consistency, or process-level tests are reported. The Discussion itself states that adherence to conservation properties (mass, energy, hydrostatic constraints) remains future work. The claim should be softened to 'statistically plausible' or supported by a quantitative physical-consistency check before it is made in the abstract.","section":"Section 3.3 / Discussion"}],"minor_comments":[{"comment":"The notation in Equations (2) and (3) appears garbled, with mismatched parentheses and an ambiguous role for σ. Please rewrite these equations carefully and define all symbols, including the FM timestep discretization (ψ_m) used in Equation (4).","section":"Section 2.3, Eqs. (2)-(3)"},{"comment":"The text states that the dataset contains 'approximately 70,000 simulated months.' Given 10 members initialized every year from 1960 to 2015, each with 120 months, the total is 56 × 10 × 120 = 67,200 months before excluding validation and test sets; please clarify how the 70,000 figure is obtained.","section":"Section 2.1"},{"comment":"There is a typo 'ISPL-DCPP' in the Figure 1 caption, which should be 'IPSL-DCPP.' Additionally, Appendix B text refers to 'repeated forcings' while Figure B1's legend says 'constant forcings'; please align the terminology.","section":"Figure 1 caption / Figure B1"},{"comment":"The sentence 'The variance is consistently higher in IPSL-DCPP across all periods' is clearer than the later claim that CRPS being lower for ArchesClimate implies better performance; the paper correctly notes the ambiguity, but the discussion would benefit from a more explicit statement that lower CRPS than the baseline is not necessarily the goal when the aim is to replicate the target distribution.","section":"Section 3.5"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is within the scope of JAMES and the method is a reasonable step toward low-cost probabilistic emulation of coupled climate models. The main issue is that the abstract and conclusions overstate the evidence: the variance deficit, the spectral shortfall at decadal timescales, the qualitative nature of the 'physically consistent' claim, and the acknowledged train/test overlap all weaken the central claims. I recommend major revision with a request to either substantially temper the abstract or add the quantitative diagnostics that would support the stronger claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you care about ML climate emulators. The genuinely new bit is coupling ocean and atmosphere in one probabilistic monthly-step model, trained on IPSL-CM6A-LR decadal hindcasts, conditioned on greenhouse gases and solar forcing. The deterministic-transformer-plus-residual-flow scheme is a sensible adaptation of ArchesWeatherGen, and the paper is admirably candid: it reports its own variance shortfall, its test/train overlap, the degraded non-overlapping split, and the qualitative nature of the 'physically consistent' check. Code and data links are provided.\n\nWhat holds up: stable 10-year rollouts, flat rank histograms for net flux, wap, hur, and ta, reasonable CRPS on several variables, and a fair qualitative match on seasonal spatial patterns. For those variables, the emulator does look usable as a cheap source of plausible members.\n\nThe soft spots are mostly in the strength of the headline. The interchangeability evidence rests on rank histograms computed on raw fields, which are dominated by the seasonal cycle. The temporal power spectra show ArchesClimate is underpowered at every period beyond the monthly/seasonal/annual peaks, and the table shows variance at roughly half of IPSL-DCPP for all variables. That is exactly the wrong place to be if the goal is decadal internal variability. The attempts to fix variance (noise scaling, per-variable scaling, spectral loss) either do little or trade accuracy for variance, so the deficit is not resolved. The paper says initial-condition memory fades, but that is asserted, not demonstrated. And the overlapping test split inflates apparent skill; Appendix C shows the honest non-overlapping split is noticeably worse. These are acknowledged, but together they mean the 'interchangeable with IPSL' claim is scoped to a subset of variables at certain scales, not the full probabilistic ensemble.\n\nI would not call this a fatal flaw. The method is plausible, the engineering is real, and the limitations are placed in plain sight. But a referee should ask for headline metrics on the non-overlapping split, an effort to quantify low-frequency variance (e.g., spectra on anomalies, variance of decadal means), and either a fix or an explicit scoping of the claim to exclude decadal variability.\n\nMy verdict: send it to peer review, but expect heavy revision. It is a serious piece of work from which the community can learn, even if the central claim needs to be pulled back to match the evidence.","headline":"A serious and unusually transparent attempt at a coupled ocean–atmosphere decadal emulator, but the 'interchangeable' claim only holds for a subset of variables at seasonal/annual scales, and the paper's own spectra show an unresolved decadal-variance deficit.","tokens_in":18283,"tokens_out":1746,"would_cite":true,"duration_ms":18690,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ArchesClimate, a flow-matching emulator, produces stable 10-year climate states whose members are interchangeable with IPSL-CM6A-LR ensemble members for several key variables.","keywords":["flow matching","climate emulator","decadal climate prediction","ensemble generation","internal variability","probabilistic emulation","coupled climate model","IPSL-CM6A-LR"],"falsifier":"Generate a 10-year ArchesClimate ensemble from the same 1969 initialization and compare the deep-ocean heat-content variable (thetaot2000) against IPSL-DCPP: if its ensemble variance stays near half the target and its rank histogram stays strongly u-shaped across all three test decades, even after the proposed loss changes, then the interchangeability claim is limited to surface and atmospheric variables rather than the full coupled state.","tokens_in":17136,"feed_emoji":"🌍","tokens_out":9218,"duration_ms":80578,"temperature":0.7,"pith_summary":"The paper sets out to show that a flow-matching generative model can emulate a full coupled climate model cheaply enough to make decadal ensemble generation practical. ArchesClimate is trained on monthly outputs of IPSL-CM6A-LR decadal hindcasts, taking two prior months of atmosphere and ocean states plus greenhouse-gas and solar forcings to predict the next month; repeating that one-month step yields 10-year rollouts. The authors claim these rollouts remain stable and physically consistent, and that for several variables including net surface flux, omega, relative humidity and 700-hPa temperature, generated members are statistically interchangeable with the IPSL-DCPP ensemble members, with CRPS comparable to or better than the model's own members. If right, the main barrier to large ensembles for near-term climate prediction, their computational cost, is substantially lowered.","feed_headline":"Flow-matching AI emulator reproduces 10-year climate ensembles","feed_subtitle":"ArchesClimate turns Gaussian noise into stable monthly states that can be swapped into a climate model's own ensembles.","key_machinery":"The load-bearing mechanism is flow matching applied to residuals. A deterministic model forecasts the mean next monthly state, and a generative model learns a vector field that transports standard Gaussian noise to the distribution of the standardized difference between the true next state and that deterministic mean. At inference the generative model is integrated over several flow-matching steps and its output is added back onto the deterministic forecast; this composite state becomes the input for the next month, which is what makes the generation autoregressive. External forcings enter every transformer block through conditional layer normalization, and the ocean and atmosphere fields are handled together in one model rather than by separate emulators.","core_discovery":"The central claim is that a probabilistic emulator trained only on output states can stand in for a coupled climate model when the question is internal variability at monthly-to-decadal scales. Starting from two months of IPSL-DCPP state, ArchesClimate predicts the next month, feeds its own prediction back in, and stays stable and physically plausible for ten years. The paper reports that for several key variables, a generated member placed inside a ten-member IPSL-DCPP ensemble shows a flat rank histogram, meaning it could have come from the climate model itself, and that CRPS is close to or below the spread among the model's own members. The authors also report that the emulator responds to changing greenhouse-gas and solar forcings over a 50-year rollout, while acknowledging that its variance is consistently lower than IPSL-DCPP for every variable tested.","pith_inferences":["The paper's interpolated train/test split leaves open whether the emulator has learned forced dynamics or memorized seen decades; a decisive extension would be to condition on SSP scenarios outside the training period and check whether the 50-year trend response still tracks a full climate model.","The consistently lower variance suggests a structural test: initialize the flow-matching noise from a distribution matched to the observed spread of IPSL-DCPP anomalies rather than unit Gaussian; if variance still lags, the missing spread comes from information the white-noise residual cannot carry, such as ocean initial-condition memory.","A frequency-weighted spectral loss, rather than the uniform 0.2 scaling used in the paper, might recover the missing variance in sea-level pressure and ocean variables without the CRPS penalty the authors report.","Adding a slowly varying ocean memory term, such as a running decadal mean of ocean heat content, is a natural next condition that could restore variance and improve Arctic persistence without changing the architecture."],"forward_implications":["A trained ArchesClimate can generate a 10-member, 10-year ensemble at a fraction of the compute of running IPSL-CM6A-LR, which would make large-ensemble studies of internal variability much cheaper.","For variables with flat rank histograms, generated members can be substituted into an IPSL-DCPP ensemble without shifting its statistical behavior, so the emulator can augment existing decadal prediction ensembles.","Because CRPS for ArchesClimate is comparable to or lower than IPSL-DCPP for most tested variables, the model can act as a probabilistic forecast system at monthly-to-decadal lead times.","The 50-year forcing experiment shows the emulator tracks a changing external-forcing trend better than repeated fixed forcings, so it is not merely replaying the seasonal cycle."],"supporting_citations":[{"why":"Provides the base architecture and the deterministic-plus-generative training scheme that ArchesClimate adapts to monthly-to-decadal climate states.","marker":"Couairon et al. (2024)"},{"why":"Introduces flow matching, the generative method used to learn the distribution of the residual between the deterministic forecast and the true next state.","marker":"Lipman et al. (2023)"},{"why":"Supplies the DCPP hindcast protocol and the 10-member initialization design that defines the dataset used for training and evaluation.","marker":"Boer et al. (2016)"},{"why":"Presents IPSL-CM6A-LR, the coupled climate model whose decadal hindcasts ArchesClimate is trained to emulate.","marker":"Boucher et al. (2020)"},{"why":"Supplies the hierarchical vision-transformer backbone with shifted-window attention that the architecture builds on.","marker":"Bi et al. (2022)"},{"why":"Documents the sea-surface temperature and salinity nudging used to construct the DCPP initial-condition ensembles.","marker":"Estella-Perez et al. (2020)"},{"why":"Supplies the rectified-flow time-sampling and training recipe used for the flow-matching loss.","marker":"Esser et al. (2024)"},{"why":"Provides the CRPS implementation used to score the probabilistic forecasts against the IPSL-DCPP members.","marker":"Rasp et al. (2024)"},{"why":"Introduces rank histograms, the diagnostic used to test whether a generated member can be interchanged with an IPSL-DCPP ensemble member.","marker":"Hamill (2001)"},{"why":"Supports the premise that initial-condition memory fades quickly, which justifies generating ensemble spread from freshly sampled noise.","marker":"Smith et al. (2019)"}],"fun_headline_variants":["Flow-matching emulator matches climate model ensembles for 10 years","AI emulator reproduces climate model ensembles for a decade","Flow matching AI emulates decadal climate ensembles cheaply","Probabilistic emulator generates 10-year climate states","Decadal AI emulator matches IPSL-CM6A ensembles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the variation that makes an ensemble useful can be regenerated from freshly sampled Gaussian noise alone, even though the model never receives the specific initial-condition perturbations or nudging that produced each target member.","fun_headline_variants_meta":{"raw":{"variants":["Flow-matching emulator matches climate model ensembles for 10 years","AI emulator reproduces climate model ensembles for a decade","Flow matching AI emulates decadal climate ensembles cheaply","Probabilistic emulator generates 10-year climate states","Decadal AI emulator matches IPSL-CM6A ensembles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000851,"raw_usage":{"total_tokens":3693,"prompt_tokens":928,"completion_tokens":2765,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":2680}},"tokens_in":544,"tokens_out":2765,"duration_ms":18641,"temperature":1.0,"reasoning_tokens":2680,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:50:01.830862+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a 10-year ArchesClimate ensemble from the same 1969 initialization and compare the deep-ocean heat-content variable (thetaot2000) against IPSL-DCPP: if its ensemble variance stays near half the target and its rank histogram stays strongly u-shaped across all three test decades, even after the proposed loss changes, then the interchangeability claim is limited to surface and atmospheric variables rather than the full coupled state.","supporting_citations":[{"cited_title":", Smith, D M","cited_arxiv_id":null,"evidence_quote":"Supplies the DCPP hindcast protocol and the 10-member initialization design that defines the dataset used for training and evaluation."},{"cited_title":", Mignot, J","cited_arxiv_id":null,"evidence_quote":"Documents the sea-surface temperature and salinity nudging used to construct the DCPP initial-condition ensembles."},{"cited_title":"APACrefauthors \\ 2001 03","cited_arxiv_id":null,"evidence_quote":"Introduces rank histograms, the diagnostic used to test whether a generated member can be interchanged with an IPSL-DCPP ensemble member."},{"cited_title":", Eade, R","cited_arxiv_id":null,"evidence_quote":"Supports the premise that initial-condition memory fades quickly, which justifies generating ensemble spread from freshly sampled noise."}],"review_version":2}