{"id":"b0230328-b566-4ba1-bcd9-c7009c4a86b6","arxiv_id":"2605.29976","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ArchesWeather and ArchesWeatherGen, when conditioned on monthly SST and SIC, produce stable multi-decadal climate simulations that reproduce ERA5 climatology, large-scale circulations, interannual variability, and distribution tails.","lead":"This paper adapts two machine learning weather forecasting models by adding monthly sea surface temperature and sea ice conditioning to run multi-decadal climate simulations following the AIMIP protocol. If the stability and fidelity claims hold, it points toward lower-cost alternatives to traditional physics-based climate models for producing long-term projections.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's limitation (abstract-only access) is acknowledged and no further load-bearing flaw is visible. The evaluation protocol described matches standard practice for testing exactly the conditioning assumption, so the UNVERDICTED status is appropriate pending full-text review.","tokens_in":1787,"tokens_out":274,"duration_ms":15375,"concrete_test":"Re-run the multi-decadal forced simulation with an additional diagnostic for global energy balance drift (TOA radiation imbalance integrated over 30+ years) and compare against the unforced baseline; if the forced configuration shows statistically significant secular drift not present in the reference AMIP runs, the stability claim requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that monthly-mean SST/SIC conditioning converts the weather models into stable forced atmospheric models under the AIMIP protocol. The abstract states that forced runs are stable, reproduce ERA5 climatology/circulations/interannual variability, and capture distribution tails, with ablations and forced/unforced comparisons presented. No internal inconsistency, hidden assumption in the conditioning mechanism, or unsupported quantitative claim is detectable from the given material. The reader's weakest_assumption correctly isolates the key condition that would have to hold; the paper positions its evaluation as a direct test of that condition.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript evaluates the adaptation of ArchesWeather (deterministic) and ArchesWeatherGen (probabilistic flow-matching) weather-forecasting models to multi-decadal climate simulation. By adding conditioning on monthly-mean SST and SIC as boundary conditions and following the AIMIP Phase 1 protocol, the authors claim that the forced configurations produce stable long-term runs with a stable annual cycle, capture drifts in many climate variables, faithfully reproduce ERA5 climatology, large-scale circulations, interannual variability and distribution tails, and outperform or match numerical climate models in key respects; the evaluation includes forced/unforced comparisons and ablation studies on design choices.","tokens_in":1897,"tokens_out":536,"duration_ms":25343,"significance":"If the quantitative results hold, the work would establish that weather-trained ML models can be repurposed for stable forced atmospheric climate simulation with only monthly SST/SIC conditioning, offering a potentially efficient route to ensemble climate runs and uncertainty quantification that complements traditional GCMs.","major_comments":[{"comment":"The central claim that monthly-mean SST/SIC conditioning alone suffices for drift-free multi-decadal stability (abstract and AIMIP protocol description) is load-bearing, yet the manuscript provides no explicit quantification of drift rates (e.g., linear trends in global-mean temperature or zonal wind with confidence intervals) or a clear statement of the exact conditioning implementation (how the monthly fields are upsampled and injected into the network).","section":"Methods / AIMIP protocol section"},{"comment":"Ablation results on the conditioning and forced vs. unforced configurations are referenced but lack tabulated metrics (e.g., RMSE or bias for key variables across periods) with error bars or statistical significance tests, making it impossible to judge whether the reported stability and ERA5 fidelity are robust or sensitive to post-hoc choices.","section":"Ablation studies and results sections"}],"minor_comments":[{"comment":"Figure captions and axis labels should explicitly state the time period, ensemble size, and variable units for all climatology and variability plots to allow direct comparison with ERA5 and CMIP-style benchmarks.","section":"Figures"},{"comment":"The abstract states that the models 'capture the tails of the distributions' but does not define the tail metric (e.g., 99th percentile bias or extreme-event frequency); this should be clarified with a precise definition and reference to the corresponding figure or table.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments and the recommendation for major revision. We will update the manuscript to provide the requested quantifications and clarifications, as detailed in our point-by-point responses below.","responses":[{"response":"We agree that providing explicit quantification of drift rates would strengthen the central claim. In the revised manuscript, we will add analyses of linear trends in global-mean temperature, zonal wind, and other key variables, including confidence intervals. We will also clarify the conditioning implementation in the Methods section by detailing how monthly SST and SIC fields are upsampled (e.g., via interpolation) and injected into the network architecture. These details were part of our experimental setup but not fully elaborated.","revision_made":"yes","referee_comment":"[Methods / AIMIP protocol section] The central claim that monthly-mean SST/SIC conditioning alone suffices for drift-free multi-decadal stability (abstract and AIMIP protocol description) is load-bearing, yet the manuscript provides no explicit quantification of drift rates (e.g., linear trends in global-mean temperature or zonal wind with confidence intervals) or a clear statement of the exact conditioning implementation (how the monthly fields are upsampled and injected into the network)."},{"response":"We acknowledge the need for more rigorous presentation of the ablation results. We will include tabulated metrics in the revised version, reporting RMSE, biases, and other relevant statistics for key variables across different time periods. These tables will incorporate error bars (e.g., standard deviations across ensemble members or runs) and statistical significance tests (e.g., t-tests or similar) to assess robustness. This will allow better evaluation of the sensitivity to design choices.","revision_made":"yes","referee_comment":"[Ablation studies and results sections] Ablation results on the conditioning and forced vs. unforced configurations are referenced but lack tabulated metrics (e.g., RMSE or bias for key variables across periods) with error bars or statistical significance tests, making it impossible to judge whether the reported stability and ERA5 fidelity are robust or sensitive to post-hoc choices."}],"tokens_in":1444,"tokens_out":450,"duration_ms":18884,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The key point is that ArchesWeather and ArchesWeatherGen, when given monthly SST and SIC as extra inputs, run stable multi-decadal simulations that match ERA5 climatology and variability reasonably well.\n\nThe new part is applying these specific models to the AIMIP Phase 1 protocol for forced atmospheric models. They test both the deterministic and the flow-matching generative version, include ablations on design choices, and contrast forced against unforced runs. This shows the models maintain a stable annual cycle and pick up drifts in climate variables without the instability one might expect from weather-only training.\n\nThey do well on the evaluation side by checking large-scale circulations, interannual variability, and distribution tails against reanalysis and numerical models. The protocol following AIMIP makes the results comparable to other efforts in the area.\n\nThe soft spots are in the level of detail. The claims about faithful reproduction and capturing tails are presented without the actual error values or plots in the summary, so it's difficult to judge how close is close enough or if there are systematic biases in certain variables or regions. The assumption that monthly mean boundary conditions suffice for physical consistency over decades could use more testing, like checking for energy balance or other conservation properties. If the full paper has those numbers and they hold, the result strengthens; if not, the stability might be overstated.\n\nThis paper is aimed at the intersection of ML weather forecasting and climate modeling. Someone working on adapting ML models for longer timescales or participating in AIMIP-style intercomparisons would find the forced/unforced analysis and the generative model extension useful.\n\nIt is worth sending to peer review. The work is grounded in an existing protocol and addresses a practical question about model repurposing. Referees can push for the missing quantitative details and any code or data availability.","headline":"These weather models stay stable over decades with monthly SST/SIC forcing and reproduce ERA5 features under AIMIP, but the strength of that match rests on details not visible in the summary.","tokens_in":2415,"tokens_out":448,"would_cite":false,"duration_ms":20702,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Conditioning weather-trained ML models on monthly SST and SIC produces stable multi-decadal climate simulations matching ERA5.","keywords":["machine learning","climate simulation","ArchesWeather","SST conditioning","AIMIP","ERA5 evaluation","atmospheric modeling","multi-decadal stability"],"falsifier":"A multi-decadal AIMIP run that shows growing unbounded drift in global temperature, circulation strength, or annual cycle amplitude beyond ERA5 levels would falsify the stability claim.","tokens_in":2698,"feed_emoji":"🌍","tokens_out":449,"duration_ms":18859,"temperature":0.7,"pith_summary":"The paper tests whether two machine learning models originally built for short-term weather prediction can be turned into usable tools for long-term climate work. It adds monthly mean sea surface temperature and sea ice cover as boundary conditions and runs them following the AIMIP experimental protocol. The central result is that the adapted models stay stable over decades, keep a consistent annual cycle, and reproduce the observed climatology, large-scale flows, year-to-year changes, and extremes from ERA5 reanalysis. A reader would care because this shows a low-cost way to repurpose existing weather models for climate-length runs instead of building new ones from scratch.","feed_headline":"Weather ML models produce stable multi-decadal climate runs","feed_subtitle":"Monthly SST and SIC inputs let ArchesWeather match ERA5 climatology, circulations and variability without drift.","key_machinery":"Monthly mean SST and SIC conditioning as boundary conditions to adapt weather models into forced atmospheric climate models under the AIMIP protocol.","core_discovery":"Despite being originally developed for weather forecasting, forced configurations of ArchesWeather and ArchesWeatherGen produce stable long-term climate simulations, have a stable annual cycle, and capture the drift of many climate variables. The models faithfully reproduce ERA5's climatology, large-scale circulations and interannual variability, and they capture the tails of the distributions.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["ArchesWeather runs stable multi-decadal climate simulations","ML models from weather forecasting yield stable climate outputs","ArchesWeather and ArchesWeatherGen match ERA5 under forced conditions","Stable long-term climate sims from weather ML with SST inputs","Weather models reproduce climatology in multi-decadal runs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Monthly mean SST and SIC conditioning alone is enough to turn weather models into stable forced climate models for multi-decadal runs.","fun_headline_variants_meta":{"raw":{"variants":["ArchesWeather runs stable multi-decadal climate simulations","ML models from weather forecasting yield stable climate outputs","ArchesWeather and ArchesWeatherGen match ERA5 under forced conditions","Stable long-term climate sims from weather ML with SST inputs","Weather models reproduce climatology in multi-decadal runs"]},"model":"grok-4.3","cost_usd":0.00465,"raw_usage":{"total_tokens":2314,"prompt_tokens":693,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":46499500,"prompt_tokens_details":{"text_tokens":693,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1542,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":693,"tokens_out":79,"duration_ms":10175,"temperature":1.0,"reasoning_tokens":1542,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T23:45:47.915830+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A multi-decadal AIMIP run that shows growing unbounded drift in global temperature, circulation strength, or annual cycle amplitude beyond ERA5 levels would falsify the stability claim.","supporting_citations":[],"review_version":1}