{"id":"46770cea-66f8-4e54-b4c0-e62ec79ba9ed","arxiv_id":"2508.03845","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CIENS is a new benchmark dataset linking twelve years of German Weather Service ensemble forecasts to station observations, explicitly spanning multiple operational model upgrades.","lead":"This paper presents CIENS, a twelve-year dataset of operational ensemble weather forecasts from the German Weather Service, mapped to the locations of 170 synoptic stations and paired with observations. It gives statistical forecasters a long record that spans multiple model updates, so they can study how post-processing methods cope when the underlying forecast model changes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dataset's flagship non-stationarity use case requires per-forecast model-version metadata; the abstract never states that such metadata is included.","rationale":"The reader's verdict is CONDITIONAL with low confidence, based only on the abstract, and identifies the forecast-to-station mapping and observation homogeneity as the weakest assumption. I agree those matter, but the abstract's own distinguishing claim is the long period spanning multiple NWP model updates. That claim is load-bearing for the paper's stated purpose: if the dataset does not include per-forecast model-version metadata or a documented model-change timeline, then users cannot attribute skill changes to model updates, and the central novelty collapses. This is a concrete, checkable requirement rather than a general concern about data fidelity. The reader's weakest assumption is related but not identical, hence 'partial'. The recommended verdict remains CONDITIONAL because the concern can be settled by inspecting the dataset schema; if the metadata is present, acceptance is straightforward, and if absent, the paper's main advertised use case is unsupported. I do not move the verdict to REJECT because the abstract alone is insufficient to establish absence, and the dataset may well include such metadata in the full manuscript.","tokens_in":23467,"tokens_out":3195,"duration_ms":41149,"concrete_test":"Upon obtaining the dataset or its schema/metadata file, check whether each forecast record or run contains a model-version field (e.g., model_version, model_update, COSMO-DE vs ICON-D2) or whether a separate changelog maps run dates to DWD model versions. Then cross-reference at least the ICON-D2 introduction date (operational around 2021, replacing COSMO-DE) against the data: if the first ICON-D2-identifiable runs do not align with the known operational switchover date, the version attribution is unreliable. If no version field or changelog exists, the flagship non-stationarity application cannot be performed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's distinguishing claim is that the long record 'encompasses multiple updates to the underlying numerical weather prediction model' and thus supports investigations into how forecasting methods can account for model changes. That use case is only executable if every forecast carries a machine-readable model-version label, or if the dataset provides documented update timestamps aligned to DWD's operational change log. The abstract describes 55 variables, lead times, station mapping, and observations, but says nothing about model-version identifiers, update dates, or a run-to-model-version mapping. Without that metadata, a user cannot separate forecast-skill changes caused by model updates from changes caused by weather regime, so the central advertised capability is not actually supported by the dataset as described. This concern is independent of the interpolation and observation-quality issues: even a perfectly interpolated, perfectly homogeneous station record would still lack the version attribution needed for the stated use case. If version metadata is present in the full paper, the concern dissolves; if absent, the dataset's main novelty claim over existing benchmarks is weakened to a long but heterogeneous record of unknown provenance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CIENS, a station-matched dataset built from DWD's operational convection-permitting ensemble forecasts (COSMO/ICON), containing 55 forecast variables mapped to 170 German synoptic stations, hourly lead times from 0 to 21 hours for the 00 UTC and 12 UTC runs, and station observations for pressure, temperature, precipitation, wind speed, wind direction, and wind gusts, covering December 2010 to June 2023. The authors emphasize the dataset's long temporal extent across multiple NWP model updates as its key distinguishing feature, and they illustrate a machine-learning ensemble post-processing use case that benefits from the rich predictor set.","tokens_in":23521,"tokens_out":3536,"duration_ms":42978,"significance":"If the dataset is as described and its construction is fully documented, CIENS would be a valuable community resource: it provides a long, station-matched, multi-variable ensemble forecast record spanning several operational model generations, which is well suited for post-processing, verification, and studies of forecast non-stationarity. The included machine-learning use case is a constructive demonstration of the dataset's potential. However, the provided manuscript text is severely corrupted in the version I received, and the abstract alone omits several load-bearing details, so the dataset's significance cannot currently be fully assessed from the materials available to me.","major_comments":[{"comment":"The paper's flagship claim is that the long temporal extent 'encompasses multiple updates to the underlying numerical weather prediction model' and supports investigations into how forecasting methods account for such changes. This use case requires a machine-readable mapping from each forecast run to the operational model version, or at least a documented update chronology aligned with DWD's operational change log. The abstract enumerates 55 variables, 170 stations, lead times, and two daily runs, but it does not mention model-version identifiers or update timestamps anywhere in the dataset description. Please state explicitly whether such metadata is included; without it, the central advertised capability is not actually supported by the dataset as described.","section":"Abstract, distinguishing feature"},{"comment":"The abstract states that forecasts are 'mapped to the locations of synoptic stations' and that additional spatially aggregated forecasts from surrounding grid points are available for a subset of variables, but it does not specify the interpolation scheme, the treatment of station elevation versus model orography, or the definition of 'surrounding grid points'. For convection-permitting output, the grid-to-point mapping strongly affects apparent forecast errors and any downstream post-processing benchmark. The corrupted full text I received does not allow me to locate such a description; please add a precise specification of the mapping and, ideally, an estimate of its representativeness uncertainty.","section":"Abstract, station mapping"},{"comment":"The dataset supplies twelve years of station observations for six variables at 170 locations, but the abstract provides no information about quality control, homogeneity adjustments, station relocations, instrument changes, or missing-data flags. Any inhomogeneity in the observational record will be inherited by every post-processing or verification study built on CIENS. Please include a dedicated observation-processing section, with station metadata and per-variable quality flags, so that users can assess and account for observational inhomogeneities.","section":"Observations and quality control"}],"minor_comments":[{"comment":"There is a typo in 'Since the forecast are mapped to the observed locations'; it should read 'Since the forecasts are mapped'.","section":"Abstract"},{"comment":"The manuscript text in the version provided to me is in an unreadable character encoding (mojibake), so I could not verify the section structure, equations, tables, or any statements beyond the abstract. Please ensure the arXiv source compiles correctly and contains readable text.","section":"Full text"},{"comment":"The abstract does not state where the CIENS dataset can be accessed, under what license, or in which file format. A dataset paper should include a persistent identifier and a clear data-availability statement.","section":"Data availability"},{"comment":"Please specify exactly which operational models and configurations are included across the 2010-2023 period (e.g., COSMO-DE versus ICON-D2 ensemble), including ensemble size and perturbation strategy, since the period straddles a major model transition.","section":"Model configuration"}],"recommendation":"major_revision","confidential_remarks":"The body text I received is entirely corrupted, so my review rests on the abstract, the reader's report, and the stress-test note. If the authors' actual submission contains the full details, several of my major comments may already be addressed; the editors should ask the authors to confirm specifically whether model-version metadata is included, and to point to the interpolation and observation-quality-control sections. As it stands, the abstract's flagship non-stationarity use case is not verifiable from the described dataset contents, and the grid-to-station mapping and observation homogeneity are load-bearing for all downstream applications. I am not recommending rejection because these are fixable documentation issues rather than fundamental errors in the dataset concept."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a dataset paper that should go to peer review, but the flagship claim about enabling studies of model changes hangs on a detail the abstract never states — per-forecast model-version metadata — and we only had the abstract to read, since the supplied full text is corrupted (mojibake, with a different arXiv ID embedded). So your confidence level is right.\n\nWhat's actually new: CIENS is a 12.5-year record of operational convection-permitting ensemble forecasts from DWD's COSMO/ICON, mapped to 170 synoptic stations, with 55 forecast variables, hourly lead times 0-21 for the 00 and 12 UTC runs, plus observations for six variables. That is a real resource for the statistical post-processing and verification community, which mostly works with shorter or non-stationary benchmark sets. The abstract is specific and internally consistent, and the use case showing how richer predictors help ML post-processing is a sensible demonstration, not an inflated prediction claim.\n\nThe soft spots are proportionate. The biggest one is the stress-test concern: the abstract's distinguishing feature is that the record spans multiple NWP model updates and 'thus supports investigations into how forecasting methods can account for such changes.' That use case only works if each forecast carries a machine-readable model-version label or the dataset provides documented update timestamps aligned with DWD's change log. The abstract lists variables, stations, lead times, and observations but says nothing about version IDs or update dates. Without that, a user cannot separate model-update effects from weather-regime effects, and the headline novelty weakens to a long but heterogeneous record of unknown provenance. This is fixable if the full paper includes the metadata; a referee needs to check.\n\nThe other soft spots are the usual dataset-paper questions: the grid-to-station interpolation isn't described in the abstract, observation quality control is not mentioned, and there's no URL or DOI. For a resource that others will build on, those need explicit answers. Also, the supplied full text was unreadable, so I could not verify the internals; the abstract alone supports a conditional judgment, not a confident accept.\n\nWho it's for: anyone doing ensemble post-processing, forecast verification, or benchmark ML for weather. It deserves a serious referee. My recommendation: send it to review, with a referee specifically asked to verify the version metadata, the station-mapping methodology, and data accessibility. If the version metadata is missing, the authors should add it or soften the non-stationarity claim.","headline":"Worth reviewing as a dataset contribution, but the abstract's flagship non-stationarity use case hinges on model-version metadata that is nowhere mentioned.","tokens_in":24154,"tokens_out":2546,"would_cite":true,"duration_ms":28433,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CIENS dataset pairs more than twelve years of operational convection-permitting ensemble forecasts with station observations at 170 German sites.","keywords":["ensemble weather forecasts","convection-permitting NWP","COSMO","ICON","ensemble post-processing","forecast verification","station observations","benchmark dataset"],"falsifier":"A single calculation could settle the central premise: compare CIENS station-mapped forecasts with raw nearest-grid-point model output for a sample of stations and lead times, and check whether residuals jump at known station relocation or instrument-change dates.","tokens_in":23161,"feed_emoji":"🌦️","tokens_out":5536,"duration_ms":68898,"temperature":0.7,"pith_summary":"The paper presents CIENS, a new dataset built from the German national weather service's operational convection-permitting ensemble forecasts, spanning December 2010 to June 2023. Forecasts for 55 meteorological variables are mapped to the locations of 170 synoptic stations, with additional spatially aggregated forecasts from surrounding grid points for a subset of variables. Hourly lead times from 0 to 21 hours are included for the 00 and 12 UTC model runs, alongside station observations of pressure, temperature, hourly precipitation, wind speed, wind direction, and wind gusts. The central claim is that the dataset's long temporal extent, which covers multiple updates to the underlying numerical weather prediction model, makes it a realistic benchmark for studying how ensemble post-processing and verification methods cope with operational model changes.","feed_headline":"12 years of German ensemble forecasts, station by station","feed_subtitle":"CIENS maps 55 operational model variables to 170 sites so forecast users can train methods on real model changes.","key_machinery":"The carrying object is the CIENS station-level forecast-observation pairing: 55 model variables are mapped or spatially aggregated onto 170 station locations and aligned with hourly lead times from 0 to 21 hours for both daily runs, with six observed variables for verification. This machinery makes the model output directly usable for post-processing and verification without grid handling, and the long record across model updates turns the archive into a test bed for changing forecast-error distributions.","core_discovery":"The central claim is that CIENS is a benchmark dataset in which operational convection-permitting ensemble forecasts, rather than reanalysis or frozen model runs, are mapped to 170 synoptic stations over more than twelve years, with six observed variables attached for direct verification. Because the archive spans several upgrades of the underlying numerical weather prediction model, it exposes the non-stationarity that real forecast post-processing must handle instead of hiding it behind a fixed model version. A use case on ensemble post-processing illustrates that machine-learning forecasting models benefit from the rich set of available model predictors.","pith_inferences":["If CIENS becomes a standard benchmark, model-update boundaries offer natural test splits: methods that degrade around those boundaries can be identified and replaced, turning operational model churn into a research asset.","The station-mapping step may introduce small-scale representativeness errors, especially at short lead times, so users studying kilometer-scale processes should compare the mapped values with raw grid output.","The long record could also be used to estimate how station relocations and observation quality-control changes affect apparent forecast skill, a confound the paper does not itself quantify.","Similar station-mapped archives for other national weather services would allow cross-country comparisons of post-processing transferability and model-change robustness."],"forward_implications":["Users can train and evaluate ensemble post-processing methods directly on forecast-observation pairs without having to handle model grid geometry.","The multi-year, multi-model-version span lets researchers test whether post-processing methods remain accurate when the underlying numerical model changes.","The 55-variable predictor set plus spatial aggregates enables machine-learning post-processing to use far more information than standard ensemble moments.","The paired observations support verification of probabilistic forecasts at the point scale for six standard variables.","Because the data format matches operational forecast use, results obtained on CIENS should transfer more readily to real forecasting practice than idealized benchmarks."],"supporting_citations":[],"fun_headline_variants":["CIENS: 12 years of operational ensemble forecasts at 170 stations","Operational ensemble forecasts spanning 12 years of model upgrades","Forecast archive reveals model-change effects at 170 German stations","12 years of station-level ensemble forecasts that span model updates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's worth rests on the station-mapped forecasts and paired observations staying consistent and unbiased over twelve years, since any interpolation bias or break in the observation record is inherited by every downstream study.","fun_headline_variants_meta":{"raw":{"variants":["CIENS: 12 years of operational ensemble forecasts at 170 stations","Operational ensemble forecasts spanning 12 years of model upgrades","Forecast archive reveals model-change effects at 170 German stations","12 years of station-level ensemble forecasts that span model updates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001172,"raw_usage":{"total_tokens":4827,"prompt_tokens":904,"completion_tokens":3923,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":3852}},"tokens_in":520,"tokens_out":3923,"duration_ms":31448,"temperature":1.0,"reasoning_tokens":3852,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:12:49.438330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single calculation could settle the central premise: compare CIENS station-mapped forecasts with raw nearest-grid-point model output for a sample of stations and lead times, and check whether residuals jump at known station relocation or instrument-change dates.","supporting_citations":[],"review_version":1}