{"id":"b7e5a2d8-b382-4a2d-b564-aefa4e2f29d8","arxiv_id":"2501.15690","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A probabilistic mixture-of-experts weighting of 13 CORDEX regional models improves High Mountain Asia precipitation estimates over equal averaging and projects wetter summers, drier winters in the western Himalayas and Karakoram, and wetter winters over the Tibetan Plateau.","lead":"This paper combines 13 regional climate model outputs over High Mountain Asia with a probabilistic weighting scheme trained on historical precipitation data, yielding more accurate historical precipitation distributions than equal model averaging. It then produces refined near- and far-future precipitation projections under RCP4.5 and RCP8.5 for a region that supplies water to more than 1.9 billion people.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The untested stationarity of historical-skill weights is the load-bearing assumption for the future projections; a temporal cross-validation split would resolve it.","rationale":"The reader's weakest assumption correctly identifies that the historical-skill weights are assumed to transfer to future climate states without any stationarity test. I agree with that assessment, and it is the most load-bearing concern for the paper's central future-projection claim. The held-out CRPS validation (1981–2005) is a genuine strength and supports the historical aggregation improvement claimed in the abstract; the 32% figure is based on that validation and is not undermined by my concern. However, the future projections are not validated by that result. The MoE weights are trained on the full historical period and then applied unchanged to RCP scenarios, so the projected spatial patterns and MoE-versus-EW differences depend entirely on the assumption that the relative skill of the 13 RCMs, as measured by Wasserstein distance to APHRODITE, is stationary under forcing. The paper itself does not test this, and the literature contains well-known examples of historical skill not guaranteeing projected-response fidelity. My proposed temporal cross-validation is a direct, feasible check: APHRODITE extends to 2015, so an early-period training set and a later-period validation set are available, and the same code release on Zenodo can be rerun with the split. This does not change the reader's verdict of CONDITIONAL; it sharpens the condition. I would not move to REJECT because the historical result is solid and reproducible artifacts are provided, but the future-refinement language in the abstract and conclusion should be presented as conditional on weight stationarity until such a test is run.","tokens_in":23150,"tokens_out":3709,"duration_ms":37884,"concrete_test":"Split the historical record into two disjoint periods: optimise the weights and Wasserstein distances on 1951–1980 only, then (1) compute CRPS and the MoE-minus-EW maps on the untouched 1981–2005 period, and (2) regenerate the RCP8.5 far-future relative-change maps using the early-period weights. If the 1981–2005 CRPS degradation relative to full-period-trained weights is within Monte Carlo sampling error and the Fig. 7 maps differ by less than the roughly ±10 percentage-point band used in Fig. 8 over most HMA cells, the stationarity assumption is supported; if the validation skill drops substantially or the projection maps shift by more than that band, the future claims are conditional on an unsupported assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that weights learned from historical agreement with APHRODITE remain the correct combination weights under changed forcing. Section 5 states that the weights are optimised over the entire historical period (1951–2005) and then applied unchanged to 2036–2095 under RCP4.5 and RCP8.5. The only validation of the MoE is CRPS on the held-out period 1981–2005; that validates historical interpolation, not transfer to a different climate state. The weighting in Eq. (11) is a softmax over Wasserstein distances between RCM surrogate distributions and APHRODITE, so it is a historical-skill weighting; nothing in the method conditions on the forced response. A model that is close to APHRODITE historically can nevertheless have a wrong response to greenhouse-gas and aerosol forcing, so the claimed refined future signals (wetter summers/drier winters in HMA1, wetter winters over the Plateau/Hengduan/Southeast Tibet, and the MoE–EW differences in Fig. 8) all inherit this untested assumption. Additionally, Fig. 6's historical REMP improvement is partly in-sample because the historical reference maps use the same 1951–2005 period used to fit the weights; this does not invalidate the held-out CRPS result but does weaken the refined-climatology framing for the historical period.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a probabilistic mixture-of-experts (MoE) framework for combining monthly precipitation from 13 CORDEX-WAS regional climate models over High Mountain Asia. Each RCM is represented by a warped Gaussian process surrogate, aggregated via Bayesian committee machines, and the RCM surrogates are weighted by a softmax function of Wasserstein distances to the APHRODITE gridded observations, with a learned 'statistical temperature'. The method is validated on a held-out period (1981-2005) after training on 1951-1980, reporting a 31% average CRPS improvement over equal weighting and a 254% improvement over any single ensemble member. The authors then re-optimize the weights over the full historical period (1951-2005) and apply them unchanged to RCP4.5 and RCP8.5 projections for 2036-2065 and 2066-2095, producing seasonal median precipitation changes that differ regionally from the equally weighted baseline, including wetter summers and drier winters over the western Himalayas and Karakoram and wetter winters over the Tibetan Plateau, Hengduan Shan, and Southeast Tibet.","tokens_in":23452,"tokens_out":5035,"duration_ms":48079,"significance":"If the held-out validation is accepted, the paper offers a broadly applicable and reproducible method for combining RCM ensembles that improves upon both equal weighting and single-model selection without discarding outliers. The temporal train/validation split is an honest design, and the public release of code and data on Zenodo is a clear strength. The central historical claim is therefore credible. However, the future projection section rests on the untested assumption that weights learned from historical distributional fidelity transfer to a different climate state, and the headline improvements are reported without uncertainty quantification. These issues affect the paper's central claims rather than merely its presentation.","major_comments":[{"comment":"The future projections assume that weights optimized for historical fidelity to APHRODITE remain optimal when applied unchanged to 2036-2095 RCP4.5/RCP8.5 simulations. Equation (11) defines the weights as a softmax over negative log Wasserstein distances between each RCM surrogate and APHRODITE, so the weights measure historical distributional agreement only; nothing conditions on the forced response of the models. The headline future signals (wetter summers/drier winters over HMA1 and wetter winters over HMA3, Fig. 8) therefore presuppose that a model closer to APHRODITE historically is also a better predictor of the change signal, which is testable but untested. I recommend an explicit sensitivity analysis, e.g., estimating weights on two disjoint historical periods and comparing the resulting projection maps, or a perfect-model experiment in which one RCM's future is withheld as a pseudo-observation. As written, the projection claims are not falsifiable within the manuscript.","section":"§5.2, Eq. (11)"},{"comment":"The headline improvements ('31% average improvement over EW' in the text, '32%' in the abstract; '254% improvement' over any single member) are reported as point estimates without confidence intervals or significance tests. Figure 5 shows large spatial heterogeneity, and Figure D2 indicates near-zero or positive differences in some months and locations. I request bootstrap confidence intervals over years or spatial blocks for the aggregate CRPS improvements, plus a statement of the fraction of grid points with statistically significant improvement. Without this, the reader cannot assess whether the 31% and 254% figures are distinguishable from noise.","section":"§4"},{"comment":"The historical REMP evaluation in Figure 6 is in-sample: the weights are optimized over the entire historical period (1951-2005) and then evaluated against APHRODITE for 1976-2005, a sub-period of the training window. The text should explicitly label these maps as diagnostics rather than validation, and the conclusion that 'the MoE makes large improvements over the EW' in Section 5.1 should be supported by the held-out CRPS result of Section 4, not by Figure 6. This does not invalidate the held-out validation, but the framing currently blurs the distinction.","section":"§5.1, Fig. 6"},{"comment":"The paper does not state which period is used to compute the 95th-percentile scaling of BCM outputs to APHRODITE before calculating Wasserstein distances. If this scaling is estimated on the same 1981-2005 validation period, the held-out CRPS partly measures the method's ability to match a bias-correction target defined from the validation data, not purely predictive skill. Please specify the period used for the scaling; if it includes the validation period, re-run the validation with the scaling estimated only from 1951-1980.","section":"§3.2.2, §4"}],"minor_comments":[{"comment":"The abstract states a '32% improvement' while Section 4 reports a '31% average improvement' over the EW; please reconcile these numbers and state explicitly that the validation is on the held-out period 1981-2005.","section":"Abstract and §4"},{"comment":"The second limitation paragraph contains a typo: 'uncertain..' should be 'uncertain.'.","section":"§6"},{"comment":"The caption for Figure F2 says 'near future (2066-2095)', but the near-future period is 2036-2065; please correct.","section":"Appendix F, Fig. F2 caption"},{"comment":"In the Snelson et al. (2003) reference, 'Vacouver' should be 'Vancouver'.","section":"References"},{"comment":"The third row label 'MOE EW REMP' has inconsistent capitalization; use 'MoE - EW' for clarity.","section":"Fig. 6 caption"},{"comment":"Please define the '254% improvement' mathematically (e.g., as the relative reduction in CRPS compared with the best individual RCM) so that the claim is unambiguous.","section":"§4"},{"comment":"With h(·) = ln(·), the weights in Eq. (11) become a power-law function of the Wasserstein distance; a sentence noting this functional form and its smoothness motivation would help the reader interpret the temperature parameter.","section":"Eq. (11)"}],"recommendation":"major_revision","confidential_remarks":"The paper's held-out validation design and open code/data are genuine strengths, and the historical improvement claim is likely sound. The main risk is that the abstract and conclusion present the future projection differences as established results, while the stationarity of the learned weights is untested. A revision that adds a stationarity test or explicit sensitivity analysis, quantifies uncertainty for the headline CRPS improvements, and clarifies the in-sample nature of the historical REMP maps would substantially strengthen the paper. I would not reject: the central methodological contribution is valuable and the concerns are addressable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious referee. It does something genuinely new: it builds a full-distribution mixture-of-experts for 13 CORDEX-WAS RCMs over High Mountain Asia, weighting each RCM by Wasserstein distance to APHRODITE with a learned statistical temperature, and it validates the aggregation on a held-out period (train 1951–1980, test 1981–2005). The headline numbers—31% CRPS improvement over equal weighting and 254% over the best single member—are supported by an honest experimental design, and the code and data are on Zenodo. That is real evidence and should count as such.\n\nThe soft spots are mostly concentrated in the future projection section. The weights are optimized over the entire historical period (1951–2005) and then applied unchanged to 2036–2095 under RCP4.5 and RCP8.5. Nothing in the method tests whether model skill relative to APHRODITE is stationary across climate states. That is the load-bearing assumption behind the refined future signals, and the paper does not flag it as strongly as it should. A temporal cross-validation split—say, refitting weights on one historical sub-period and applying to another with different forcing—would help, though it cannot fully resolve the transfer problem. Second, the 31% and 254% improvements are reported without uncertainty bars; the reader cannot tell whether the gains are statistically distinguishable from zero at regional or seasonal scales. Third, the historical REMP comparison in Fig. 6 is partly in-sample because the same full historical period is used to fit the weights and to evaluate against APHRODITE. That weakens the historical framing but does not invalidate the held-out CRPS result.\n\nI agree with the paper's own stated limitations: the GP surrogates assume unimodality, the weighting function is somewhat arbitrary, and the 95th-percentile scaling is an observation-tuned step. These are minor relative to the central contribution. The literature engagement is fair; the citation pattern looks appropriate, and the comparison to Sanjay et al. is a sensible anchor.\n\nWho is this for? Anyone working on multi-model ensemble aggregation, particularly for precipitation or regional climate downscaling. It deserves peer review. My main request would be to make the stationarity assumption explicit and add uncertainty quantification to the headline skill statistics. The historical claim is solid; the future claim needs a caveat.\n\nFor the record: I would cite this for the method and would bring it to a reading group, but I would not yet use the future projections for planning without an independent test of weight transfer.","headline":"A solid methodological contribution with honest validation, but the future projections rest on an untested stationarity assumption that should be flagged.","tokens_in":23981,"tokens_out":1393,"would_cite":true,"duration_ms":14201,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A mixture-of-experts blend of 13 climate models beats equal-weight averaging by 32% and the best single model by 254%, making High Mountain Asia's projected summers wetter and its western Himalayan winters drier.","keywords":["High Mountain Asia","precipitation","ensemble learning","regional climate models","probabilistic","Gaussian process","CORDEX-WAS"],"falsifier":"Re-fit the mixture weights on two non-overlapping historical windows, such as 1951–1970 and 1971–1990, and compare the resulting weight maps: if the optimal weights shift substantially between windows, or if weights fit on one window fail to beat equal weighting on the other, the fixed-weight future projections rest on an unsupported assumption. A complementary check is to score the same learned weights against a second gridded precipitation product not used in training, since the reported 32% gain is measured only against APHRODITE.","tokens_in":22958,"feed_emoji":"🏔️","tokens_out":21221,"duration_ms":159711,"temperature":0.7,"pith_summary":"This paper tries to establish that precipitation over High Mountain Asia is better estimated by a learned blend of regional climate models than by the standard equal-weight average. The blend is a mixture of experts: each of 13 CORDEX models is converted into a probabilistic surrogate, and the surrogates are weighted at every location and month by how closely each model's precipitation distribution matches the APHRODITE gridded observations, with a statistical temperature tuned on historical data. On held-out years the blend is 32% closer to observations than the equal-weight average and 254% closer than the single best model. Applied to RCP4.5 and RCP8.5, the refined blend projects wetter summers but drier winters over the western Himalayas and Karakoram and wetter winters over the Tibetan Plateau, Hengduan Shan, and South East Tibet, compared with previous estimates. The stakes are concrete: precipitation is the largest source of uncertainty in modelling the water resources of a region whose frozen stores feed more than 1.9 billion people.","feed_headline":"Blending 13 climate models cuts Asia precipitation error 32 percent","feed_subtitle":"The blend also beats the best single model by 254 percent, reshaping High Mountain Asia's seasonal precipitation outlook.","key_machinery":"The carrying object is the mixture of experts (MoE): at each month and location a convex combination $p(y)=\\sum_{r=1}^{R} w_r\\,p_r(y)$ of 13 surrogate distributions. Each surrogate $p_r$ is a Gaussian process fitted to one RCM's monthly precipitation and then transformed back to skew: a Box-Cox warp makes the GP's Gaussian output match the log-normal-like distribution of precipitation, chained Gaussian processes let the variance as well as the mean be learned spatial fields, and a Bayesian committee machine splits the data into per-month domains so training cost scales as $O(JM^3)$ instead of $O(N^3)$. The weight formula carries the argument: $w_r \\propto \\exp(-\\ln W_r / T)$, where $W_r$ is the Wasserstein distance between the $r$-th surrogate and APHRODITE at that location and month, and $T$ is a statistical temperature fitted by maximum likelihood on the 1951–2005 training period; $T=0$ would pick only the closest model, $T\\to\\infty$ reproduces equal weighting, and the fitted temperature keeps the weights smooth across space and months while never letting any model's weight vanish.","core_discovery":"The paper's central claim is that a mixture of experts, a weighted sum of 13 regional climate model precipitation distributions, produces more faithful precipitation distributions over High Mountain Asia than equally weighted averaging or any single member of the CORDEX-WAS ensemble. Weights are computed from the Wasserstein distance between each model's surrogate distribution and the APHRODITE historical dataset, with a likelihood-optimised statistical temperature that interpolates between best-model selection and equal weighting; poorly performing models are down-weighted, never discarded. On the held-out period 1981–2005 the mixture improves the continuous rank probability score by 32% on average over equal weighting (the validation section reports 31%), and by 254% over the best individual surrogate, with the largest gains over the Karakoram and Himalayan arc where member skill differs most and near-zero gains where the learned weights revert to near-equal values. When those weights are applied unchanged to future scenarios, the mixture projects summer monsoon median precipitation increases of roughly 15% to 27% by 2066–2095 under RCP8.5 across the three study regions, and winter changes that are drier over the inner Himalayas and Karakoram (median differences down to −31% relative to the equal-weight estimate for the West Himalaya and Karakoram) but wetter over the Tibetan Plateau and Southeast Tibet (up to +62%).","pith_inferences":["Because the mixture is a convex combination of the 13 member distributions, it cannot produce precipitation values its members never span; the heavy tail of the blended distribution is inherited, not generated, so updating flood and drought return levels would require an explicit extremes model on top of the mixture.","The paper fixes the weights once on 1951–2005 and applies them to 2036–2095 without checking that model skill rankings are stable; recomputing the optimal weights on separate historical decades would show whether the fixed weights are a liability for projections.","Training the weights on APHRODITE and validating the resulting blend against an independent observational product, or against station gauges, would separate genuine ensemble skill from adherence to one reference dataset's particular biases.","If the projected seasonal shift is correct, wetter summers and drier winters in the western Himalayas and Karakoram would alter glacier accumulation-season forcing and the timing of downstream river flow for the 1.9 billion people served by High Mountain Asia's rivers, consequences the paper flags but does not quantify."],"forward_implications":["The weighted blend reproduces historical precipitation more accurately than the equal-weight average, with annual average CRPS differences of about 0.34 mm per day over the West Himalaya and Karakoram, 0.24 mm per day over the Central and East Himalaya, and 0.14 mm per day over the Hengduan Shan and Southeast Tibet.","Under RCP8.5 in the far future (2066–2095), the mixture projects summer monsoon median precipitation increases of about 15% over the West Himalaya and Karakoram, 23% over the Central and East Himalaya, and 27% over the Hengduan Shan and Southeast Tibet.","For winter, the blend's projections diverge from the equal-weight baseline by up to −31% in median relative change over the inner Himalayas and Karakoram and up to +51% to +62% over the Tibetan Plateau and Southeast Tibet, implying different seasonal contributions to frozen water stores.","Differences between the blended and equal-weight projections reach tens of percent at the 95th percentile as well, which the paper reads as a higher possible probability of very high precipitation events, such as floods and landslides, than previously estimated.","The authors argue the framework transfers to other grids, model ensembles, and climate variables, because the Gaussian process surrogates can be evaluated at arbitrary points."],"supporting_citations":[{"why":"Supplies the study domain, time periods, and the 13-model CORDEX-WAS ensemble used throughout, with the equal-weight aggregation serving as the baseline that the 32% improvement claim is measured against.","marker":"Sanjay et al. (2017)"},{"why":"Defines APHRODITE, the gridded observational precipitation dataset against which the surrogate Wasserstein distances and all validation scores are computed.","marker":"Yatagai et al. (2012)"},{"why":"Provides the continuous rank probability score (CRPS) used to quantify the 32% and 254% improvement claims.","marker":"Hersbach (2000)"},{"why":"Supplies the Gaussian process regression machinery from which every regional climate model surrogate is built.","marker":"Rasmussen and Williams (2006)"},{"why":"Introduces warped Gaussian processes, the mechanism that gives the surrogates their skewed, log-normal-like precipitation distributions.","marker":"Snelson et al. (2003)"},{"why":"Introduces chained Gaussian processes, used here so that both the mean and the variance of precipitation become learned spatial fields.","marker":"Saul et al. (2016)"},{"why":"Provides the distributed (robust Bayesian committee machine) combination of Gaussian processes that makes fitting over the full domain computationally tractable.","marker":"Deisenroth and Ng (2015)"},{"why":"Supplies the Wasserstein distance formalism whose per-model values directly parametrise the mixture-of-experts weights.","marker":"Panaretos and Zemel (2019)"}],"fun_headline_variants":["Climate model blend beats best single model by 254%","Probabilistic blend sharpens Asia precipitation forecasts","13 climate models, one smarter forecast for Asia's water","Ensemble learning cuts Asia precipitation error 32%","Smarter blend rewrites High Mountain Asia's rain future"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that weights optimised on how well each model matched APHRODITE precipitation over 1951–2005 remain the correct weights for 2036–2095, because Section 5 fixes the weights once on the whole historical record and applies them unchanged to future projections without any test that a model's ranking by precipitation fidelity is stationary across climate states.","fun_headline_variants_meta":{"raw":{"variants":["Climate model blend beats best single model by 254%","Probabilistic blend sharpens Asia precipitation forecasts","13 climate models, one smarter forecast for Asia's water","Ensemble learning cuts Asia precipitation error 32%","Smarter blend rewrites High Mountain Asia's rain future"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000591,"raw_usage":{"total_tokens":2822,"prompt_tokens":1044,"completion_tokens":1778,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":1700}},"tokens_in":660,"tokens_out":1778,"duration_ms":11949,"temperature":1.0,"reasoning_tokens":1700,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:03:09.670395+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-fit the mixture weights on two non-overlapping historical windows, such as 1951–1970 and 1971–1990, and compare the resulting weight maps: if the optimal weights shift substantially between windows, or if weights fit on one window fail to beat equal weighting on the other, the fixed-weight future projections rest on an unsupported assumption. A complementary check is to score the same learned weights against a second gridded precipitation product not used in training, since the reported 32% gain is measured only against APHRODITE.","supporting_citations":[],"review_version":1}