{"id":"36613094-9f2d-406c-a4c6-cb4d8d5d7d87","arxiv_id":"2501.02814","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An autoencoder-based analogue forecast system with optimised feature weights improves Hong Kong daily rain-class forecasts, particularly for heavy rain, compared with the existing operational system.","lead":"The paper describes an upgraded analogue forecast system at the Hong Kong Observatory that uses a deep autoencoder to compress 60 weather fields and then finds similar historical days to forecast daily rain class up to nine days ahead. It reports better heavy-rain detection than the current operational system during 2019 to 2022, but the evaluation partly overlaps the data used to tune the system.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline performance comparison is compromised because the Optuna feature-weight optimisation (Section 2.2.3) used Day-1 forecasts from 2017-2020, and the stated verification period (May 2019-Apr 2022) overlaps by about 20 months, so the reported superiority is partly in-sample rather than an…","rationale":"The reader's weakest_assumption correctly identifies the core problem: the feature-weight optimisation period (2017-2020) overlaps with the verification period (May 2019-Apr 2022), so the enhanced AFS's superiority is partly in-sample. This is the most load-bearing issue because every headline number in Section 3.1 - the heavy-rain POD above 0.5, CSI near 0.4, and the factor-of-two improvement in correct heavy-rain forecasts - is drawn from a sample that was used to fit the weights. The concern is not that the method is unsound or the authors are misleading; it is that the reported verification design does not provide an independent test of the central claim. The paper also has no uncertainty quantification, and with roughly 73 heavy-rain days the difference between 38 and 16 correct forecasts may not be statistically robust, but the overlap alone is sufficient to justify a conditional verdict. The paper otherwise contains useful operational detail, a deployed system, and honest discussion of limitations in Section 4, so I do not see grounds for rejection. The reader's CONDITIONAL verdict remains appropriate: the claim should be accepted only after a clean temporal split and ideally confidence intervals are provided.","tokens_in":12605,"tokens_out":3520,"duration_ms":35564,"concrete_test":"Recompute the Day-1 to Day-9 confusion matrices and heavy-rain CSI/POD using only the 12 disjoint verification months May 2021-Apr 2022, which fall entirely after the 2017-2020 Optuna optimisation window. Also retrain the feature weightings on 2017-2018 only and verify on May 2019-Apr 2022. If the enhanced AFS no longer outperforms the existing AFS on the disjoint months, the headline claim is not supported; if it does, the overlap concern is mitigated. Accompany either comparison with bootstrap 95% confidence intervals on the heavy-rain CSI difference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the enhanced AFS 'demonstrates a consistent and superior performance over the existing AFS, especially in capturing heavy rain cases' over May 2019-Apr 2022. For this to be credible, the verification must measure generalisable skill, not skill that was tuned on the verification days themselves. Section 2.2.3 states that Optuna minimises the MSE between the rain class of each day in 2017-2020 and the rain class of its best analogue, using Day-1 HRES forecasts from 2017-2020. The verification period in Section 2.1.2 is May 2019 to April 2022, so roughly 20 of the 36 verification months (May 2019-Dec 2020) also participated in fitting the 60 feature weightings. The optimiser is directly rewarded for selecting analogues whose historical rain class matches the observed rain class on those days. Therefore the Day-1 heavy-rain hit rate of 38/73 (52%), CSI near 0.4, and the 'more than double' improvement over the existing AFS are at least partly in-sample. The existing AFS was not re-optimised on this period, so the comparison is biased in favour of the enhanced AFS. No confidence intervals are provided, and with only about 73 heavy-rain events in the three-year window the sampling uncertainty is material. Because the paper's strongest quantitative claim rests on this overlapping optimisation and verification sample, the reported superiority is not yet established as an unbiased operational skill estimate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes an enhanced analogue forecast system (AFS) for daily precipitation prediction in Hong Kong. The system uses an autoencoder trained on ERA5 reanalysis to compress 60 meteorological fields into feature vectors, applies Optuna-optimised feature weightings to select analogues from ERA5, and forms an ensemble of the 25 most similar analogues to produce a rain-class forecast. The enhanced AFS is verified over May 2019 to April 2022 and compared with the existing HKO AFS, forecaster bulletins, and ECMWF direct output. The authors report that the enhanced AFS outperforms the existing AFS, especially for heavy rain, with a Day-1 CSI near 0.4, POD above 0.5, and 38 of 73 heavy-rain days correctly predicted.","tokens_in":13032,"tokens_out":4181,"duration_ms":38502,"significance":"If the reported skill holds, the enhanced AFS is a practically valuable upgrade to an operational precipitation guidance system, with a clear methodological advance in using autoencoder feature extraction and a detailed, reproducible description of the workflow. The paper also provides a useful comparison against existing operational practice. However, the central verification claim is compromised by the overlap between the optimisation period (2017-2020) and the verification period (May 2019-April 2022), which affects roughly 20 of the 36 verification months. The reported superiority over the existing AFS is therefore partially in-sample, and the absence of uncertainty quantification further weakens the quantitative claims. The strengths are the complete system description, the use of a long ERA5 archive (1979-2020), the real-time operational deployment, and the transparent confusion matrices.","major_comments":[{"comment":"The verification period (May 2019-April 2022) overlaps with the feature-optimisation period (2017-2020) by approximately 20 months. Section 2.2.3 states that Optuna minimises the MSE between the rain class of each day in 2017-2020 and the rain class of its best analogue, using Day-1 HRES forecasts; hence days from May 2019 through December 2020 directly participated in selecting the 60 weightings in Table 3. The Day-1 heavy-rain statistics (38/73 hits, CSI ≈ 0.4, POD > 0.5) and the 'more than double' improvement over the existing AFS are therefore partly in-sample, and the comparison with the existing AFS (whose parameters were not re-optimised on this period) is biased in favour of the enhanced AFS. To establish the headline claim, the authors should either report skill on the subset of the verification period outside the optimisation period (e.g., January 2021-April 2022), or apply a cross-validated/nested optimisation and show that the conclusion is unchanged.","section":"Sections 2.1.2, 2.2.3, 3.1"},{"comment":"The paper reports point-value verification metrics without confidence intervals or significance tests. The central comparison rests on only 73 heavy-rain events over three years; with this sample size the difference between 38 and 16 hits, or the CSI difference near 0.4, may not be statistically robust. Please provide interval estimates (e.g., bootstrap or a proper score test) for at least the Day-1 heavy-rain metrics, and ideally for the other lead times and rain classes.","section":"Section 3.1, Table 5"},{"comment":"Several key design choices are justified only as 'empirical' or 'based on tested results': the Gaussian radius r0 = 7 in Eq. (1), the weighted-score formula with the sample-to-population ratio threshold in Eq. (2), and the fixed ensemble size of 25 analogues. The paper presents no sensitivity analysis or ablation for these choices. This matters because Section 4.2 itself warns against overfitting the 'empirical equation' to the verification period; a sensitivity study would show whether the reported heavy-rain advantage is robust to these choices or an artefact of the specific settings.","section":"Sections 2.2.3, 2.2.4, 4.2"}],"minor_comments":[{"comment":"The text contains a typo: 'predication of moderate rain' should be 'prediction of moderate rain'.","section":"Section 3.1"},{"comment":"The acronym 'ECMWF' is misspelled as 'ECWMF' in the sentence 'the averaged ECWMF direct model output' and in the figure captions.","section":"Section 3.1 and Figures 6-8"},{"comment":"The phrase 'Confusion Comparison' in the figure titles appears to be a typo; it should likely read 'Comparison' or 'Performance Comparison'.","section":"Figures 6-8"},{"comment":"The row label 'Heavy ≤' is ambiguous; it would be clearer to use 'Heavy or above' to match the rain-class grouping in Table 2.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The optimisation/verification overlap is the principal substantive concern and must be addressed before the quantitative claims can be accepted. The paper's own caution in Section 4.2 about overfitting makes the absence of a clean hold-out verification particularly noticeable. If a hold-out evaluation confirms the skill, the paper would be a solid contribution to an operational meteorology venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Name],\n\nThe short version: the paper describes a concrete, operational analogue forecasting system for Hong Kong daily rain classes, and it does a good job explaining the engineering. The skill claim, however, is not yet established, because the verification period overlaps with the optimization period.\n\nWhat's new: combining autoencoder feature extraction on 60 ERA5/HRES fields, per-feature weights tuned with Optuna, and a 25-analogue weighted ensemble. Each piece is known, but the specific system is new and it's running at HKO, which gives it practical grounding. The paper also publishes confusion matrices and the usual scores, so you can recompute the numbers. The case study of 11–13 May 2022 is a fair illustration, not a proof.\n\nThe soft spots are in the verification. Weights are optimized on Day-1 HRES forecasts from 2017–2020, minimizing the rain-class MSE between each day's observed class and its best analogue. The verification window is May 2019–April 2022, so roughly 20 of 36 verification months also helped set the 60 weights. The Day-1 heavy-rain scores (CSI ~0.4, POD ~0.5) are therefore partly in-sample, and the existing AFS, which was not re-optimized on this period, is a biased comparator. No confidence intervals are given, and with only about 73 heavy-rain days in the window, the sampling uncertainty is not trivial. The comparison also changes several other things at once (archive length, resolution, search window, ensemble size), so you can't attribute the improvement specifically to the autoencoder. Finally, the optimization target is single-best-analogue skill, while the actual forecast uses the 25-analogue ensemble; that mismatch is unaddressed.\n\nNone of this makes the paper worthless. The method is plausible, the system is deployed, and the authors are honest about the inherent limits of analogues for unprecedented extremes. But the headline claim needs a cleaner test. I'd ask for a fully held-out verification period (e.g., train on 2017–2019, test on 2020–2022, or at least report the post-2020 portion separately) and bootstrap confidence intervals. If the edge survives that, it's a useful operational contribution; until then, treat the superiority as plausible but unproven.\n\nRecommendation: send it to review. The flaw is fixable and the system deserves a proper evaluation.\n\n[Your name]","headline":"A genuinely operational autoencoder-based analogue forecast system, but the headline skill claim is compromised by overlapping optimization and verification periods.","tokens_in":13514,"tokens_out":3553,"would_cite":false,"duration_ms":33998,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An enhanced analogue forecast system correctly predicted over half of Hong Kong's heavy-rain days at Day 1, more than double the old system.","keywords":["Analogue forecast","Daily precipitation forecast","Autoencoder feature extraction","Heavy rain","Hong Kong","Numerical weather prediction","Machine learning"],"falsifier":"Re-score the enhanced and existing systems on a period that contributed no days to weight optimisation, for example forecasts issued from May 2022 onward, and compare heavy-rain CSI and POD; if the enhanced system no longer beats the existing system on this independent sample, the central claim collapses. A cleaner experiment is to re-optimise the weights using only 2017–2018 days and verify only on 2021–2022 days.","tokens_in":12458,"feed_emoji":"🌧️","tokens_out":11086,"duration_ms":101049,"temperature":0.7,"pith_summary":"The paper argues that replacing hand-crafted weather-pattern comparisons with autoencoder-learned feature vectors makes analogue forecasting markedly better for daily rain classes in Hong Kong. The new system compresses 60 gridded meteorological fields into compact vectors, optimises a weight for each field, and blends the 25 closest historical analogues into a rain-class forecast. Over a three-year verification period, more than half of observed heavy-rain days were correctly identified on Day 1, more than double the old system's count, and the heavy-rain advantage persisted through Day 9. If true, forecasters gain earlier, more consistent guidance on the damaging heavy-rain events, using the same global model output they already receive.","feed_headline":"Autoencoder analogue system doubles heavy-rain forecast hits","feed_subtitle":"Deep-learning feature extraction plus optimised weights gives forecasters sharper heavy-rain guidance to Day 9.","key_machinery":"The load-bearing component is the convolutional autoencoder. For each day, 60 fields (ten variables at six pressure levels) are remapped to a one-degree grid, normalised, and multiplied by a Gaussian weight $G(r_k)=\\exp(-r_k^2/(1.2 r_0)^2)$ centred on Hong Kong; the encoder compresses each field into a vector, and the decoder is trained to reconstruct the input, so each vector is meant to be a compact representation of the weather pattern. Similarity between a forecast day and an archive day is scored by the mean squared error of the extracted vectors, combined with per-feature weights optimised to minimise the rain-class error of the best matching analogue. The final rain class comes from a weighted mean of the 25 closest analogues, with each analogue's contribution scaled by the ratio of its rain class's occurrence in the 25-member sample to its past population frequency.","core_discovery":"The paper's central claim is that autoencoder-based feature extraction improves the analogue method for daily precipitation prediction. During the May 2019 to April 2022 verification, the enhanced AFS correctly predicted over half of the observed heavy-rain days at Day 1, with a critical success index near 0.4 and a probability of detection above 0.5, and it held a consistent heavy-rain edge over the existing AFS from Day 1 through Day 9 while keeping comparable performance for lighter rain classes. The authors interpret this as evidence that the autoencoder captures the synoptic patterns relevant to local rain more effectively than the previous hand-defined similarity scores, and that the optimised per-feature weights plus the ensemble of 25 analogues produce forecasts that are more accurate and more stable.","pith_inferences":["An independent test on forecasts issued after the weights were frozen, say from May 2022 onward, would give a cleaner estimate of the operational gain than the 2019–2022 verification.","Because the autoencoder is trained to reconstruct fields rather than to separate rain classes, a supervised or contrastive training objective could make the features even more precipitation-specific.","The pipeline should transfer to other regions with a long reanalysis archive and deterministic model output; the main requirements are enough historical cases and a sensible seasonal restriction on analogues.","Under climate change, the fixed historical archive may lack good analogues for unprecedented extremes, so the system would benefit from an explicit flag when no analogue is close enough."],"forward_implications":["Heavy-rain forecasting is the largest gain: at Day 1 the enhanced AFS reaches a critical success index near 0.4 and a probability of detection above 0.5, outperforming the existing AFS, forecaster bulletins, and direct model output for that class.","The advantage is not a one-day effect: the enhanced AFS keeps higher heavy-rain skill from Day 1 through Day 9 and nearly eliminates Day 1 forecasts that miss the observed rain class by more than one step.","The 25-analogue ensemble gives forecasters a traceable set of historical scenarios and a range of plausible rainfall amounts, not just a single class, reducing run-to-run fluctuation.","False alarms of light rain on dry days increase slightly, so users should expect a small wet bias accompanying the improved heavy-rain detection.","The system has been running in real-time operations since May 2022, meaning its reported behaviour can be checked against daily use.","The paper itself warns against overfitting the empirical scoring equation to the verification period; the same caution applies to the overlapping optimisation and verification years, though the paper does not directly address that overlap."],"supporting_citations":[{"why":"The existing AFS that serves as the baseline and source of the Gaussian weighting and original predictor design.","marker":"Chan et al., 2014"},{"why":"Defines atmospheric analogues as similar states, the concept that motivates searching historical cases.","marker":"Lorenz, 1969"},{"why":"Grounds the analogue method for quantitative precipitation forecasts, the methodological basis of the system.","marker":"Hamill & Whitaker, 2006"},{"why":"Supplies the definition of autoencoders and representation learning used to justify the feature extractor.","marker":"Bank et al., 2020"},{"why":"Provides the hyperparameter optimisation framework used to tune the per-feature weightings.","marker":"Akiba et al., 2019"},{"why":"Provides the Tree-structured Parzen Estimator algorithm used by that weight optimisation.","marker":"Bergstra et al., 2011"}],"fun_headline_variants":["Autoencoder sharpens analogue rain forecasts","Deep learning tunes Hong Kong's rain analogues","AI feature extraction improves heavy-rain prediction","Autoencoder-enhanced analogue system for HK rain","Enhanced analogue forecast lifts heavy-rain hits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 2019–2022 verification period fairly represents future operational performance, even though days in 2019–2020 also helped set the optimised feature weightings; if that overlap inflates the apparent skill, the reported superiority may not hold on genuinely new forecasts.","fun_headline_variants_meta":{"raw":{"variants":["Autoencoder sharpens analogue rain forecasts","Deep learning tunes Hong Kong's rain analogues","AI feature extraction improves heavy-rain prediction","Autoencoder-enhanced analogue system for HK rain","Enhanced analogue forecast lifts heavy-rain hits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1488,"prompt_tokens":975,"completion_tokens":513,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":448}},"tokens_in":591,"tokens_out":513,"duration_ms":5416,"temperature":1.0,"reasoning_tokens":448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:03:30.545446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the enhanced and existing systems on a period that contributed no days to weight optimisation, for example forecasts issued from May 2022 onward, and compare heavy-rain CSI and POD; if the enhanced system no longer beats the existing system on this independent sample, the central claim collapses. A cleaner experiment is to re-optimise the weights using only 2017–2018 days and verify only on 2021–2022 days.","supporting_citations":[],"review_version":1}