{"id":"d49e0571-91f2-4497-963f-5fa9022ab85d","arxiv_id":"2411.10921","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Adding attention modules to a ConvLSTM cloud forecasting model improves downstream hour-ahead solar power forecast skill scores, especially for high altitude clouds.","lead":"This paper tests whether attention-based deep learning models improve cloud movement forecasts from satellite images and whether those improved cloud forecasts lead to better short-term solar power forecasts at 50 rooftop PV sites in Australia. It reports that attention-based cloud forecasts improve solar forecast skill scores by about 6 percentage points under high altitude cloud conditions compared with a standard ConvLSTM baseline.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 5.86% high-altitude gain rests on one arbitrary pixel-value threshold with no uncertainty quantification; the claim may be threshold-dependent.","rationale":"The reader's verdict is CONDITIONAL, and my analysis supports that disposition but for a somewhat different reason. The reader identified the weakest assumption as the vertical alignment between the satellite pixel and the cloud actually shading the PV site. That concern is acknowledged by the authors and is worth testing, but it is symmetric across all cloud-input methods: if the extracted pixel is misaligned, the noise affects ground-truth clouds, persistence clouds, ConvLSTM clouds, and attention-based clouds alike. It could weaken absolute skill scores or blur the high/low altitude distinction, but it is not the most direct threat to the relative attention-versus-ConvLSTM improvement that constitutes the headline claim. The more load-bearing issue is that the headline number is computed on a post-hoc subset (pixel value > 50) with no sensitivity analysis and no uncertainty quantification. Since the authors themselves state that cloud altitude is not explicitly calculated, the exact cut is arbitrary. With only 16.24% of test samples in the high-altitude class, the difference between 5.86% and a smaller or non-significant effect could easily arise from the choice of threshold or from a few influential sites. A threshold sweep with per-site paired bootstrap intervals would settle whether the effect is real or an artifact of the chosen cut. Because this is an addressable empirical check rather than a demonstrated error, the appropriate verdict remains CONDITIONAL rather than ACCEPT or REJECT; the condition should be that the authors provide threshold-robustness and significance evidence for the headline improvement.","tokens_in":22972,"tokens_out":4745,"duration_ms":53880,"concrete_test":"Recompute Tables 8 and Figures 8-10 for a sweep of altitude thresholds (pixel > 20, 35, 50, 65, 80, 100) and also with the pixel value treated continuously, using per-site paired differences and 95% bootstrap confidence intervals for CBAMConvLSTM/SAConvLSTM minus ConvLSTM. If the high-altitude improvement is not consistently positive and statistically significant across neighboring thresholds, the headline 5.86% claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that attention-based cloud forecasts improve solar skill scores by 5.86% or more under high-altitude clouds. This is established only on a conditional subset defined by an infrared pixel-value cutoff of 50, with no sensitivity analysis around that cutoff and no error bars or significance tests. The authors acknowledge in Section 6.4 that cloud altitude is not explicitly calculated, making the cutoff a post-hoc proxy. High-altitude samples are only 16.24% of the test set, so the reported point estimates are averages over a small subset; small changes in the threshold, a few outlier sites, or an unlucky data split could materially change the 5.86-6.67 percentage-point improvements. The reader's vertical-alignment concern is legitimate, but it applies symmetrically to all compared cloud inputs and therefore is less likely to overturn the relative attention-vs-ConvLSTM comparison. Threshold dependence, by contrast, directly targets the quantitative headline: if the advantage shrinks or reverses at neighboring thresholds, the claim '5.86% or more' is not robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an attention-augmented cloud forecasting pipeline for hour-ahead solar generation forecasting at 50 rooftop PV sites in Perth, Australia, using Himawari-8 infrared satellite imagery. It introduces CBAMConvLSTM, applies the existing SAConvLSTM, and compares both against a standard ConvLSTM as cloud forecasters; cloud forecasts are then fed into MLP, CNN, and LSTM solar power predictors, with additional benchmark conditions using ground-truth clouds, persistence clouds, and no clouds. The central empirical claim is that attention-based cloud forecasts yield RMSE skill score improvements of 5.86% or more over ConvLSTM-derived cloud forecasts under high-altitude cloud conditions.","tokens_in":23109,"tokens_out":2841,"duration_ms":33863,"significance":"If the central claim is robust, the paper provides a useful large-scale empirical evaluation: 50 PV sites, three solar forecasting architectures, three cloud forecasting models, and multiple cloud-input benchmark scenarios. The explicit two-stage evaluation and the decomposition into clear-sky, high-altitude, and low-altitude conditions are valuable for practitioners, and the reported point estimates consistently favor the attention-based methods under cloudy conditions. The strengths are the breadth of the study, the inclusion of persistence and no-cloud baselines, and the investigation of downstream forecast impact rather than only image-quality metrics. The main caveat is that the headline quantitative claim rests on a post-hoc sample stratification whose robustness is not demonstrated.","major_comments":[{"comment":"The headline claim of 5.86% or greater skill score improvement depends entirely on the threshold of pixel value > 50 used to define high-altitude cloud conditions in Section 6.2. This threshold is introduced without justification, and Section 6.4 explicitly acknowledges that cloud altitude is not calculated in the work, making the threshold a proxy. Because high-altitude samples are only 16.24% of the test set, small changes in the threshold or a few outlier sites could materially change the reported 5.86-6.67 percentage-point improvements. I request a sensitivity analysis over a range of plausible thresholds, together with per-site confidence intervals or paired significance tests, before the abstract's quantitative claim can be accepted.","section":"Section 6.2, Table 8 and Figures 8-10"},{"comment":"The training, validation, and testing split is not described beyond sample counts: no information is given on whether the split is temporal, random, site-level, or contiguous. Since solar power and satellite image data are 10-minute time series, a random split could place temporally adjacent samples from the same day in both training and testing sets, potentially inflating skill scores. Please specify the exact splitting procedure and confirm that no temporal overlap exists between training, validation, and test periods.","section":"Section 5, Table 4"},{"comment":"All reported skill scores are point estimates averaged over 50 PV sites, with no measure of cross-site variability, standard errors, confidence intervals, or significance tests. Differences of 1-2 percentage points between attention-based and non-attention methods may be within site-level noise, and the conclusion that attention-based methods provide 'meaningful benefits' is currently supported only by the magnitudes of the averaged differences. I request site-level error bars or paired statistical tests for the attention-versus-ConvLSTM comparison, at least for the high-altitude subsample that anchors the headline claim.","section":"Tables 5-10"}],"minor_comments":[{"comment":"The caption contains typographical errors: 'N 4 2020' should be 'Nov 4 2020' and 'F b14 2021' should be 'Feb 14 2021'.","section":"Section 3, Figure 4"},{"comment":"Several typos should be corrected: 'refereed' should be 'referred', 'investiage' should be 'investigate', and 'huperparameters' should be 'hyperparameters'.","section":"Section 4.2.1"},{"comment":"The introduction to Table A.11 states that it shows the average RMSE skill score, but the table reports MAE skill scores; the text should refer to MAE consistently.","section":"Appendix A"},{"comment":"Figure 8 labels one scenario 'Lower altitude clouds only'; for consistency with Section 6.2 and Tables 8-9, this should be 'low altitude clouds only'.","section":"Section 6.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of an applied machine-learning journal, and the empirical breadth is a genuine strength. However, the abstract's precise quantitative claim is contingent on a threshold and a data split that are currently not justified or tested for robustness, so major revision is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a useful, honestly reported empirical study of whether attention-based cloud movement forecasting (CBAMConvLSTM and SAConvLSTM) beats a plain ConvLSTM when the predicted cloud pixels are fed into hour-ahead solar forecasts. Across 50 PV sites in Perth and three solar forecast networks, attention models do better almost everywhere, and the biggest gains appear under high-altitude cloud conditions. I think the direction of the result is probably right. The exact headline number—5.86% or more skill-score improvement—is less solid than it looks, because it comes from a post hoc pixel-value threshold of 50, covers only about 16% of test samples, and has no error bars or significance tests. I'd believe the qualitative claim before quoting the number.\n\nWhat is new: CBAMConvLSTM is a straightforward insertion of CBAM into the ConvLSTM gates, and SAConvLSTM is an existing video-prediction architecture, but neither has been applied to cloud forecasting with a downstream solar forecast evaluation. The experiment is broader than prior work: 50 sites, a year of 10-minute Himawari-8 infrared imagery plus PV generation, three solar forecasters, and sensible baselines (persistence clouds, ground-truth clouds, no-cloud inputs). The limitation section is unusually honest. The authors admit that cloud altitude is not explicitly calculated and that the cloud vertically above a site may not be the cloud shading the panel. That honesty counts.\n\nThe soft spots are in the magnitude, not the direction. The high/low altitude split at pixel value 50 is not justified, and there is no sensitivity analysis around it. Since high-altitude samples are 16.24% of the test set, a few outlier sites or a small shift in the cutoff could move the 5.86–6.67 percentage-point gains meaningfully. The train/test split is not described, so I cannot tell whether the cloud forecasting networks overlap the test period. The solar forecasting networks are trained with ground-truth future cloud pixels, which is fine as an oracle upper bound, but it means the realized benefit depends on how closely forecast clouds approach that oracle. Error bars, site-level variance, and code/data release would address the main worries. The vertical-alignment issue is real but applies symmetrically to all cloud inputs, so it is unlikely to flip the attention-vs-ConvLSTM comparison. The baseline tables look internally consistent, and the citation pattern is normal.\n\nThis paper deserves a serious referee. It should not be desk rejected, but it needs revision: uncertainty quantification, threshold sensitivity, and better data documentation. I would not cite it as a reliable quantitative result yet, but it is a fair empirical contribution.","headline":"Solid empirical comparison of attention-based cloud forecasting for solar, but the headline skill-score gain rests on one arbitrary altitude threshold with no uncertainty quantification.","tokens_in":23679,"tokens_out":2940,"would_cite":false,"duration_ms":30602,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By adding attention to the cloud-movement forecaster and feeding those cloud forecasts into solar power predictors, the paper finds hour-ahead solar forecast skill improves by 5.86% or more under high-altitude cloud conditions compared…","keywords":["solar photovoltaic forecasting","cloud movement prediction","attention mechanism","ConvLSTM","infrared satellite imagery","satellite image forecasting","deep neural networks","distributed solar generation"],"falsifier":"For a sample of test timestamps, identify the cloud pixel that actually shades each PV site using solar azimuth, zenith angle, and cloud-top height, then rerun the same solar forecasting networks with those corrected cloud values; if the attention-based cloud forecasts no longer beat the ConvLSTM forecasts by 5.86% or more in high-altitude conditions, the central claim's mechanism is undermined.","tokens_in":22747,"feed_emoji":"☀️","tokens_out":5798,"duration_ms":53693,"temperature":0.7,"pith_summary":"This paper asks whether attention mechanisms, which let a neural network focus on the informative parts of an image, improve the forecasting of cloud movement and thereby improve hour-ahead solar generation forecasts. To answer it, the authors build a two-stage pipeline: an attention-enhanced ConvLSTM (CBAMConvLSTM) and an existing self-attention ConvLSTM predict future infrared satellite images, and then three separate neural networks (MLP, CNN, LSTM) use the predicted cloud pixel above each of 50 rooftop PV sites to forecast solar power. Across all sites and test conditions, the cloud forecasts from both attention-based networks lead to higher solar forecast skill scores than forecasts from a standard ConvLSTM. The largest and most consistent gains are for high-altitude cloud conditions, where the attention-based cloud inputs improve RMSE skill score by 5.86% or more relative to the non-attention baseline. The paper concludes that attention-based cloud movement forecasting is a useful ingredient for short-term distributed solar forecasting, particularly when high-altitude clouds are present.","feed_headline":"Attention improves cloud-based solar forecasts by 5.86%","feed_subtitle":"On 50 PV sites, attention-based cloud forecasts beat standard ConvLSTM; biggest gains for high-altitude clouds.","key_machinery":"The carrying object is the attention-augmented convolutional LSTM. The paper's CBAMConvLSTM applies the convolutional block attention module (CBAM), a channel-and-spatial attention module that reweights feature maps, after each convolution inside a standard ConvLSTM cell, so the cell emphasizes informative cloud regions while ignoring non-cloud areas. The second model, SAConvLSTM, is an existing self-attention ConvLSTM that computes pairwise pixel correlations and keeps a separate memory of long-range spatial dependencies. Both predict a 60x60 infrared image sequence six time steps (one hour) ahead in an autoregressive loop, and the pipeline then extracts the pixel corresponding to each PV site as the cloud input to the solar forecasters. The cloud forecasts are trained with a structural similarity (SSIM) loss, and the solar networks are trained to minimize mean squared error, with performance measured as RMSE and MAE skill scores against a persistence forecast.","core_discovery":"The central claim is that cloud forecasts produced by attention-based sequence predictors translate into measurably better hour-ahead solar power forecasts than cloud forecasts from a standard ConvLSTM. The proposed CBAMConvLSTM inserts a convolutional block attention module after each convolution inside a ConvLSTM, while SAConvLSTM adds a self-attention memory module; both are trained to predict the next six infrared satellite images. When the single cloud pixel above a PV site is extracted from these predicted images and fed, together with past power values, into MLP, CNN, or LSTM solar forecasters, the attention-produced cloud inputs consistently beat the ConvLSTM-produced inputs. In the high-altitude cloud test subset, the RMSE skill score improvements over ConvLSTM are 5.97%, 5.86%, and 6.67% for MLP, CNN, and LSTM, respectively, and SAConvLSTM achieves 5.97%, 5.44%, and 6.29%. The same comparison under low-altitude cloud conditions yields smaller improvements, between 0.89 and 1.54 percentage points, indicating that the benefit of attention is concentrated where the cloud signal in the infrared image is strongest.","pith_inferences":["If the vertical-pixel assumption is corrected using solar azimuth, zenith angle, and cloud height, attention-based methods may show different (possibly larger or partially reduced) gains; testing that correction would separate prediction quality from input alignment effects.","Combining infrared imagery with visible-band imagery, which carries cloud-thickness information, could extend the attention benefit to low-altitude clouds, where the current improvements are smaller.","The concentration of gains in high-altitude clouds suggests attention networks are particularly good at reproducing near-white, high-pixel-value regions; a pixel-wise error analysis by cloud brightness would test this directly.","The same two-stage pipeline could be applied to other image-driven renewable forecasting tasks, such as wind gust nowcasting or irradiance forecasting, whenever the relevant atmospheric feature moves across a satellite image."],"forward_implications":["For cloudy periods, feeding attention-based cloud forecasts into any of the three solar forecasting networks yields higher RMSE skill scores than feeding ConvLSTM cloud forecasts or persistence cloud images.","The largest attention benefit appears specifically in high-altitude cloud test samples, where LSTM solar forecasts gain 6.67 percentage points in skill score compared with ConvLSTM cloud inputs.","Under clear-sky conditions, cloud forecasting method makes little difference, so the attention value is tied to cloudy-sky samples.","Including ground-truth cloud values still beats all forecasted-cloud variants, so further gains in cloud prediction accuracy should keep improving solar forecasts."],"supporting_citations":[{"why":"Defines the standard ConvLSTM structure used as the non-attention baseline cloud forecasting network.","marker":"[31]"},{"why":"Supplies the convolutional block attention module (CBAM) that the paper integrates into ConvLSTM to form CBAMConvLSTM.","marker":"[37]"},{"why":"Proposes the self-attention ConvLSTM (SAConvLSTM) that the paper adapts from video prediction to cloud movement forecasting.","marker":"[21]"},{"why":"Establishes the indirect pipeline of forecasting satellite cloud images and extracting the pixel above a PV site for solar forecasting.","marker":"[11]"},{"why":"Demonstrates the two-step satellite-image-based solar nowcasting approach that the paper extends with attention-based cloud forecasting.","marker":"[10]"},{"why":"Provides the cloud-type and altitude categories that motivate interpreting infrared pixel brightness as high- versus low-altitude cloud conditions.","marker":"[7]"}],"fun_headline_variants":["Attention-based cloud forecasts boost solar skill by 5.86%","High-altitude clouds show attention's solar forecast edge","Attention beats ConvLSTM for solar forecast skill on 50 sites","Cloud attention lifts solar forecast skill 5.86% for high clouds","Attention-based cloud prediction sharpens solar forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that the cloud pixel directly above a PV site in the satellite image represents the clouds that actually shade that site, although sunlight reaches the ground at an angle and the shading cloud can be displaced horizontally.","fun_headline_variants_meta":{"raw":{"variants":["Attention-based cloud forecasts boost solar skill by 5.86%","High-altitude clouds show attention's solar forecast edge","Attention beats ConvLSTM for solar forecast skill on 50 sites","Cloud attention lifts solar forecast skill 5.86% for high clouds","Attention-based cloud prediction sharpens solar forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000812,"raw_usage":{"total_tokens":3617,"prompt_tokens":1060,"completion_tokens":2557,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":2473}},"tokens_in":676,"tokens_out":2557,"duration_ms":17787,"temperature":1.0,"reasoning_tokens":2473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:08:53.637816+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a sample of test timestamps, identify the cloud pixel that actually shades each PV site using solar azimuth, zenith angle, and cloud-top height, then rerun the same solar forecasting networks with those corrected cloud values; if the attention-based cloud forecasts no longer beat the ConvLSTM forecasts by 5.86% or more in high-altitude conditions, the central claim's mechanism is undermined.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the standard ConvLSTM structure used as the non-attention baseline cloud forecasting network."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the convolutional block attention module (CBAM) that the paper integrates into ConvLSTM to form CBAMConvLSTM."},{"cited_title":"A Moment in the Sun: Solar Nowcasting from Multispectral Satellite Data using Self-Supervised Learning","cited_arxiv_id":"2112.13974","evidence_quote":"Demonstrates the two-step satellite-image-based solar nowcasting approach that the paper extends with attention-based cloud forecasting."},{"cited_title":"Barbieri, S","cited_arxiv_id":null,"evidence_quote":"Provides the cloud-type and altitude categories that motivate interpreting infrared pixel brightness as high- versus low-altitude cloud conditions."}],"review_version":1}