{"id":"4340653e-f96c-4d1f-a0b2-c6354dc7bb12","arxiv_id":"2607.12954","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Under physically constrained, heteroscedastic NWP perturbations, sequence models filter noise better than a strong tabular baseline and shift reliance toward history and physical priors.","lead":"This paper tests how well deep learning models for solar power forecasting hold up when weather forecasts are wrong, using a physics-aware simulation setup. It matters because real grid and plant operations never get perfect weather inputs, so robustness under realistic forecast error is what decides whether a model is usable.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the abstract-only limit already flagged by the Reader; the simulation-proxy assumption is the natural soft spot but cannot be stress-tested without methods, metrics, or artifacts.","rationale":"The Reader’s UNVERDICTED / LOW-confidence stance is the only defensible position given an abstract-only review. The strongest claim is an engineering ranking under a carefully designed but still synthetic noise model; the weakest assumption is precisely the transferability of that ranking. No further load-bearing flaw (e.g., an algebraic inconsistency, an unstated unboundedness condition, or a circular metric) can be extracted from the abstract. Credit is due for the physically motivated design choices (heteroscedasticity modulated by clear-sky, Erbs radiation consistency, multi-objective Pareto view) that already go beyond the “simplistic perturbations” the authors criticize. Until full methods, tables, and artifacts are available, the correct action is to leave the verdict and the identified soft spot unchanged rather than manufacture a deeper objection.","tokens_in":2114,"tokens_out":543,"duration_ms":5323,"concrete_test":"Obtain the full paper (or request the authors’ perturbation-generation code and virtual-PV simulator). Re-run the medium-to-high disturbance regime on at least one real plant time series with historically observed NWP errors (instead of the synthetic clear-sky/Erbs process). If the ranking of PatchTST/GRU/N-HITS versus LightGBM reverses or the SHAP/IG reallocation pattern disappears, the proxy assumption fails and the central claim does not transfer.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Reader correctly isolates the load-bearing assumption: that virtual PV power under clear-sky-modulated heteroscedastic NWP noise plus Erbs-consistent radiation is a faithful enough proxy for real NWP error structure and plant response that robustness rankings (sequence models > LightGBM in medium-to-high regimes) and SHAP/IG reallocation patterns transfer to operational forecasting. From the abstract alone this assumption is neither confirmed nor refuted; no quantitative error metrics, perturbation calibration details, real-vs-virtual validation, or released code/data exist to inspect. Because the paper is explicitly a controlled simulation isolating input-uncertainty propagation, the claim is internally coherent as far as the abstract states it. No additional internal inconsistency, circularity, or hidden mathematical failure can be demonstrated without the full text. The concern therefore remains exactly the one the Reader already elevated, and it is not yet load-bearing in a falsifying sense—only in an unverifiability sense.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes a physically constrained, simulation-based robustness evaluation framework for PV power forecasting models under NWP input errors. Virtual PV power is used as a controlled response to isolate input-uncertainty propagation from plant-level confounders. Six ML/DL models (including PatchTST, GRU, N-HITS, and LightGBM) are compared under dynamic NWP perturbations with clear-sky-modulated heteroscedasticity and Erbs-consistent radiation reconstruction. The abstract claims that sequence models provide stronger noise filtering and temporal resilience than LightGBM in medium-to-high disturbance regimes; that SHAP and Integrated Gradients indicate case-level feature reallocation from corrupted future forecasts toward historical observations and physical priors; and that a Pareto analysis of clean-condition accuracy, robustness, and latency yields engineering guidance for model selection under forecast uncertainty.","tokens_in":2285,"tokens_out":826,"duration_ms":16797,"significance":"If the quantitative results hold under full scrutiny, the work would address a genuine operational gap: most PV forecast evaluations still rely on perfect-forecast or simplistic-noise assumptions that do not capture temporally correlated, state-dependent, and physically coupled NWP errors. A controlled simulation that preserves radiation consistency, combined with multi-model comparison, attribution (SHAP/IG), and Pareto trade-offs, would be practically useful for robustness-aware model selection. The explicit isolation of input-error propagation is a methodological strength relative to purely observational benchmarks that confound plant and weather effects.","major_comments":[{"comment":"Only the abstract is available for this review. The central ranking—sequence models (PatchTST, GRU, N-HITS) outperforming LightGBM under medium-to-high disturbance—cannot be verified: no error metrics, confidence intervals, sample sizes, site diversity, statistical tests, or ablation tables are inspectable. Without those results the load-bearing claim remains unsubstantiated.","section":null},{"comment":"The framework’s fidelity assumption is load-bearing: that virtual PV under clear-sky-modulated heteroscedastic NWP noise plus Erbs-consistent radiation is a faithful enough proxy for real NWP error structure and plant response that robustness rankings transfer to operations. The abstract states the isolation intent but provides no calibration against real NWP archives, real-vs-virtual residual diagnostics, or sensitivity to the free perturbation schedule and plant parameters. Transferability therefore cannot be assessed from the available text.","section":null},{"comment":"The SHAP/IG feature-reallocation claim is presented as case-level evidence of a shift from corrupted future forecasts toward historical observations and deterministic physical priors. Without systematic aggregation across disturbance levels, controls under clean inputs, or comparison of attribution stability, it is unclear whether the pattern is robust or anecdotal; the abstract alone does not establish it as a general mechanism.","section":null}],"minor_comments":[{"comment":"The abstract lists models incompletely (“six representative… including”); the full set and selection rationale should be named for reproducibility.","section":null},{"comment":"Key experimental design quantities (number of sites/scenarios, horizon, metrics, whether real NWP error statistics were used to set the heteroscedasticity schedule) are absent from the abstract and would help readers gauge scope.","section":null},{"comment":"Clarify whether code, virtual-plant parameters, and perturbation generators will be released; reproducibility is especially important for a simulation-driven robustness study.","section":null}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; full text, figures, tables, and any supplementary code/data were not available. The assessment is therefore provisional. The simulation-proxy assumption flagged by the reader is the natural soft spot but is not yet a demonstrated internal inconsistency—only an unverifiability concern. I recommend a full re-review once the complete manuscript is provided; until then a firm accept/reject decision is not warranted."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: this is a controlled simulation framework that tries to evaluate PV forecasting models under NWP errors that are temporally correlated, state-dependent, and radiation-consistent, rather than under perfect forecasts or ad-hoc noise. That is the real contribution, not a new architecture.\n\nWhat they do well is package the evaluation carefully. Virtual PV power as the response isolates input-error propagation from plant confounders. Clear-sky-modulated heteroscedasticity plus Erbs-consistent radiation is a more physical perturbation design than most of the literature uses. Comparing sequence models (PatchTST, GRU, N-HITS) against LightGBM under medium-to-high disturbance, then adding SHAP/IG evidence of feature reallocation toward history and physical priors, and closing with a Pareto of accuracy–robustness–latency, is a coherent engineering story. Circularity looks modest: they are not claiming to predict a fitted identity; they are stress-testing models under controlled input noise.\n\nThe soft spot is exactly the one the reader flagged and the stress-test did not invent: whether the virtual plant plus this noise model is faithful enough that the ranking (sequence models more robust than LightGBM) and the reallocation pattern transfer to real NWP error and real plants. From the abstract alone we cannot check calibration against actual NWP residuals, site diversity, error bars, or ablations. Free parameters (perturbation schedule, virtual conversion model) exist but are not automatically fatal; they just need to be documented and sensitivity-tested. No internal contradiction or load-bearing circularity is visible in the abstract.\n\nThis paper is for people who select or deploy PV forecast models under NWP uncertainty—operators, applied ML for renewables, and anyone tired of perfect-forecast benchmarks. It deserves a serious referee if the full text supplies the metrics, calibration, and preferably code/data. I would not desk-reject it. Bring it to reading group only if someone is actively working NWP-error robustness; otherwise it is a solid methods note, not a must-read. I would cite it if the full results hold and I am writing on forecast robustness under realistic weather-model error.","headline":"Solid applied methods package for PV forecast robustness under realistic NWP error; abstract-only so rankings and transfer remain unchecked.","tokens_in":2934,"tokens_out":529,"would_cite":false,"duration_ms":4238,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["92.60.Wc","89.60.-k","07.05.Mh"],"model":"grok-4.5","headline":"Sequence models filter dynamic NWP forecast errors better than strong tabular baselines for PV power prediction, reallocating reliance to history and physical priors.","keywords":["photovoltaic power forecasting","NWP forecast errors","robustness evaluation","sequence models","PatchTST","LightGBM","SHAP","Integrated Gradients"],"falsifier":"Apply the same model suite and perturbation protocol to real multi-site PV plants with measured NWP error archives; if sequence models no longer outperform LightGBM on medium-to-high disturbance days or SHAP/IG no longer show the reallocation pattern, the central claim fails.","tokens_in":2946,"feed_emoji":"☀️","tokens_out":868,"duration_ms":7049,"temperature":0.7,"pith_summary":"This paper argues that engineering-grade AI for photovoltaic power forecasting must stay predictable when numerical weather prediction inputs are wrong in realistic ways: temporally correlated, state-dependent, and physically coupled. The authors build a controlled simulation that isolates how those input errors propagate into power forecasts, using virtual PV power as the response so plant-level confounders do not cloud the comparison. Under dynamic, clear-sky-modulated heteroscedastic NWP noise and Erbs-consistent radiation reconstruction, sequence models such as PatchTST, GRU, and N-HITS show stronger noise filtering and temporal resilience than a strong tabular baseline (LightGBM) once disturbance levels move into the medium-to-high range. Case-level SHAP and Integrated Gradients evidence indicates that the sequence models reallocate predictive weight away from corrupted future forecasts toward historical observations and deterministic physical priors. A Pareto view of clean accuracy, robustness, and latency then turns the rankings into concrete guidance for model selection under forecast uncertainty.","feed_headline":"Sequence models beat tabular baselines under realistic NWP noise","feed_subtitle":"Controlled PV simulations show they filter dynamic forecast errors and shift weight to history and physical priors","key_machinery":"A simulation-based robustness framework that treats virtual PV power as a controlled response variable, injects dynamic NWP perturbations with clear-sky-modulated heteroscedasticity, and reconstructs radiation via the Erbs model so physical consistency is preserved while input-uncertainty propagation is isolated from plant-level confounders.","core_discovery":"Under physically constrained, dynamic NWP perturbations that preserve radiation consistency via Erbs reconstruction and clear-sky-modulated heteroscedasticity, sequence models (PatchTST, GRU, N-HITS) deliver stronger noise filtering and temporal resilience than LightGBM in medium-to-high disturbance regimes, with SHAP/IG showing case-level feature reallocation from corrupted future forecasts toward historical observations and deterministic physical priors.","pith_inferences":["The same controlled-simulation plus physically consistent perturbation recipe could be reused for wind or load forecasting where NWP errors are likewise correlated and state-dependent.","If the reallocation pattern is reliable, hybrid models that hard-wire clear-sky and historical channels may further improve robustness without large accuracy loss.","Operational monitoring could flag days when feature attribution drifts back toward corrupted future NWP as an early warning that the forecast should be down-weighted.","Latency-aware Pareto selection may push edge deployments toward lighter sequence variants once robustness thresholds are met."],"forward_implications":["Robustness rankings under realistic NWP error, not only clean-data accuracy, become a required dimension of model selection for operational PV forecasting.","Sequence architectures can be preferred when medium-to-high forecast disturbance is expected, while tabular baselines may remain competitive under low-noise regimes.","Explainability tools (SHAP/IG) can be used operationally to monitor whether a model is shifting reliance toward stable history and physical priors under degraded NWP.","Pareto trade-offs among clean accuracy, robustness, and latency supply an engineering checklist for choosing models under forecast uncertainty."],"fun_headline_variants":["Sequence models filter NWP noise better than LightGBM in PV forecasts","Deep sequence models show resilience to dynamic weather forecast errors","PatchTST and GRU beat tabular baselines under realistic NWP perturbations","Sequence nets shift weight to history under corrupted NWP inputs","Controlled PV sims show sequence models handle NWP heteroscedasticity"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That virtual PV power driven by clear-sky-modulated heteroscedastic NWP noise and Erbs radiation reconstruction is a faithful enough proxy for real NWP error structure and plant response that the robustness rankings transfer to operational forecasting.","fun_headline_variants_meta":{"raw":{"variants":["Sequence models filter NWP noise better than LightGBM in PV forecasts","Deep sequence models show resilience to dynamic weather forecast errors","PatchTST and GRU beat tabular baselines under realistic NWP perturbations","Sequence nets shift weight to history under corrupted NWP inputs","Controlled PV sims show sequence models handle NWP heteroscedasticity"]},"model":"grok-4.5","effort":"low","cost_usd":0.00303,"raw_usage":{"total_tokens":1075,"prompt_tokens":807,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":30300000,"prompt_tokens_details":{"text_tokens":807,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":201,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":807,"tokens_out":67,"duration_ms":3268,"temperature":1.0,"reasoning_tokens":201,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T02:03:04.706970+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Apply the same model suite and perturbation protocol to real multi-site PV plants with measured NWP error archives; if sequence models no longer outperform LightGBM on medium-to-high disturbance days or SHAP/IG no longer show the reallocation pattern, the central claim fails.","supporting_citations":[],"review_version":1}