REVIEW 3 major objections 3 minor
Where you put uncertainty in a smart-building load forecaster depends on the backbone: integrated quantile learning wins with the Temporal Fusion Transformer, while post-hoc residual quantiles fail when inputs are reconstructed.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 03:46 UTC pith:H6B36452
load-bearing objection Useful applied comparison of post-hoc vs in-model uncertainty under reconstructed inputs, but abstract-only so the backbone ranking stays unverifiable. the 3 major comments →
Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Uncertainty placement is backbone-dependent: when inference inputs are reconstructed, an integrated in-model quantile-learning scheme paired with a Temporal Fusion Transformer yields the most reliable one-day-ahead probabilistic load forecasts (2.2–3.6 % MAPE, 28–83 W RMSE, intervals ~5× narrower than modular residual quantiles at closest-to-nominal coverage), while post-hoc residual quantiles remain limited under the same reconstruction pipeline.
What carries the argument
A unified one-day-ahead probabilistic forecasting pipeline that (i) aligns temporal resolution, reconstructs unavailable inputs, and derives causal features, then (ii) contrasts modular post-hoc residual-quantile uncertainty against integrated in-model quantile learning across three mid-scale DL backbones under identical inputs, horizon, preprocessing, and training budgets.
Load-bearing premise
That equalizing inputs, horizon, preprocessing and training budgets across three mid-scale backbones, together with residual-quantile modularization as the sole post-hoc baseline, constitutes a fair and representative test of where uncertainty should sit once reconstruction is required.
What would settle it
Re-run the identical reconstruction-plus-forecasting pipeline on a new mid-scale building dataset; if the TFT with integrated quantiles no longer produces intervals ~5 imes narrower at closest-to-nominal coverage, or if a simple residual-quantile wrapper on a recurrent backbone matches or beats it, the placement claim fails.
If this is right
- Demand-response schedulers that receive reconstructed hourly inputs should prefer TFT-style models with built-in quantile heads rather than attaching residual quantiles after a point forecast.
- Post-hoc residual-quantile modules cannot be assumed to capture reconstruction-induced uncertainty; their intervals stay nearly constant while Quantile Score roughly doubles.
- Ranking of uncertainty schemes is backbone-dependent, so architecture choice and uncertainty placement must be co-optimized rather than treated as independent design decisions.
- Diebold-Mariano tests and seasonal hold-outs already support the TFT ranking, giving a concrete statistical criterion for future backbone comparisons under reconstruction.
Where Pith is reading between the lines
- The same reconstruction-sensitivity pattern is likely to appear in other cyber-physical forecasting domains (grid edge, EV charging, district heating) that train on dense SCADA and deploy on sparse telemetry.
- A natural next experiment is to replace residual quantiles with conformal or ensemble post-hoc methods under the identical reconstruction pipeline to test whether any modular scheme can close the gap to in-model quantiles.
- If reconstruction error dominates, end-to-end training that jointly optimizes the reconstructor and the quantile heads may further shrink the observed Quantile-Score inflation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies one-day-ahead probabilistic load forecasting for smart buildings under a deployment mismatch: models trained on dense multivariate high-frequency data must run on reconstructed, feature-limited hourly inputs. It presents a unified pipeline that aligns temporal resolution, reconstructs unavailable features, and derives causal inputs, then compares modular post-hoc residual quantiles with integrated in-model quantile learning on three mid-scale DL backbones (recurrent, hybrid recurrent, TFT) under matched inputs, horizon, preprocessing, and training budgets. Reported results claim backbone-dependent uncertainty placement: TFT with in-model quantiles yields 2.2–3.6% MAPE and 28–83 W RMSE with intervals ~5× narrower than modular residual quantiles at closest-to-nominal coverage; Diebold–Mariano tests support the ranking; a reconstruction-sensitivity test shows QS +106% with nearly unchanged width; and non-DL baselines plus seasonal hold-outs are said to corroborate the ranking.
Significance. If the matched multi-backbone comparison and the reconstruction-sensitivity result survive full methodological scrutiny, the paper would give practical guidance on where to place uncertainty when inference depends on reconstructed inputs—an operationally relevant issue for demand-response calibration. The backbone-dependent ranking and the QS-vs-width dissociation under reconstruction are concrete, falsifiable claims. Use of Diebold–Mariano tests, seasonal hold-outs, and non-DL baselines are appropriate strengths when fully documented. On the available abstract alone these contributions cannot be verified, so significance remains conditional on the missing methods, data splits, and calibration diagnostics.
major comments (3)
- Only the abstract is available for review. The central claim that uncertainty placement is backbone-dependent (TFT + in-model quantiles most reliable; modular residual quantiles limited under reconstruction) rests on matched training budgets, preprocessing, residual modularization, and the reconstruction pipeline. None of these can be audited from the abstract; without methods, tables, and calibration plots the ranking and the QS +106% sensitivity result remain unverifiable. Full manuscript is required before any accept/revise decision.
- Abstract claim of intervals “about 5× narrower … at the closest-to-nominal coverage level”: the precise coverage targets, calibration diagnostics (e.g., PICP/PINAW or reliability diagrams), and how “closest-to-nominal” is selected across schemes are not stated. This choice is load-bearing for the width comparison; if coverage is not matched on a common nominal level the 5× factor is not interpretable.
- Abstract reconstruction-sensitivity result (QS +106%, width nearly unchanged): whether the reconstruction model is trained only on training data, whether residual quantiles see reconstruction error at training time, and whether any leakage of deployment features into the reconstruction step exists cannot be checked. These details determine whether the experiment actually isolates reconstruction-induced uncertainty or confounds it with capacity/training differences.
minor comments (3)
- Abstract: report the dataset identity, number of buildings/meters, sampling rate, and exact forecast horizon (e.g., 24 hourly steps) so the MAPE/RMSE ranges can be contextualized.
- Abstract: name the three backbones more explicitly (architecture families and approximate parameter counts) and state whether residual quantiles are fitted on the same held-out residual distribution for all three.
- Abstract: clarify the non-DL baselines used in the robustness checks (e.g., ARIMA/quantile regression/gradient boosting) so the ranking claim can be scoped.
Circularity Check
No significant circularity: empirical backbone comparison with held-out metrics; no self-definitional or fitted-as-prediction reductions visible in the abstract.
full rationale
This is an abstract-only empirical ML study comparing modular post-hoc residual quantiles versus integrated in-model quantile learning across three DL backbones (recurrent, hybrid recurrent, TFT) under matched inputs, horizon, preprocessing, and training budgets, with reconstruction of missing features. The load-bearing claims are experimental rankings (MAPE/RMSE, interval width and coverage, Diebold–Mariano tests, reconstruction-sensitivity QS +106% with nearly unchanged width, non-DL and seasonal robustness checks). Nothing in the abstract defines a quantity in terms of the reported result, renames a fit as a prediction, imports a uniqueness theorem from the authors, or smuggles an ansatz via self-citation. Mild design-choice risk (author-chosen budgets/preprocessing/reconstruction) is a fairness/verification concern, not circularity under the enumerated patterns. With only the abstract available, no equation-level reduction to inputs can be exhibited; the honest finding is score 0 with empty steps.
Axiom & Free-Parameter Ledger
free parameters (2)
- Training budgets / hyperparameters of the three DL backbones
- Reconstruction model parameters for missing deployment features
axioms (3)
- domain assumption Deployment inputs are feature-limited and temporally coarser than training data, so missing features must be reconstructed and can inject error into the forecaster.
- ad hoc to paper Post-hoc residual quantiles and in-model quantile learning are the two relevant uncertainty-placement schemes to compare under identical inputs and horizon.
- domain assumption Standard probabilistic scoring (MAPE, RMSE, Quantile Score, coverage/width, Diebold–Mariano) is adequate to rank the schemes.
read the original abstract
Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs. Missing features must then be reconstructed, and their errors can propagate through the model. If this input uncertainty is not reflected, prediction intervals may become miscalibrated, affecting demand-response scheduling. Our work examines where uncertainty should be placed once inference inputs are reconstructed. We develop a unified one-day-ahead probabilistic forecasting framework that aligns temporal resolution, reconstructs the unavailable inputs, and derives causal features, and we compare a modular post-hoc residual-quantile scheme with an integrated in-model quantile-learning scheme. The comparison uses three mid-scale Deep Learning (DL) backbones: recurrent, hybrid recurrent, and attention-based Temporal Fusion Transformer (TFT) models, under identical inputs, forecasting horizon, preprocessing rules, and training budgets. Results show that uncertainty placement is backbone-dependent. Integrated quantile learning is most reliable with the TFT, yielding 2.2-3.6% MAPE and 28-83W RMSE on the labeled test window, while producing intervals about 5x narrower than the modular intervals at the closest-to-nominal coverage level. Diebold-Mariano tests support the TFT ranking and the mixed behavior of the recurrent backbones. A reconstruction-sensitivity test shows that reconstructed inputs increase the Quantile Score (QS) by 106% while interval width remains nearly unchanged, indicating that the model does not automatically absorb reconstruction-induced uncertainty. Robustness checks against non-DL baselines and seasonal hold-out weeks support this ranking. Our results expose the limits of post-hoc residual quantiles when inference depends on reconstructed inputs.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.