{"id":"aca76af0-74b8-42f9-9a26-dad26163d21c","arxiv_id":"2411.13921","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"Neural Basis Models for Location, Scale and Shape (NBMLSS) achieve probabilistic electricity price forecast accuracy comparable to distributional neural networks, while revealing per-feature shape functions across the forecast horizon.","lead":"The paper introduces NBMLSS, a neural model for probabilistic electricity price forecasting that combines shared basis functions with distributional regression, and shows it matches the accuracy of black-box distributional neural networks while exposing how each input feature shapes the predicted price distribution. A generalist should read it because interpretable forecasting models can be competitive with opaque ones in a volatile, high-stakes market.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The interpretability claim is not yet supported: the paper's own figures show shape functions varying strongly across seeds, and no quantitative stability or identifiability analysis is provided, so the 'more insights' contribution remains fragile.","rationale":"The reader's verdict is CONDITIONAL, and the reader's weakest-assumption analysis identifies exactly the same load-bearing concern: the interpretability contribution depends on learned shape functions being stable and meaningful, yet the paper's own figures show substantial cross-seed variation. My stress-test agrees with that assessment. The accuracy claim is comparatively well supported: the tables show NBMLSS achieving CRPS values on par with or better than DDNN across multiple markets and test periods, and the paper includes recalibration, ensembles, and DM tests. No serious internal inconsistency appears in the forecasting methodology. The decisive weakness is the interpretability claim, which is the paper's primary novelty. The shape functions in Equations (1)-(3) are the vehicle for insight, but the paper demonstrates that different seeds produce quite different functions while yielding similar predictive distributions. This is textbook non-identifiability, here driven by concurvity and underspecification, and the authors acknowledge it. Consequently, the claim that NBMLSS provides 'more insights' into feature-to-distribution-parameter mechanisms is not yet established. The proposed synthetic test would settle whether the instability is inherent to the architecture or a data-level problem, and the real-data stability metric would quantify how much the displayed insights can be trusted. Because the reader's conditional recommendation already asks for a more rigorous treatment of shape-function stability, my read does not change the verdict; it reinforces it. I therefore recommend UNCHANGED, with the condition remaining that the authors provide code and a quantitative identifiability or stability analysis before the interpretability claim is accepted at face value.","tokens_in":14060,"tokens_out":5219,"duration_ms":58736,"concrete_test":"Construct a synthetic experiment with independent features generated from a known GAMLSS/JSU data-generating process whose marginal distributions match the DE market, then fit the same J-NBMLSS pipeline with the same hyperparameters and 15 seeds. Compare each recovered shape function f_i to the true feature effect using a normalized L2 metric across the feature support. If recovery error is small and stable across seeds, the real-data instability is attributable to feature concurvity and could be addressed by regularizers; if recovery is poor even without concurvity, the architecture itself does not identify shape functions and the interpretability claim should be downgraded. Apply the same stability metric to the real DE data shown in Figures 2, 3, and 7 to quantify the magnitude of the problem on the actual forecasting task.","verdict_should_be":"UNCHANGED","load_bearing_attack":"NBMLSS's central novelty is the claim of more insight via learned feature shape functions f_i(x_d,i) in Eqs. (1)-(3). For that claim to hold, each f_i must be a reliable representation of the conditional effect of feature i on a given distribution parameter. The paper's own evidence contradicts this: Figures 2, 3, and 7 show large differences in f_i across random recalibrations, while Figure 4 shows nearly identical predictive distributions produced by those heterogeneous functions. The authors explicitly attribute this to approximate concurvity and underspecification, citing D'Amour et al. and Siems et al., and they report 'sensible heterogeneity' even in the reduced-feature masked configuration. This is a non-identifiability problem: many combinations of shape functions yield equivalent forecasts, so the extracted maps do not uniquely describe the learned relationships. Without a quantitative stability metric, a synthetic identifiability check, or a regularizer that enforces unique shape functions, the 'more insights' claim is an aspiration rather than a demonstrated result. The accuracy-side claim, CRPS parity with DDNN, is not the weak point: Tables 1-8 support it, and the paper's limitations are honestly stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes NBMLSS, a neural basis model for distributional regression in day-ahead electricity price forecasting. It combines a shared basis-function network with linear projections to produce feature-specific shape functions that map inputs to the parameters of a parametric distribution (e.g., Johnson's SU). The model is evaluated against distributional neural networks on four European markets (BE, DE, ES, SE) over two out-of-sample test periods (TS23, TS24), with weekly recalibration, ensemble aggregation, and statistical tests. The authors report comparable or better CRPS for NBMLSS in several settings, and they inspect the learned shape functions as an interpretability analysis.","tokens_in":14297,"tokens_out":5882,"duration_ms":53348,"significance":"Strengths of the paper: a sound out-of-sample forecasting pipeline (multiple markets, two test periods, train/validation split, weekly recalibration, 5-member ensembles, Kupiec and Diebold-Mariano tests), honest reporting of limitations, and the use of open datasets. The accuracy-side claim is credible: Tables 1-8 support comparable or better CRPS for NBMLSS relative to DDNN in several configurations. The interpretability-side claim, which is the paper's main novelty, is not yet established quantitatively: the authors themselves show that shape functions vary strongly across random initializations (Figures 2, 3, 7) and attribute this to concurvity and underspecification. The forecasting contribution alone would be incremental; the interpretability contribution is potentially valuable but requires additional evidence to be convincing.","major_comments":[{"comment":"The central claim of \"more insights\" is not supported by the presented evidence. The paper shows that shape functions differ substantially across random recalibrations while the corresponding predictive distributions are nearly identical (Figure 4). The authors explicitly attribute this behavior to approximate concurvity and underspecification, citing refs [25]-[27]. This is a non-identifiability problem: many combinations of shape functions yield equivalent forecasts, so the extracted maps do not uniquely describe the learned relationships. Without a quantitative stability metric, a synthetic-data identifiability check, or a regularizer that enforces unique shape functions, the interpretability claim remains an aspiration. This issue is load-bearing because interpretability is the paper's primary contribution.","section":"§3, Figures 2-4 and 7"},{"comment":"The interpretability analysis is qualitative and restricted to one market (DE) and two distribution parameters (location and skewness). The masked-feature experiments were intended to mitigate the non-identifiability problem, but the paper reports that \"sensible heterogeneity\" remains in the shape functions (Figure 7). The paper does not demonstrate that the shape functions correspond to true conditional relationships (e.g., on simulated data with known ground truth) or provide any quantitative measure of agreement across runs. To support the claim, the authors should either: (i) report a seed-variability metric (e.g., variance or range of shape functions relative to their magnitude) for all features and parameters, (ii) validate shape-function recovery on synthetic data with known additive structure, or (iii) apply a concurvity regularizer and show that stable functions are obtained while maintaining CRPS.","section":"§3, Figures 5-7 and Table 9"}],"minor_comments":[{"comment":"The sentence describing the training loss contains a duplicated article: \"using the the negative log-likelihood\".","section":"Section 2"},{"comment":"The author list appears to include \"F. Nov\"; this is likely a typo for \"F. Ziel\" (Ziel is the usual co-author of the cited distributional neural network paper). Please verify.","section":"Reference [8]"},{"comment":"The title \"Conformai pid control\" should be \"Conformal PID control\".","section":"Reference [23]"},{"comment":"For J-DNN ES under TS24, the number of units is listed as \"762\", which is likely a typo for \"768\".","section":"Table A.11"},{"comment":"The phrase \"share basis\" in the conclusions should be \"shared basis\".","section":"Section 4"},{"comment":"In the interpretability discussion, \"renawable\" should be \"renewable\".","section":"Section 3"},{"comment":"The caption of Figure 4 should specify more precisely what is plotted (e.g., deciles vs. percentiles) for clarity.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The accuracy comparison is sound and the empirical setup is careful. The interpretability claim is central to the paper's novelty and is not yet supported; however, it is a fixable issue if the authors add the quantitative stability/identifiability analysis described in the major comments. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core contribution here is the combination of Neural Basis Models with the NAMLSS framework, applied to multi-horizon probabilistic electricity price forecasting across several European markets, with RevIN and ensemble aggregation. That specific combination is new, and the empirical work is careful: multiple markets, two out-of-sample periods, proper validation split, weekly recalibration, five random restarts per ensemble, Kupiec coverage tests, and Diebold-Mariano tests. The headline accuracy claim—NBMLSS matches DDNN on CRPS, sometimes better—is supported by the tables and is honestly hedged. I also give credit for the JSU parameterization results and for openly discussing the limitations of the approach.\n\nThe soft spot is exactly what the stress-test note flags. The interpretability story rests on the shape functions f_i being meaningful per-feature conditional effects. The paper's own figures (Figures 2, 3, 7) show substantial variation across random recalibrations, and the authors trace this to concurvity and underspecification. They even show that the heterogeneous functions produce nearly identical predictive distributions. That means the 'more insights' claim is not yet a demonstrated result. A quantitative stability metric, a synthetic identifiability check, or an explicit concurvity regularizer would begin to fix it. Without one, the interpretability contribution is an aspiration. This is a real caveat, not a fatal flaw: the forecasting accuracy comparison stands on its own.\n\nTwo smaller points. First, no code or data are released. The data are public via ENTSO-E, so the setup is reproducible in principle, but the lack of code slows validation. Second, the parameter search for NBMLSS used a reduced range for nu and nz, which is fine but makes the comparison to DDNN slightly less equal than claimed.\n\nOverall, this is a useful paper for the probabilistic energy forecasting community. A serious referee should engage with it, but the revision should require code and a more rigorous treatment of shape-function stability. I'd bring it to a reading group if anyone works on interpretability or distributional regression. I wouldn't cite it in my own work unless I needed a recent example of NBM-style distributional forecasting.\n\nRecommendation: send to peer review, with expectations of revision.","headline":"A solid, well-executed empirical application of neural basis models to probabilistic price forecasting, but the interpretability payoff is weaker than the framing suggests and needs quantitative support.","tokens_in":14818,"tokens_out":1394,"would_cite":false,"duration_ms":17035,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that NBMLSS, an interpretable neural basis model for distributional regression, matches or exceeds distributional neural networks in day-ahead electricity price forecasting while exposing how each input feature shapes the…","keywords":["Neural Networks","GAMLSS","Time series","Electricity price","Forecasting","Day-ahead market","Probabilistic forecasting","Neural Basis Models"],"falsifier":"Feed NBMLSS synthetic data generated from a known additive relationship between inputs and distribution parameters; if the recovered per-feature shape functions do not converge to the true generating functions as the training sample grows, the interpretability claim is falsified.","tokens_in":13855,"feed_emoji":"⚡","tokens_out":7398,"duration_ms":69932,"temperature":0.7,"pith_summary":"The paper asks whether an interpretable neural network for distributional regression can match a black-box distributional neural network in day-ahead electricity price forecasting. It introduces NBMLSS, a model that learns a small set of shared basis functions and combines them with linear projections so that every input feature gets its own shape function for each distribution parameter at each forecast horizon. Tested on four European markets over two out-of-sample years, NBMLSS achieves continuous ranked probability scores comparable to, and on several markets better than, distributional neural networks under the same settings. The authors read this as evidence that interpretable additive architectures do not have to give up accuracy, while cautioning that the learned shape functions vary across random initializations because of concurvity and underspecification.","feed_headline":"Transparent forecaster matches black-box electricity price models","feed_subtitle":"Per-feature shape maps stay readable while CRPS scores match distributional neural networks.","key_machinery":"The load-bearing object is the shared basis decomposition. A single two-layer network computes $n_z$ basis functions $z_k(x_{d,i})$ from each feature value $x_{d,i}$; a learned matrix $w_{i,k}$ blends these bases into a per-feature shape function $f_i(x_{d,i})$, and a second projection $v^p_{h,i}$ with bias $\\beta^p_h$ maps the shape functions into the $p$-th parameter of the predictive density at horizon $h$, passed through link functions such as Softplus for the scale and tailweight of the Johnson's SU density. Because the expensive nonlinear network is shared across all 147 features and only the linear projections are feature-specific, the GAMLSS-style additive interpretation becomes computationally practical for 24-hour multi-horizon forecasting.","core_discovery":"The central claim is that a Neural Basis Model for Location, Scale and Shape can reach the forecast quality of distributional neural networks while delivering per-feature, per-parameter, per-horizon shape functions that expose how inputs shape the predicted price distribution. The paper demonstrates this on Germany, Spain, Belgium, and Sweden, using the Johnson's SU parameterization and consistent training settings; in several cases NBMLSS improves on the DDNN's CRPS, particularly with Reversible Instance Normalization. The learned shape functions reveal economically sensible relations, such as load pushing the location parameter up at peak hours and renewable generation shrinking location while inversely affecting skewness. The authors also report that the shape functions are not unique: different random initializations produce functionally different but predictively equivalent maps, which they trace to approximate concurvity and underspecification.","pith_inferences":["A reader could infer that single-run shape functions should not be treated as causal estimates; a more reliable protocol is to report distributions of shape functions over ensemble members, or to apply concurvity regularization before interpretation.","The architecture is not price-specific, so the same shared-basis design could carry to other multi-horizon distributional forecasting problems, such as load or renewable generation, where GAMLSS-style interpretability is wanted.","The scalability claim invites a direct benchmark: comparing NBMLSS against a per-feature NAMLSS under identical data would quantify the speed-up that motivates the shared basis.","The masked-exogenous experiments suggest the full model uses cross-hour information beyond same-hour renewables and load; a targeted ablation could identify which of those cross-hour features actually carry predictive value and shrink the conditioning set."],"forward_implications":["Forecasters can choose NBMLSS over DDNN without expecting a systematic loss of CRPS, gaining per-feature shape functions for location, scale, tailweight, and skewness across the 24-hour horizon.","Reversible Instance Normalization should be considered a standard component for both NBMLSS and DDNN, since it lowers CRPS and MAE on average across the four markets and both test periods.","The Johnson's SU parameterization beats the Normal form in most volatile test conditions, so flexible distributional forms remain valuable even inside an additive architecture.","The shape functions provide concrete diagnostic information, such as load having a steeper positive effect on location in peak hours and renewable generation showing a shrinking influence on location with an inverse relation to skewness.","Ensembles of NBMLSS components, or hybrids combining NBMLSS with DDNN, present a practical route to combine interpretability with robustness, since individual shape functions are seed-dependent."],"supporting_citations":[{"why":"Supplies the DDNN baseline, ensemble aggregation scheme, and Johnson's SU parameterization that NBMLSS is compared against.","marker":"[8]"},{"why":"Introduces Neural Additive Models, the interpretable architecture class that NBMLSS builds on.","marker":"[11]"},{"why":"Extends NAMs to distributional regression as NAMLSS, giving the GAMLSS-style parameterization that NBMLSS scales up.","marker":"[12]"},{"why":"Guides the hour-wise masking experiment that isolates each feature's contribution to the JSU parameters.","marker":"[13]"},{"why":"Supplies the shared basis decomposition that lets one network serve all features instead of one network per feature.","marker":"[14]"},{"why":"Provides the open multi-market datasets and benchmark structure used for the TS23 and TS24 evaluations.","marker":"[15]"},{"why":"Introduces Reversible Instance Normalization, which the paper adds to both NBMLSS and DDNN to handle distribution shift.","marker":"[17]"},{"why":"Documents concurvity in differentiable GAMs, used by the paper to explain non-uniqueness of shape functions across initializations.","marker":"[25]"},{"why":"Establishes underspecification as a general machine-learning issue, used to motivate why different random initializations yield different but equally valid shape functions.","marker":"[27]"}],"fun_headline_variants":["Interpretable price forecaster rivals black-box neural nets","Neural basis model exposes price drivers, matches deep nets","NBMLSS: transparent forecasts, deep-net accuracy","See prices clearly: NBMLSS rivals deep nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the learned per-feature shape functions accurately represent the true conditional relationships between inputs and distribution parameters; the paper itself shows these maps vary strongly across random initializations, so if this assumption fails the interpretability benefit weakens even though forecast accuracy may still hold.","fun_headline_variants_meta":{"raw":{"variants":["Interpretable price forecaster rivals black-box neural nets","Neural basis model exposes price drivers, matches deep nets","NBMLSS: transparent forecasts, deep-net accuracy","See prices clearly: NBMLSS rivals deep nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00107,"raw_usage":{"total_tokens":4418,"prompt_tokens":817,"completion_tokens":3601,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":3537}},"tokens_in":433,"tokens_out":3601,"duration_ms":26535,"temperature":1.0,"reasoning_tokens":3537,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:43:31.749226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed NBMLSS synthetic data generated from a known additive relationship between inputs and distribution parameters; if the recovered per-feature shape functions do not converge to the true generating functions as the training sample grows, the interpretability claim is falsified.","supporting_citations":[{"cited_title":"Agarwal, L","cited_arxiv_id":null,"evidence_quote":"Introduces Neural Additive Models, the interpretable architecture class that NBMLSS builds on."},{"cited_title":"Thielmann, R.-M","cited_arxiv_id":null,"evidence_quote":"Extends NAMs to distributional regression as NAMLSS, giving the GAMLSS-style parameterization that NBMLSS scales up."},{"cited_title":"Radenovic, A","cited_arxiv_id":null,"evidence_quote":"Supplies the shared basis decomposition that lets one network serve all features instead of one network per feature."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces Reversible Instance Normalization, which the paper adds to both NBMLSS and DDNN to handle distribution shift."},{"cited_title":"Siems, K","cited_arxiv_id":null,"evidence_quote":"Documents concurvity in differentiable GAMs, used by the paper to explain non-uniqueness of shape functions across initializations."},{"cited_title":"D’Amour, K","cited_arxiv_id":null,"evidence_quote":"Establishes underspecification as a general machine-learning issue, used to motivate why different random initializations yield different but equally valid shape functions."}],"review_version":1}