REVIEW 2 major objections 4 minor 69 references
Nonparametric regression models can be embedded exactly into stochastic programs whose uncertainty depends on both decisions and covariates, with consistency guarantees and a finite-convergence algorithm for kNN.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
ER-DD-SAA with exact MIP embeddings of kNN/CART/ReLU NNs is consistent and asymptotically optimal, and BD-CG solves the kNN two-stage case to global optimality in finite iterations.
T0 review reviewed 2026-07-14 challenge →
load-bearing objection Solid, usable MIP encodings and a proved Benders+CG algorithm for residual-based decision-dependent SAA with nonparametric regressors; theory is standard but clean, experiments synthetic-only. the 2 major comments →
Contextual Stochastic Optimization with Decision-Dependent Uncertainty via Nonparametric Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The authors show that an empirical-residuals decision-dependent sample-average approximation (ER-DD-SAA) built from kNN, CART or ReLU-network regressors is statistically consistent and asymptotically optimal for a general decision-dependent contextual stochastic program, and that the resulting mixed-integer programs can be solved to global optimality—via a finite-convergent Benders-constraint-generation algorithm when the regressor is kNN and the problem is two-stage.
What carries the argument
ER-DD-SAA: form decision-dependent scenarios by adding empirical residuals of a trained nonparametric regressor to its point prediction at the candidate decision and observed covariate, then embed that construction inside a mixed-integer program (pairwise-distance or compact bilevel form for kNN; leaf-indicator form for CART; ReLU big-M form for networks).
Load-bearing premise
Uncertainty is assumed to be a true regression function of decision and covariate plus additive zero-mean noise, and the nonparametric estimator must converge uniformly over the whole decision-covariate domain.
What would settle it
On a problem whose demand truly depends on both price (or facility openings) and a covariate, check whether out-of-sample cost of ER-DD-SAA with a nonparametric regressor fails to improve on a correctly specified parametric model, or whether the BD-CG algorithm fails to close the gap to global optimality within a finite iteration budget on moderate kNN instances.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies decision-dependent contextual stochastic programs (DD-CSP) in which the uncertainty Y depends jointly on endogenous decisions z and exogenous covariates w. It approximates the problem by an empirical residuals-based decision-dependent SAA (ER-DD-SAA) that embeds nonparametric regressors (kNN, CART, ReLU NNs) trained on historical data and adds empirical residuals to the point predictions. Exact MIP/MINLP reformulations are derived for each regressor (pairwise-distance and compact bilevel forms for kNN; leaf-selection for CART; standard big-M ReLU encoding). For two-stage problems with kNN a Benders-plus-constraint-generation algorithm (BD-CG) is proposed and proved to reach a global optimum in finitely many iterations under relatively complete and expensive recourse. Consistency and asymptotic optimality of ER-DD-SAA are established under Lipschitz costs, weak LLN, and uniform consistency of the three regressors. Synthetic experiments on a pricing newsvendor and a two-stage facility-location problem show improved out-of-sample performance relative to linear regression and improved tractability of the proposed reformulations and BD-CG.
Significance. The work fills a clear gap between contextual stochastic optimization and decision-dependent uncertainty by giving a unified residual-based SAA framework that accommodates continuous and discrete decisions and three standard nonparametric models. The exact MIP encodings (especially the compact bilevel kNN form) and the finite-convergence BD-CG algorithm are concrete algorithmic contributions that make the approach usable beyond pure continuous settings. The statistical guarantees (Theorem 2) rest on classical uniform-consistency results for kNN/CART/ReLU and standard residual-SAA transfer arguments; they are correctly scoped and non-circular. The synthetic experiments, while limited in scope, illustrate both the statistical and computational claims. Overall the paper is a solid, self-contained contribution that should be of interest to the data-driven optimization community.
major comments (2)
- The additive-error model Y = Q*(z,w) + ε with zero-mean noise (Section 2.1) and the uniform consistency Assumption 5(i) are the load-bearing statistical hypotheses. Theorems 5–8 correctly import known results, but the paper never discusses how sensitive the finite-sample OOS gains in Tables 3 and 6 are to misspecification of this additive structure (e.g., multiplicative or heteroscedastic noise). A short remark or a single misspecified experiment would strengthen the claim that nonparametric ER-DD-SAA is robust beyond the synthetic DGPs of Section 6.
- All numerical evidence is synthetic and generated from the same functional forms used to motivate the models (Eqs. (19) and (21)). While this is standard for methodological papers, the absence of any real-data instance or public code makes it harder to assess practical impact. At minimum the authors should release the synthetic generators and the MIP formulations so that the computational claims (Tables 4–5) can be reproduced.
minor comments (4)
- Table 1 claims that ReLU NNs yield an MINLP under objective uncertainty; this is correct for the newsvendor (bilinear p·ŷ), but the text could note more clearly that the same model remains MILP under pure RHS uncertainty (as later confirmed by Model (34)).
- In Algorithm 1 the termination test F^m ⊆ S^m is slightly weaker than equality when ties occur; a one-sentence clarification that the objective is still optimal under the objective-driven tie-breaking of Section 3.1.1 would help.
- Several big-M constants (M_1^{ij}, M_2, δ) are stated to be “sufficiently large/small”; giving the explicit tight values already used in the text (triangle inequality, N-k) in a single place would improve reproducibility.
- Typos: “Bender’s” should be “Benders’” throughout; “locaion” in Appendix D.2; arXiv date 13 Jul 2026 appears to be a placeholder.
Circularity Check
No significant circularity: consistency and finite convergence rest on standard residual-SAA transfer plus classical nonparametric uniform-consistency results; self-citation supplies the ER-DD-SAA template but is not load-bearing for the new claims.
specific steps
-
self citation load bearing
[Section 5, Theorem 2 and surrounding text]
"The proof of Theorem 2 mainly follows from Theorem 1 in Sun et al. [2026] and we omit it here."
The residual-SAA transfer argument that converts uniform consistency of Q̂_N into consistency of the ER-DD-SAA value/solution sets is imported from the authors' own prior paper rather than re-derived. This is a mild self-citation of a modeling template; the present paper still supplies independent content (MIP encodings, BD-CG, and verification of Assumption 5 for three nonparametric models via external classical results). It is not load-bearing for the new algorithmic or formulation claims and does not force the target result by definition.
full rationale
The paper's central theoretical claims are Theorem 2 (consistency and asymptotic optimality of ER-DD-SAA under Assumptions 3–5 for kNN/CART/ReLU) and Theorem 1 (finite global convergence of BD-CG under Assumptions 1–2). Theorem 2 is an asymptotic transfer result: under Lipschitz costs, weak LLN, and uniform consistency of the nonparametric estimator (Assumption 5), the ER-DD-SAA value and solution sets converge in probability to those of the true DD-CSP. The proof is deferred to Sun et al. [2026, Theorem 1] for the residual-SAA transfer step, while the paper itself verifies Assumption 5 for the three estimators via classical external results (Biau–Devroye for kNN, Bertsimas–McCord for CART, Imaizumi for ReLU NNs) restated as Theorems 5–8. Those citations are independent of the present authors' prior work and do not encode the target optimization result. Theorem 1 is a self-contained finite-convergence argument for a Benders + constraint-generation scheme on a finite pool of pairwise-distance and dual-extreme-point cuts; it does not rely on fitted parameters or self-citation. Hyperparameters (k, tree depth, network width) are chosen by cross-validation and are outside the consistency theorems. The only self-citation of note is the ER-DD-SAA template from Sun et al. [2026]; it is used as a modeling framework, not as a uniqueness theorem that forces the present conclusions. Numerical experiments use synthetic DGPs that match the additive-error model, so they do not create a fitted-input-called-prediction loop. Overall the derivation chain is non-circular; score 1 reflects a single non-load-bearing self-citation of the authors' prior application paper.
Axiom & Free-Parameter Ledger
free parameters (3)
- k (number of neighbors)
- CART max_depth / min_samples_split / min_samples_leaf
- ReLU network architecture (depth, width, alpha, solver)
axioms (5)
- domain assumption Additive error model Y = Q*(z,w) + ε with E[ε]=0
- domain assumption Z compact and LP/MILP-representable; Proj_Y MILP-representable
- domain assumption Relatively complete and sufficiently expensive recourse (Assumptions 1–2)
- standard math Lipschitz continuity of c(z,·) uniformly in z (Assumption 3) and weak LLN + continuity + domination (Assumption 4)
- domain assumption Uniform consistency of the three nonparametric estimators over Z×W (Assumption 5)
Cite this review
Pith. "Pith review of Contextual Stochastic Optimization with Decision-Dependent Uncertainty via Nonparametric Learning." pith.science (2026). https://pith.science/paper/IYQ4NE3Z
@misc{pith2026260711714,
author = {Pith},
title = {Pith review of: Contextual Stochastic Optimization with Decision-Dependent Uncertainty via Nonparametric Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/IYQ4NE3Z}},
note = {Machine review of arXiv:2607.11714}
}
read the original abstract
We study a general decision-dependent contextual stochastic program (DD-CSP) in which uncertainty depends on both exogenous contextual information and endogenous decisions. To learn the potentially complex dependence of uncertainty on decisions and contextual information, we employ several nonparametric regression models, including k nearest neighbors (kNN), classification and regression trees (CART), and ReLU neural networks. To account for estimation errors in predicting the uncertainty, we adopt an empirical residuals-based decision-dependent sample average approximation (ER-DD-SAA) framework, which adds empirical residuals to the point predictions from the learned regression models. For each nonparametric regression model, we develop exact mixed-integer programming (MIP) representations that can be seamlessly embedded within the ER-DD-SAA framework. For two-stage ER-DD-SAA problems with kNN, we further propose a tailored decomposition algorithm, named BD-CG, that combines Bender's decomposition with constraint generation. Under suitable assumptions, we prove that the proposed BD-CG converges to a global optimum within a finite number of iterations. From a statistical perspective, we establish the consistency and asymptotic optimality of ER-DD-SAA with all three nonparametric regression models under mild regularity conditions. Numerical experiments on a newsvendor problem with pricing and a two-stage facility location problem demonstrate that the ER-DD-SAA model with nonparametric learning consistently outperforms a parametric benchmark in out-of-sample performance and the proposed reformulations and algorithm substantially improve computational tractability.
Reference graph
Works this paper leans on
-
[1]
Journal of Operations Research , volume =
Smith, John , title =. Journal of Operations Research , volume =
-
[2]
INFORMS Mathematics of Operations Research , volume =
Jones, Sarah , title =. INFORMS Mathematics of Operations Research , volume =
-
[3]
Brown, David , title =
-
[4]
Multistage distributionally robust mixed-integer programming with decision-dependent moment-based ambiguity sets , volume =
Yu, Xian and Shen, Siqian , date-added =. Multistage distributionally robust mixed-integer programming with decision-dependent moment-based ambiguity sets , volume =. Mathematical Programming , number =
-
[5]
Nonparametric regression using deep neural networks with ReLU activation function , author=
-
[6]
Production and Operations Management , volume=
Learning Newsvendor Problems With Intertemporal Dependence and Moderate Non-stationarities , author=. Production and Operations Management , volume=. 2024 , publisher=
2024
-
[7]
Management Science , volume =
Qi, Meng and Cao, Ying and Shen, Zuo-Jun (Max) , title =. Management Science , volume =
-
[8]
arXiv preprint arXiv:2406.20004 , year=
Residuals-Based Contextual Distributionally Robust Optimization with Decision-Dependent Uncertainty , author=. arXiv preprint arXiv:2406.20004 , year=
-
[9]
INFORMS Journal on Computing , year=
Production planning under demand and endogenous supply uncertainty , author=. INFORMS Journal on Computing , year=
-
[10]
University of West Bohemia in Pilsen , year=
Optimization under exogenous and endogenous uncertainty , author=. University of West Bohemia in Pilsen , year=
-
[11]
Operations Research , volume=
Strong optimal classification trees , author=. Operations Research , volume=. 2025 , publisher=
2025
-
[12]
Machine Learning , volume=
Optimal classification trees , author=. Machine Learning , volume=. 2017 , publisher=
2017
-
[13]
Mathematical Programming , volume=
Strong mixed-integer programming formulations for trained neural networks , author=. Mathematical Programming , volume=. 2020 , publisher=
2020
-
[14]
arXiv preprint arXiv:2304.13924 , year=
Solving data-driven newsvendor pricing problems with decision-dependent effect , author=. arXiv preprint arXiv:2304.13924 , year=
-
[15]
arXiv preprint arXiv:2505.01985 , year=
Optimization over trained (and sparse) neural networks: A surrogate within a surrogate , author=. arXiv preprint arXiv:2505.01985 , year=
-
[16]
International Conference on the Integration of Constraint Programming, Artificial Intelligence, and Operations Research , pages=
Optimization over trained neural networks: Taking a relaxing walk , author=. International Conference on the Integration of Constraint Programming, Artificial Intelligence, and Operations Research , pages=. 2024 , organization=
2024
-
[17]
The Annals of Probability , volume=
A maximal inequality and dependent strong laws , author=. The Annals of Probability , volume=. 1975 , publisher=
1975
-
[18]
Computational tradeoffs of optimization-based bound tightening in
Badilla, Fabian and Goycoolea, Marcos and Mu. Computational tradeoffs of optimization-based bound tightening in. arXiv preprint arXiv:2312.16699 , year=
-
[19]
arXiv preprint arXiv:2512.24295 , year=
Optimization over Trained Neural Networks: Going Large with Gradient-Based Algorithms , author=. arXiv preprint arXiv:2512.24295 , year=
-
[20]
arXiv preprint arXiv:2307.04042 , year=
Sup-norm convergence of deep neural network estimator for nonparametric regression by adversarial training , author=. arXiv preprint arXiv:2307.04042 , year=
-
[21]
Stochastic optimization of power market forecast using non-parametric regression models , year=
Shenoy, Saahil and Gorinevsky, Dimitry , booktitle=. Stochastic optimization of power market forecast using non-parametric regression models , year=
-
[22]
INFORMS Journal on Computing , volume=
A simulation optimization approach for the appointment scheduling problem with decision-dependent uncertainties , author=. INFORMS Journal on Computing , volume=. 2022 , publisher=
2022
-
[23]
European Journal of Operational Research , volume=
Distributionally robust facility location problem under decision-dependent stochastic demand , author=. European Journal of Operational Research , volume=. 2021 , publisher=
2021
-
[24]
European Journal of Operational Research , volume=
Robust alternative fuel refueling station location problem with routing under decision-dependent flow uncertainty , author=. European Journal of Operational Research , volume=. 2023 , publisher=
2023
-
[25]
Adaptive Distributionally Robust Planning for Renewable-Powered Fast Charging Stations Under Decision-Dependent
Li, Yujia and Qiu, Feng and Chen, Yixuan and Hou, Yunhe , journal=. Adaptive Distributionally Robust Planning for Renewable-Powered Fast Charging Stations Under Decision-Dependent. 2024 , publisher=
2024
-
[26]
Mathematics of Operations Research , volume=
Coupled learning enabled stochastic programming with endogenous uncertainty , author=. Mathematics of Operations Research , volume=. 2022 , publisher=
2022
-
[27]
Management Science , volume=
Distribution-free algorithms for learning enabled optimization with non-parametric estimation , author=. Management Science , volume=
-
[28]
A Data-Driven Methodology for Contextual Unit Commitment Using Regression Residuals , year=
Yurdakul, Ogun and Bayraksan, Güzin , journal=. A Data-Driven Methodology for Contextual Unit Commitment Using Regression Residuals , year=
-
[29]
Mathematical Programming , volume=
Residuals-based distributionally robust optimization with covariate information , author=. Mathematical Programming , volume=. 2024 , publisher=
2024
-
[30]
Computational Management Science , pages=
Predictive stochastic programming , author=. Computational Management Science , pages=. 2022 , publisher=
2022
-
[31]
Manufacturing & Service Operations Management , volume=
Dynamic procurement of new products with covariate information: The residual tree method , author=. Manufacturing & Service Operations Management , volume=. 2019 , publisher=
2019
-
[32]
Operations Research , volume=
The big data newsvendor: Practical insights from machine learning , author=. Operations Research , volume=. 2019 , publisher=
2019
-
[33]
Data-driven optimization: A reproducing kernel
Bertsimas, Dimitris and Koduri, Nihal , journal=. Data-driven optimization: A reproducing kernel. 2022 , publisher=
2022
-
[34]
Management Science , volume=
Optimal robust policy for feature-based newsvendor , author=. Management Science , volume=. 2024 , publisher=
2024
-
[35]
2023 , publisher=
Dynamic optimization with side information , author=. 2023 , publisher=
2023
-
[36]
Predict, then Optimize
Smart “Predict, then Optimize” , author=. Management Science , volume=. 2022 , publisher=
2022
-
[37]
Advances in Neural Information Processing Systems , volume=
Generalization bounds in the predict-then-optimize framework , author=. Advances in Neural Information Processing Systems , volume=
-
[38]
INFORMS Journal on Optimization , volume=
Smart predict-then-optimize for two-stage linear programs with side information , author=. INFORMS Journal on Optimization , volume=. 2023 , publisher=
2023
-
[39]
International Conference on Machine Learning , pages=
DNNR: Differential nearest neighbors regression , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[40]
Tutorials in Operations Research: Emerging and Impactful Topics in Operations , pages=
Integrating prediction/estimation and optimization with applications in operations management , author=. Tutorials in Operations Research: Emerging and Impactful Topics in Operations , pages=. 2022 , publisher=
2022
-
[41]
European Journal of Operational Research , volume=
A survey of contextual optimization methods for decision-making under uncertainty , author=. European Journal of Operational Research , volume=
-
[42]
The American Statistician , volume=
An introduction to kernel and nearest-neighbor nonparametric regression , author=. The American Statistician , volume=. 1992 , publisher=
1992
-
[43]
arXiv preprint arXiv:2505.07298 , year=
Adaptive Learning-based Surrogate Method for Stochastic Programs with Implicitly Decision-dependent Uncertainty , author=. arXiv preprint arXiv:2505.07298 , year=
-
[44]
2016 , publisher=
Deep learning , author=. 2016 , publisher=
2016
-
[45]
2017 , publisher=
Classification and regression trees , author=. 2017 , publisher=
2017
-
[46]
arXiv preprint arXiv:1904.11637 , year=
From predictions to prescriptions in multistage optimization problems , author=. arXiv preprint arXiv:1904.11637 , year=
Pith/arXiv arXiv 1904
-
[47]
IEEE Transactions on Information Theory , year=
Uniform convergence of deep neural networks with Lipschitz continuous activation functions and variable widths , author=. IEEE Transactions on Information Theory , year=
-
[48]
Mathematical Programming , volume=
Computability of global solutions to factorable nonconvex programs: Part I—Convex underestimating problems , author=. Mathematical Programming , volume=. 1976 , publisher=
1976
-
[49]
Operations Research , note=
Data-driven sample average approximation with covariate information , author=. Operations Research , note=
-
[50]
arXiv preprint arXiv:2602.07286 , year=
Solving contextual chance-constrained programming under decision-dependent uncertainty , author=. arXiv preprint arXiv:2602.07286 , year=
-
[51]
Proceedings of the 27th international conference on machine learning (ICML-10) , pages=
Rectified linear units improve restricted boltzmann machines , author=. Proceedings of the 27th international conference on machine learning (ICML-10) , pages=
-
[52]
Management Science , volume=
From predictive to prescriptive analytics , author=. Management Science , volume=. 2020 , publisher=
2020
-
[53]
2015 , publisher=
Lectures on the nearest neighbor method , author=. 2015 , publisher=
2015
-
[54]
Transportation Research Part B: Methodological , volume=
Contextual stochastic optimization for determining electric vehicle charging station locations with decision-dependent demand learning , author=. Transportation Research Part B: Methodological , volume=. 2026 , publisher=
2026
-
[55]
2021 , publisher=
Lectures on stochastic programming: modeling and theory , author=. 2021 , publisher=
2021
-
[56]
Pattern Recognition , volume=
On Euclidean norm approximations , author=. Pattern Recognition , volume=. 2011 , publisher=
2011
-
[57]
Gurobi Machine Learning , year =
-
[58]
Ohio Supercomputer Center , year =
-
[59]
Optimization Online, URL: https://optimization-online
Statistical Inference of Contextual Stochastic Optimization with Endogenous Uncertainty , author=. Optimization Online, URL: https://optimization-online. org/2021/1 0/8634 , year=
2021
-
[60]
1998 , publisher=
Theory of linear and integer programming , author=. 1998 , publisher=
1998
-
[61]
1990 , publisher=
Applied nonparametric regression , author=. 1990 , publisher=
1990
-
[62]
Dimensionality reduction with unsupervised nearest neighbors , pages=
K-nearest neighbors , author=. Dimensionality reduction with unsupervised nearest neighbors , pages=. 2013 , publisher=
2013
-
[63]
The annals of statistics , pages=
Consistent nonparametric regression , author=. The annals of statistics , pages=. 1977 , publisher=
1977
-
[64]
Annual meeting of the society for academic emergency medicine in San Francisco, California , volume=
An introduction to classification and regression tree (CART) analysis , author=. Annual meeting of the society for academic emergency medicine in San Francisco, California , volume=
-
[65]
Wiley interdisciplinary reviews: data mining and knowledge discovery , volume=
Classification and regression trees , author=. Wiley interdisciplinary reviews: data mining and knowledge discovery , volume=. 2011 , publisher=
2011
-
[66]
Ecology , volume=
Classification and regression trees: a powerful yet simple technique for ecological data analysis , author=. Ecology , volume=. 2000 , publisher=
2000
-
[67]
arXiv preprint arXiv:1803.08375 , year=
Deep learning using rectified linear units (relu) , author=. arXiv preprint arXiv:1803.08375 , year=
-
[68]
2018 , publisher=
An introduction to neural networks , author=. 2018 , publisher=
2018
-
[69]
2012 , publisher=
Neural networks: an introduction , author=. 2012 , publisher=
2012
This paper was first reviewed by grok-4.5 on July 14, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.