{"id":"b8af6d34-21df-4e4c-b14b-77fe5383e2ad","arxiv_id":"2412.13159","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A conformal bias-correction step (global training with local calibration) is proposed to improve quantile predictions for newsvendor decisions under model misspecification, with claimed conditional coverage guarantees that lack proofs.","lead":"This paper applies conformal prediction to the feature-based newsvendor inventory problem, adding a bias-correction step from held-out data to any quantile forecasting model. It reports large measured loss reductions on simulated and bike-sharing data, but its core new theory is stated without proofs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Conditional coverage rests on a deterministic gap-function bound (Assumption 1) that standard quantile learners fail with positive probability; the paper never controls this failure event.","rationale":"The reader's weakest assumption correctly identifies Assumptions 1 and 2 as the load-bearing supports of the conditional guarantees. My concern sharpens this: Assumption 1 is not merely unverified for the real data; it is stated as a deterministic inequality about a random estimator, and such inequalities fail with positive probability for essentially any non-trivial learner. The paper's own citation for kappa is a high-probability bound, which cannot be converted into an almost-sure bound without additional structure. Hence Theorems 3 and 4 are conditional on a random event whose probability is not controlled, and the practical, data-driven estimation of kappa in Appendix A uses a loss-difference proxy that is not valid for the pinball loss. This is a concrete correctness risk in the paper's central theoretical contribution, not merely a missing empirical check. Theorem 1 itself is a standard split-conformal result and appears correct, but the paper's advertised 'conditional guarantees' and the quality-quantity trade-off results (Theorems 3-5, Propositions 1-3) all rely on the problematic Assumption 1. The reader's REJECT verdict is therefore unchanged: the central novel claims are not established as stated.","tokens_in":23367,"tokens_out":18393,"duration_ms":167933,"concrete_test":"Simulate the well-specified linear quantile regression model Y_i = x_i^T theta* + epsilon_i with standard normal epsilon, n1=500, d=10, x_i ~ U[0,1]^d. For 10,000 replications, compute M = max_{i,j on a held-out grid} |(qhat(x_i)-q*(x_i))-(qhat(x_j)-q*(x_j))|. Record the fraction of replications where M exceeds C sqrt(xi^nu/n1) for, say, C=1, nu=1 and xi equal to the grid diameter. If that fraction is not negligible (e.g., >5%), Assumption 1 is not a property of the estimator, so Theorem 3's conditional guarantee does not apply to the algorithm used in the experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of rigorous conditional coverage is only as strong as Assumption 1, which states that for every realization of the n1 training points and every x1,x2, |(qhat_{n1,alpha}(x1)-q*_alpha(x1))-(qhat_{n1,alpha}(x2)-q*_alpha(x2))| <= kappa(n1, xi(x1,x2)). But qhat is a random function; unless the estimator is exactly Lipschitz with a fixed constant, this inequality cannot hold for all realizations. The cited Pan-Zhou result is only high-probability, not almost sure, and does not imply the existence of a deterministic kappa. Consequently Theorems 3 and 4 hold only on the event that Assumption 1 is satisfied; the paper gives no bound on P(Assumption 1 fails). The abstract's claim of guarantees 'independent of the correctness of the underlying model' is therefore not established. The paper itself acknowledges (Section 4.2) that kappa and the margin functions are unknown in practice and defers to Appendix A, but the loss-based estimator of kappa in Appendix A does not control the quantile error: the pinball loss is only Lipschitz, not strongly convex, so small loss differences can accompany large quantile shifts when the conditional density is small. Thus the conditional guarantees are not merely unvalidated on the bike-sharing data; their hypotheses are not satisfied with probability one for standard learners.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conformalized quantile prediction framework (CQPC/GTLC) for feature-based newsvendor problems under model misspecification. The method splits data into a training set and a calibration set, computes conformity scores as residuals of a base quantile predictor, and adds an empirical quantile of calibration residuals to the base prediction. The paper claims an unconditional marginal coverage guarantee (Theorem 1), conditional coverage guarantees under a gap-function assumption and a margin condition (Theorems 2--4), a confidence interval for the true conditional quantile (Theorem 4), and a data-pooling strategy that trades off data quality and quantity (Section 5). Numerical experiments on simulated data and the Capital Bikeshare dataset report consistent reductions in empirical pinball loss.","tokens_in":23710,"tokens_out":5112,"duration_ms":49455,"significance":"The problem is well-motivated: feature-based newsvendor decisions are sensitive to model misspecification, and conformal calibration offers a plausible remedy. The GTLC algorithm is simple, computationally attractive, and empirically effective in the reported experiments. If the conditional coverage guarantees were rigorously established, the paper would be a useful contribution to data-driven decision-making under misspecification. The strengths are the clean algorithmic idea, the use of a standard split-conformal mechanism, and the breadth of numerical comparisons across four quantile-regression algorithms. However, the current manuscript does not substantiate its central theoretical claims: the proofs are absent, the key gap-function assumption is not satisfied by standard learners with probability one, and the proposed estimator for that assumption is not valid. As written, the conditional-guarantee results rest on unverified structural conditions, so the paper's main novelty is not yet established.","major_comments":[{"comment":"Assumption 1 asserts a deterministic inequality, holding for every pair (x1,x2) in X, between the prediction errors of the sample-dependent quantile estimator qhat_{n1,alpha}. For any standard quantile-regression learner, the left-hand side is a random function of the training data; no finite-sample concentration result of the cited Pan--Zhou type (which is high-probability, not almost sure) implies that such an inequality holds for all realizations. Consequently, Theorems 3 and 4 hold only on the event that Assumption 1 is satisfied, and the paper provides no bound on the probability of that event's failure. This undermines the claimed guarantee that the conformalized quantile is valid 'independent of the correctness of the underlying model.' The assumption should be reformulated as a high-probability condition with the failure probability entering the bounds, or the authors should prove that a specific estimator class satisfies the deterministic inequality.","section":"§4.2, Assumption 1"},{"comment":"The main theoretical results other than Theorem 1 are stated without proofs. The manuscript contains no proof section; Appendices A and B are devoted to the data-driven selection method and additional numerical results. In particular, the text near Proposition 1 refers to 'the proof' without providing it, and the derivations of the bound phi(Delta,B) in Theorem 3 and of the two-approximation result in Proposition 2 are not shown. A reader cannot verify the central claims of the paper in its current form, which is a load-bearing deficiency for a journal submission.","section":"§4.2--§5.1"},{"comment":"The proposed estimator of the gap function kappa approximates the difference in quantile prediction errors by the difference in empirical pinball losses, justified by the Lipschitz continuity of L. This justification is not sufficient: the pinball loss is Lipschitz but not strongly convex, so a small loss difference can accompany a large shift in the optimal quantile when the conditional density near the quantile is small. Thus the estimated kappa is not an upper bound for the left-hand side of Assumption 1, and the data-driven selection of the pooling diameter using Equation (8) has no theoretical support. At minimum, the authors need to prove a quantitative relation between loss gaps and quantile gaps under Assumption 2, or replace the estimator with one that controls the quantity in Assumption 1 directly.","section":"Appendix A, κ Estimation"},{"comment":"The specific form kappa(n1,xi)=C sqrt(xi^nu/n1) in Equation (7) is introduced as 'reasonable' without derivation, and the constants C and nu are free parameters that are not estimated or validated against any quantile-regression estimator. In Proposition 1, the additional assumptions n1(B_xi)=rho n xi^iota and n2(B_xi)=(1-rho)n xi^iota are ad hoc and not justified by any metric-space structure or sampling model. The claimed optimal pooling diameter therefore depends on untested functional forms, weakening the paper's stated contribution on balancing data quality and quantity.","section":"§5.1, Eq. (7) and Proposition 1"}],"minor_comments":[{"comment":"The abstract reports loss reductions of 'up to 40% on the simulated data and 25% on the real-world dataset,' whereas Section 6.2 reports reductions of 38.6% (simulated) and 47.3% (real data). Please reconcile these numbers.","section":"Abstract and §6.2"},{"comment":"The regularizer lambda_n is introduced in the quantile-regression objective but is never defined or discussed; please clarify whether it is a fixed constant, a tuning parameter, or simply omitted from the definition.","section":"§3.1, Eq. (2)"},{"comment":"In Example 6, the text says 'suppose X follows exponential distribution'; this should presumably be Y. Also, the displayed interval for hbar(Delta) and h(Delta) appears to use the same expression for both bounds; the lower and upper bounds should involve gamma(X) and bar{gamma}(X), respectively.","section":"§4.2, Example 6"},{"comment":"The text refers to 'For CPRP, we subdivide...' which appears to be a typo for the proposed CQPC method. Additionally, the use of the nearest 50 calibration points is described but not connected to the method in Section 5.2.2, where the parameter is denoted m; please align the terminology.","section":"§6.1, Experiment Setup"},{"comment":"The caption of Figure 10 reads 'Empirical pinball loss for MA model,' but the figure presents results for the Capital Bikeshare real-world dataset. Please correct the caption.","section":"§6.2, Figure 10"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an early version that does not yet contain proofs for its main theoretical claims; the appendices are algorithmic rather than mathematical. The conditional-coverage results require a substantive reworking of Assumption 1 into a high-probability statement with explicit failure control, or a restriction to estimator classes that satisfy a deterministic gap bound. The empirical section shows promising loss reductions but does not directly validate the coverage properties claimed by Theorems 1--4, so additional experiments reporting coverage would help. Given the scope of the needed changes, a major revision is appropriate; the paper should not be rejected outright because the algorithmic framework and the unconditional guarantee (Theorem 1) are sound and potentially useful, and the authors may be able to repair the theoretical foundations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Junyu — brief read on arXiv:2412.13159. The estimator is Q_c(X)=qhat(X)+Q_α(residuals), which is distributional conformal prediction with the sign flipped to a quantile decision. Theorem 1 recovers the standard marginal coverage bound and is fine, but it's known — the paper itself cites Chernozhukov et al. (2021). The claimed contribution is the conditional coverage results (Thms 2–4) and the pooling analysis.\n\nThat's where the trouble is. Assumption 1 asserts a deterministic gap function κ such that, for every realization of the training set, the difference of prediction errors at any two contexts is bounded by κ(n1, ξ). For a random quantile-regression estimator, that's false with positive probability. The cited Pan–Zhou bound is high-probability, not almost sure. Theorems 2–4 therefore hold only on the event that Assumption 1 holds, and the paper gives no bound on P(failure). The abstract's claim of guarantees 'independent of the correctness of the underlying model' is not established.\n\nWorse, there are no proofs anywhere. Theorems are stated and 'proven' is claimed, but a preprint with zero proofs of the main results cannot support a rigorous-guarantees narrative. This alone would force heavy revision.\n\nWhat is genuinely useful: The GTLC scheme (train once globally, calibrate on a local neighborhood) is simple and effective. The experiments show consistent pinball-loss reductions across four base learners, on simulations and the Bikeshare data. The comparison is only against uncalibrated versions of the same learners, and there are no error bars, so the headline percentages (38.6% simulated, 47.3% real) should be read as upper bounds on what a careful benchmark would show. Still, the direction is credible.\n\nThe pooling analysis (Thms 3 and 5, Prop. 1) is a reasonable formal framework, but it inherits the Assumption 1 problem. Appendix A's κ estimator is also ad hoc: it proxies quantile-error differences by pinball-loss differences. The pinball loss is only Lipschitz, not strongly convex, so tiny loss gaps can go with large quantile shifts near sparse conditional densities; the η correction is a fudge factor. The abstract/full-text discrepancy (40/25 vs 38.6/47.3 percent) is minor but worth fixing.\n\nBottom line: this is a paper with a good practical idea and an unproven, likely flawed theoretical superstructure. I would bring it to a reading group for the GTLC results and the failure-event critique. I would not cite it for the conditional guarantees. It does deserve a serious referee — the problem is important and the flaws are identifiable — but the referee should treat the theory section as needing major revision, possibly down to the marginal guarantee plus an honest conditional statement under explicitly quantified assumptions.\n\nRecommendation: send to review, with instructions that the proof of Thms 2–4 and a resolution of the deterministic-κ issue are prerequisites.","headline":"Unproven conditional guarantees rest on a deterministic gap-function assumption that standard estimators fail with positive probability; the practical local-calibration scheme is the real contribution.","tokens_in":24226,"tokens_out":5103,"would_cite":false,"duration_ms":47992,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G15","90B05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding the empirical residual quantile to any fitted quantile predictor yields a conformalized critical quantile with finite-sample coverage, and under regularity conditions conditional coverage and confidence intervals, even under model…","keywords":["feature-based newsvendor","conformal prediction","quantile regression","model misspecification","critical quantile","data quality and quantity","conditional coverage","inventory decision-making"],"falsifier":"Simulate data with a true quantile function that changes abruptly between nearby feature values by more than any gap function of the form $\\kappa(n_1,\\xi)=C\\sqrt{\\xi^\\nu/n_1}$, then measure the conditional coverage $P(Y_{n+1}\\leq\\hat{q}^c_\\alpha(X_{n+1})\\mid X_{n+1})$ at test points near the abrupt change. If the empirical coverage leaves the interval $\\alpha\\pm\\varphi(\\Delta,B)$ from Theorem 3, or if the coverage gap does not shrink with $n_2$ as predicted, the paper's conditional claim is falsified.","tokens_in":23113,"feed_emoji":"📦","tokens_out":14625,"duration_ms":118527,"temperature":0.7,"pith_summary":"The paper studies the feature-based newsvendor problem, where a retailer chooses an order quantity after seeing features such as weather or seasonality, balancing overage and underage costs. Its proposal is a two-phase recipe: fit any quantile predictor on a training set, then add to its output an empirical quantile of the prediction residuals computed on a separate calibration set. The central claim is that this conformalized critical quantile has a finite-sample guarantee on the probability that demand falls at or below the chosen order quantity, and the guarantee holds even when the trained model is misspecified. Under additional regularity assumptions on how prediction error varies between contexts and how demand density behaves near the quantile, the paper also derives conditional guarantees and a confidence interval for the true critical quantile. The method is validated on simulated data and Washington D.C. bike-sharing data, reporting substantially lower newsvendor loss than benchmarks.","feed_headline":"Conformal fix keeps newsvendor orders on target when models fail","feed_subtitle":"Adding a residual-quantile correction to any demand model yields coverage guarantees and up to 47% lower inventory loss.","key_machinery":"The load-bearing object is the conformalized critical quantile $\\hat{q}^c_\\alpha(X_0)=\\hat{q}_\\alpha(X_0)+Q_\\alpha(s,\\mathcal{I}_2)$, where the correction term $Q_\\alpha$ is the empirical quantile of signed calibration residuals at level $\\alpha(1+1/|\\mathcal{I}_2|)$. Because the scores are signed, a model that systematically overestimates the quantile receives a negative correction and one that underestimates receives a positive correction. Conditional guarantees are carried by two structural assumptions: the gap function $\\kappa(n_1,\\xi(x_1,x_2))$ bounds the difference in quantile-prediction error between contexts separated by distance $\\xi$, and the margin functions $\\underline{h}(\\Delta),\\bar{h}(\\Delta)$ control how much demand probability mass sits within a $\\Delta$-neighborhood of the true quantile. These combine into the local bound $\\varphi(\\Delta,B)=\\bar{h}(\\Delta+\\kappa(n_1(B),\\xi(B)))+\\exp(-2n_2(B)\\underline{h}(\\Delta)^2)$ that Theorem 3 places on conditional coverage, and they determine the optimal data-pooling ball.","core_discovery":"The core discovery is the additive conformalization identity $\\hat{q}^c_\\alpha(X_0) = \\hat{q}_\\alpha(X_0) + Q_\\alpha(s,\\mathcal{I}_2)$, where $s_i = Y_i - \\hat{q}_\\alpha(X_i)$ are signed residuals on the calibration set and $Q_\\alpha(s,\\mathcal{I}_2)$ is the empirical quantile of those residuals at level $\\alpha(1+1/|\\mathcal{I}_2|)$. Theorem 1 states that for i.i.d. data with almost surely distinct scores, $\\alpha \\leq P(Y_{n+1} \\leq \\hat{q}^c_\\alpha(X_{n+1})) \\leq \\alpha + 1/(n_2+1)$, so the corrected quantile lies between the true $\\alpha$-quantile and the true $(\\alpha+1/(n_2+1))$-quantile regardless of whether the underlying demand model is correct. Under a gap-function assumption and a margin condition, Theorems 2 and 3 convert this marginal guarantee into conditional coverage bounds, and Theorem 4 provides a confidence interval for the true critical quantile whose width decreases as training and calibration sample sizes grow.","pith_inferences":["Editorial extension: The same additive residual-quantile correction can likely be applied to any decision problem whose optimal action is a quantile, such as capacity or staffing decisions, since the correction only requires signed residuals from a fitted quantile.","Editorial extension: The data-driven estimation of the gap function in the appendix substitutes empirical pinball-loss differences for true quantile-error differences; whether that substitution preserves the theorem's deterministic bound is not formally analyzed and would be a natural stress test.","Editorial extension: For nonstationary demand, a rolling-window version of the calibration step would restore a form of local exchangeability, but the paper does not quantify how much coverage degrades when exchangeability is only approximate."],"forward_implications":["Any quantile regression algorithm can be plugged into the training phase, and the conformalized output still satisfies the finite-sample coverage bound of Theorem 1.","Conditional coverage error decays exponentially with calibration-set size, so the calibration phase directly tightens the guarantee as more data arrive.","The optimal pooling region balances quality and quantity: adding more local data tightens the calibration term, while pushing the region wider increases the gap-function term; in the big-data limit the optimal diameter shrinks to zero.","The confidence interval from Theorem 4 lets a manager choose an optimistic or conservative order quantity within a stated uncertainty range.","Numerical results on simulated data and the Capital Bikeshare dataset report reductions in empirical newsvendor loss of up to 38.6% and 47.3% from adding local calibration."],"supporting_citations":[{"why":"Introduces split conformal prediction and the residual-quantile argument that the paper adapts to the newsvendor quantile.","marker":"Shafer and Vovk 2008"},{"why":"Defines quantile regression and the pinball loss used to train the base predictor in the training phase.","marker":"Koenker and Bassett Jr 1978"},{"why":"Develops conformalized quantile regression, the closest precursor whose score construction and local-length motivation the paper modifies.","marker":"Romano et al. 2019"},{"why":"Provides the non-asymptotic quantile-regression estimation error bound used to justify the parametric gap-function form used in the paper.","marker":"Pan and Zhou 2021"},{"why":"Supplies the Capital Bikeshare dataset used for the real-world validation of the calibration method.","marker":"Fanaee-T and Gama 2014"},{"why":"Establishes the feature-based newsvendor problem as a setting and a baseline against which the paper positions its conformal approach.","marker":"Ban and Rudin 2019"}],"fun_headline_variants":["Conformal residual quantile fixes newsvendor model errors","Newsvendor loss cut up to 40% with conformal residual tuning","Model-free newsvendor: conformal calibration guarantees coverage","Add a conformal residual quantile to any demand model","Conformal correction gives coverage regardless of demand model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conditional theorems assume that the difference between the quantile-prediction errors at two contexts is bounded by a known function of their distance, and that the probability of demand falling just above or below the true quantile changes in a controlled way; if either control is absent or inaccurate, the local coverage guarantees do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Conformal residual quantile fixes newsvendor model errors","Newsvendor loss cut up to 40% with conformal residual tuning","Model-free newsvendor: conformal calibration guarantees coverage","Add a conformal residual quantile to any demand model","Conformal correction gives coverage regardless of demand model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001182,"raw_usage":{"total_tokens":4922,"prompt_tokens":1023,"completion_tokens":3899,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":3818}},"tokens_in":639,"tokens_out":3899,"duration_ms":27898,"temperature":1.0,"reasoning_tokens":3818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:22:57.383149+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data with a true quantile function that changes abruptly between nearby feature values by more than any gap function of the form $\\kappa(n_1,\\xi)=C\\sqrt{\\xi^\\nu/n_1}$, then measure the conditional coverage $P(Y_{n+1}\\leq\\hat{q}^c_\\alpha(X_{n+1})\\mid X_{n+1})$ at test points near the abrupt change. If the empirical coverage leaves the interval $\\alpha\\pm\\varphi(\\Delta,B)$ from Theorem 3, or if the coverage gap does not shrink with $n_2$ as predicted, the paper's conditional claim is falsified.","supporting_citations":[{"cited_title":"Journal of Machine Learning Research 9(3)","cited_arxiv_id":null,"evidence_quote":"Introduces split conformal prediction and the residual-quantile argument that the paper adapts to the newsvendor quantile."},{"cited_title":"Econometrica: journal of the Econometric Society 33--50","cited_arxiv_id":null,"evidence_quote":"Defines quantile regression and the pinball loss used to train the base predictor in the training phase."},{"cited_title":"Advances in neural information processing systems 32","cited_arxiv_id":null,"evidence_quote":"Develops conformalized quantile regression, the closest precursor whose score construction and local-length motivation the paper modifies."},{"cited_title":"Information and Inference: A Journal of the IMA 10(3):813--861","cited_arxiv_id":null,"evidence_quote":"Provides the non-asymptotic quantile-regression estimation error bound used to justify the parametric gap-function form used in the paper."},{"cited_title":"Progress in Artificial Intelligence 2:113--127","cited_arxiv_id":null,"evidence_quote":"Supplies the Capital Bikeshare dataset used for the real-world validation of the calibration method."}],"review_version":1}