REVIEW 2 major objections 5 minor 30 references
Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Conditional on the realized localization center, RLCP bounds the conditional-coverage gap and the oracle-relative length error uniformly over the localization ball with probability at least 1−δ over calibration, at the classical…
desk verdict Finite-sample uniform-local guarantees for RLCP with an honest fixed-score analysis; the length bound is vacuous when the localized score density vanishes, and the abstract should say so. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three devices carry the argument. The reverse-law representation: after the auxiliary center $\tilde{x}$ is drawn from the kernel-smoothed law, the test covariate, conditional on $\tilde{x}$, has exactly the localized distribution $\Lambda_{\tilde{x},h}(du) = (w_{\tilde{x},h}(u)/Z_{\tilde{x},h})P_X(du)$, so RLCP is weighted conformal prediction under covariate shift and inherits finite-sample marginal validity. The calibration event $\Omega^{\mathrm{cal}}_{n,\tilde{x},h,\delta}$: a single event of probability at least $1-\delta$ on which the kernel-weighted empirical score CDF concentrates uniformly and the test-point weight is uniformly small, so every bound holds simultaneously for all $x \in B(\tilde{x}, h) \cap \mathcal{X}$. The inversion devices: the localized score-density floor $\kappa^{(S)}_{\alpha,\tilde{x},h}$ (the essential infimum of the localized score density on the quantile window) converts level error into quantile error, the length map $q \mapsto |C^{(S)}(x,q)|$ converts quantile error into length error, and Hölder continuity of the conditional score CDF converts the mismatch between the local mixture and the target covariate into the $h^\beta$ bias.
What would settle it
Run RLCP at the balanced bandwidth $h_\star \asymp (\log(4n)/(n p_{\min}))^{1/(2\beta+d)}$ on a score that satisfies the paper's A1–A3 but whose conditional score density vanishes on a band around the $(1-\alpha)$-quantile (for instance, a residual score with a gap in the noise density): if the realized oracle-relative length error does not shrink at the predicted $n^{-\beta/(2\beta+d)}$ rate, or stays bounded while the coverage gap shrinks, then the length claim fails precisely in the regime where $\kappa^{(S)}_{\alpha,\tilde{x},h} = 0$ makes the theorem's bound vacuous.
Extended reading notes
Core claim
With probability at least $1-\delta$ over the calibration sample, conditional on the realized localization center $\tilde{x}$, and for calibration sizes and bandwidths above explicit thresholds, the RLCP threshold deviates from the conditional score-oracle quantile by $O\big((A^{\mathrm{cal}}_{n,h}(\delta) + h^\beta)/\kappa^{(S)}_{\alpha,\tilde{x},h}\big)$, uniformly over $B(\tilde{x}, h) \cap \mathcal{X}$; from this, the conditional-coverage gap is bounded by $O(A^{\mathrm{cal}}_{n,h}(\delta) + h^\beta)$ (Theorem 3) and the oracle-relative length error by $O\big(L_{\mathrm{len}}(x)\,(A^{\mathrm{cal}}_{n,h}(\delta) + h^\beta)/\kappa^{(S)}_{\alpha,\tilde{x},h}\big)$ (Theorem 2), where $A^{\mathrm{cal}}_{n,h}(\delta) = \sqrt{\log(4/\delta)/(n p_{\min} h^d)} + \log(4/\delta)/(n p_{\min} h^d)$ captures the local effective sample size and $\kappa^{(S)}_{\alpha,\tilde{x},h}$ is the localized score-density floor. The first term in the rates is the localization bias coming from Hölder regularity of the conditional score law; the others are calibration error, and balancing them at $h_\star \asymp (\log(4n)/(n p_{\min}))^{1/(2\beta+d)}$ gives the classical pointwise nonparametric rate $n^{-\beta/(2\beta+d)}$ up to logarithmic factors. When the score is learned on an independent fold and targets a pivotal population score — one whose $(1-\alpha)$-quantile is the same at every covariate, as with conformalized quantile regression and the probability-integral-transform score — the localization bias vanishes, and Theorems 6 and 7 deliver the same uniform local control with the error split into calibration error plus a linearly entering uniform score-estimation error.
Load-bearing premise
The load-bearing premise is that the localized score distribution keeps strictly positive density across the quantile window; the paper's own passage before Theorem 2 concedes that this is not guaranteed by the main assumptions, and when that density floor is zero the stated length bound is formally true but vacuous, so the oracle-length guarantee rests on the separate shape-function condition of Lemma 1.
Editorial extensions
If this is right
- Balancing the two error sources pins the optimal bandwidth at $h_\star \asymp (\log(4n)/(n p_{\min}))^{1/(2\beta+d)}$, where both the coverage gap and the length error attain the classical rate $n^{-\beta/(2\beta+d)}$ up to logarithmic factors (equations (19)–(20)).
- Because the calibration event is sample-specific and common to the whole ball, the guarantee is simultaneous: no union bound over test covariates and no averaging over the auxiliary randomization is required for uniformity over $B(\tilde{x}, h) \cap \mathcal{X}$.
- Under a pivotal target score the localization bias disappears and the bandwidth enters only through the effective calibration size $n p_{\min} h^d$, so the score-estimation error enters linearly: improving the learned score sharpens the localized oracle comparison at exactly the training rate.
- The assumptions, including the density-floor condition via Lemma 1, are verified for residual, conformalized-quantile-regression, and distributional scores, so the bounds apply to the standard score constructions in practice.
Reading between the lines
- A design lesson the paper leaves implicit: when a strong, near-pivotal score is available, the bandwidth should be pushed as large as the density-floor constraint allows rather than tuned to the fixed-score optimum, since localization then mainly buys effective sample size while the score itself carries covariate adaptivity.
- The vacuous-kappa caveat points to a concrete stress test: scores whose conditional density is tiny at the target quantile (for example, extreme-quantile regions in heavy-tailed responses) should show realized length errors far above the predicted rate, and mapping where the degradation begins would delineate the exact class of scores for which the oracle-tracking claim holds.
- The calibration-versus-localization decomposition is the same structure that governs locally weighted regression, suggesting the $n^{-\beta/(2\beta+d)}$ rate is the natural minimax benchmark for any locally weighted conformal procedure; the pivotal-score result further predicts that with a well-learned score, wide-bandwidth (near-marginal) calibration is nearly optimal, since localization adds no
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops finite-sample guarantees for randomly localized conformal prediction (RLCP) targeted at the realized prediction set rather than at marginal or averaged quantities. For a fixed score, under Hölder regularity of the conditional score CDF, kernel and design conditions, and a length-regularity assumption, Theorems 2 and 3 give, with probability at least 1−δ over the calibration sample and conditional on the realized localization center, uniform bounds over the localization ball for the conditional-coverage gap and for the oracle-relative length error. The bounds combine a calibration term A_cal_{n,h}(δ) with a localization bias h^β; the length bound also carries the inverse of the localized score-density floor κ. For a data-split learned score that targets a pivotal population score, Theorems 6 and 7 give analogous bounds in which calibration error and score-estimation error enter separately and linearly. The residual, CQR, and distributional scores are shown to satisfy the structural assumptions, and numerical experiments illustrate the predicted rates and the coverage-length trade-off in a disclosed real-data hybrid implementation.
Significance. If the results are taken with the required positivity of the localized score-density floor, the paper is a substantial contribution: it gives high-probability, realized-center oracle comparisons for RLCP, explicitly separates calibration error from localization bias, and clarifies when learned pivotal scores remove the localization bias. The proofs are detailed and appear internally consistent, with explicit lemmas for concentration, quantile inversion, and score perturbation, and the experimental sections are thorough and transparent, including exact seeds, bandwidth grids, and disclosure that the real-data procedure is a hybrid rather than formal RLCP. The main caveat is that the advertised length guarantee is non-vacuous only under an additional density-minorization condition that is not part of A1–A3; the paper itself notes this after Lemma 1, but the abstract and introduction do not.
major comments (2)
- [Abstract and Section 3.2, Theorem 2] The abstract's claim that, for any fixed score, the paper proves finite-sample bounds for the length error relative to the oracle is not supported as stated. The right-hand side of Theorem 2 contains the factor 1/κ(S)_{α,x-tilde,h}, and under the convention 1/0=∞ the bound is formally true but vacuous whenever κ=0. Assumptions A1–A3 do not imply κ>0: A1 is an upper bound on the score density, A2 is a Hölder condition on the score CDF, and A3 concerns length regularity. The only sufficient condition supplied, Lemma 1's inequality (17), requires a covariate-local lower bound on p_{S|X} through a shape function, and it is verified for the residual, CQR, and distributional scores but not for arbitrary fixed scores. The text after Lemma 1 acknowledges the vacuity, so the theorem itself is formally correct, but the abstract and the introduction should qualify the length guarantee by κ>0 or by an equivalent density-floor assumption.
- [Section 4.2, Theorems 6 and 7 and Proposition 18] The same κ-positivity issue propagates to the learned-score results. The bounds in (28) and (29) are multiplied by 1/κ(S⋆)_{α,x-tilde,h}, yet the hypothesis list of Theorem 6 does not state κ(S⋆)>0. Assumptions A1(S⋆), A4(α), and A5 do not imply this positivity; Lemma 15's identification of the localized quantile already assumes κ(S⋆)>0, and Proposition 18 includes it as a hypothesis in the text. As written, Theorem 6 is non-vacuous only for target scores whose localized density floor is positive. The theorem statements should include this condition explicitly, or the theorems should be qualified as informative only when κ(S⋆)>0.
minor comments (5)
- [Abstract] The phrase 'for any fixed score' should be replaced by a formulation that explicitly conditions on the localized density minorization condition (13) for the length result.
- [Section 1, Eq. (1) and following text] The description of the length bound as a 'constant multiple' is imprecise because the factor 1/κ(S)_{α,x-tilde,h} is not a universal structural constant; it depends on the realized center, the bandwidth, and the score, and it can be infinite under the stated convention.
- [Section 3.3, Proposition 4 and Lemma 1] For the fixed-score examples, positivity of κ is verified only for the symmetric level choice (15), whereas Theorems 2 and 3 are stated for arbitrary α−<α<α+. The paper should clarify whether the main fixed-score length theorem is intended for those symmetric levels or whether the user must verify (13) separately for other level choices.
- [Section 5 and Appendix M] The real-data procedure is carefully disclosed as a hybrid, but referring to it simply as 'RLCP' in the decile figures and tables may mislead readers; a label such as 'RLCP-hybrid' would make the distinction from formal RLCP clearer.
- [Figure 5 caption] In the displayed expression for the decomposition proxy, the placeholder 'slow' appears where κ(S⋆) is intended; this should be corrected.
Circularity Check
No circular derivation: bounds follow from stated structural conditions and independent concentration results; the κ>0 caveat is a vacuity limitation, not a circular step.
full rationale
The paper's fixed-score analysis derives Theorem 3's coverage gap from Bernstein/Bousquet/Talagrand concentration (Proposition 11), A2 Hölder bias, and direct CDF comparisons; Theorem 2 additionally converts threshold error to set-length error via A3 and the quantile-inversion cost 1/κ. None of these steps restates its conclusion: κ(S)_{α,x̃,h} is defined as an essential infimum of the localized score density (13), not as a fitted parameter, and the theorem's RHS is not used to define any input. Lemma 1 gives a sufficient, unverified-for-arbitrary-scores condition for κ>0, and the paper explicitly records that the length bound is vacuous when κ=0 ('While Theorem 2 below remains valid for κ=0, the resulting upper bound is vacuous'), which is a limitation rather than circularity. The learned-score results (Theorems 6, 7) take A5's uniform score-estimation rate as an assumption and propagate it linearly through Lemmas 16-17; the rate is not fitted to the same outcomes being bounded. No load-bearing self-citations appear: references to Hore-Barber, Min et al., Romano et al., etc. are external prior results, and the reverse-law representation is used as a starting point, not as the paper's own uniqueness claim. The simulations are explicitly rate diagnostics and do not enter the theorem chain. Thus there is no exhibited reduction of any claimed prediction to its own inputs.
Assumptions & free parameters
assumptions (8)
- domain assumption P1: marginal density p_X bounded below by p_min and X is (c_reg, r0)-regular.
- domain assumption K1: kernel K is bounded, symmetric, integrates to one, supported on B(0,1) and lower-bounded by kappa on B(0,eta).
- domain assumption A1(S): conditional score density bounded above by sigma_up.
- domain assumption A2(S): conditional score CDF is beta-Holder in the covariate, uniformly in t.
- domain assumption A3(S): the map q -> |C(S)(x,q)| is L_len(x)-Lipschitz.
- domain assumption A4(alpha): target score S* is (1-alpha)-pivotal.
- domain assumption A5(S): uniform high-probability bound on the essential supremum of the score estimation error.
- domain assumption Lemma 1 condition (17): conditional score density is lower-bounded by sigma_psi(u,h) psi(F_{S|X}(s|z)).
Cite this review
Pith. "Pith review of Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction." pith.science (2026). https://pith.science/paper/Q3D7DPP5
@misc{pith2026260806206,
author = {Pith},
title = {Pith review of: Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q3D7DPP5}},
note = {Machine review of arXiv:2608.06206}
}
abstract
Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable. Randomly localized conformal prediction (RLCP) mitigates this gap by calibrating near the test point while preserving marginal coverage. Existing theory, however, lacks finite-sample guarantees for the realized localized set that jointly control conditional validity and oracle efficiency. We provide such guarantees. For any fixed score, under H\"older regularity of the conditional score CDF and standard density and kernel assumptions, we prove high-probability bounds, uniform over a realized localization neighbourhood, for the conditional-coverage gap and the length error relative to the oracle. The bounds decompose into an $O(h^\beta)$ localization bias and a calibration term decreasing with calibration size, clarifying the bandwidth bias-variance tradeoff and when RLCP tracks the oracle. We also analyze data-split learned scores: when the score targets a pivotal score, as in conformalized quantile regression, uniform local guarantees decompose into fixed-score calibration and uniform score-estimation errors, showing that improved learning sharpens localized guarantees.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Jean-Yves Audibert and Alexandre B. Tsybakov. Fast learning rates for plug-in classifiers.The Annals of Statistics, 35(2):608–633, 2007. doi: 10.1214/009053606000001217
-
[2]
Rina Foygel Barber and Ryan J. Tibshirani. Unifying different theories of conformal prediction. Electronic Journal of Statistics, 20(1):1428–1474, 2026. doi: 10.1214/26-EJS2513. URL https://doi.org/10.1214/26-EJS2513
-
[3]
Candès, Aaditya Ramdas, and Ryan J
Rina Foygel Barber, Emmanuel J. Candès, Aaditya Ramdas, and Ryan J. Tibshirani. The limits of distribution-free conditional predictive inference.Information and Inference: A Journal of the IMA, 10(2):455–482, 2021. doi: 10.1093/imaiai/iaaa017. URLhttps://doi.org/10.1093/ imaiai/iaaa017
-
[4]
Sergey G. Bobkov and Michel Ledoux.One-Dimensional Empirical Measures, Order Statistics, and Kantorovich Transport Distances, volume 261 ofMemoirs of the American Mathematical Society. American Mathematical Society, 2019. doi: 10.1090/memo/1259
-
[5]
Oxford University Press, Oxford, 2013
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart.Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, Oxford, 2013. ISBN 9780199535255. doi: 10.1093/acprof:oso/9780199535255.001.0001. URL https://doi.org/10. 1093/acprof:oso/9780199535255.001.0001. 59 Beyond Marginal ValidityConrad et al
arXiv 2013
-
[6]
Thomas Brooks, D. Pope, and Michael A. Marcolini. Airfoil self-noise. UCI Machine Learning Repository, 1989. URLhttps://doi.org/10.24432/C5VW2C
doi:10.24432/c5vw2c 1989
-
[7]
Probal Chaudhuri. Nonparametric estimates of regression quantiles and their local Bahadur representation.The Annals of Statistics, 19(2):760–777, 1991. doi: 10.1214/aos/1176348119
arXiv 1991
-
[8]
Victor Chernozhukov, Iván Fernández-Val, Blaise Melly, and Kaspar Wüthrich. Generic inference onquantileandquantileeffectfunctionsfordiscreteoutcomes.Journal of the American Statistical Association, 115(529):123–137, 2020. doi: 10.1080/01621459.2019.1611581
arXiv 2020
Show all 30 references
-
[9]
Distributional conformal prediction
Victor Chernozhukov, Kaspar Wüthrich, and Yinchu Zhu. Distributional conformal prediction. Proceedings of the National Academy of Sciences, 118(48):e2107794118, 2021. doi: 10.1073/pnas. 2107794118. URLhttps://arxiv.org/abs/1909.07889
2021 arXiv
-
[10]
Dhillon, George Deligiannidis, and Tom Rainforth
Guneet S. Dhillon, George Deligiannidis, and Tom Rainforth. On the expected size of conformal prediction sets. InProceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 ofProceedings of Machine Learning Research, pages 1549–1557. ...
2024
-
[11]
Uwe Einmahl and David M. Mason. Uniform in bandwidth consistency of kernel-type function estimators.The Annals of Statistics, 33(3):1380–1403, 2005. doi: 10.1214/009053605000000129. URLhttps://doi.org/10.1214/009053605000000129
2005 doi
-
[12]
A note on generalized inverses.Mathematical Methods of Operations Research, 77(3):423–432, 2013
Paul Embrechts and Marius Hofert. A note on generalized inverses.Mathematical Methods of Operations Research, 77(3):423–432, 2013. doi: 10.1007/s00186-013-0436-7
2013 doi
-
[13]
Bike sharing
Hadi Fanaee-T. Bike sharing. UCI Machine Learning Repository, 2013. URLhttps://doi. org/10.24432/C5W894
2013 doi
- [14]
-
[15]
Localized conformal prediction: A generalized inference framework for conformal prediction.Biometrika, 110(1):33–50, 2023
Leying Guan. Localized conformal prediction: A generalized inference framework for conformal prediction.Biometrika, 110(1):33–50, 2023. doi: 10.1093/biomet/asac040. URL https: //doi.org/10.1093/biomet/asac040
2023 doi
-
[16]
Uniform bias study and Bahadur representation for local polynomial estimators of the conditional quantile function.Econometric Theory, 28(1):87–129,
Emmanuel Guerre and Camille Sabbah. Uniform bias study and Bahadur representation for local polynomial estimators of the conditional quantile function.Econometric Theory, 28(1):87–129,
-
[17]
Conformal prediction with local weights: Ran- domization enables robust guarantees.Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(2):549–578, 2025
Rohan Hore and Rina Foygel Barber. Conformal prediction with local weights: Ran- domization enables robust guarantees.Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(2):549–578, 2025. doi: 10.1093/jrsssb/qkae103. URL https: //doi.org/10.1093/jrss...
2025 doi
-
[18]
Rafael Izbicki, Gilson Shimizu, and Rafael B. Stern. Cd-split and hpd-split: Efficient conformal regions in high dimensions.Journal of Machine Learning Research, 23(87):1–32, 2022. URL https://jmlr.org/papers/v23/20-797.html
2022
-
[19]
Pappas, and Hamed Hassani
Shayan Kiyani, George J. Pappas, and Hamed Hassani. Length optimization in conformal prediction. InAdvances in Neural Information Processing Systems, volume 37, 2024. doi: 10.52202/079017-3158. URLhttps://doi.org/10.52202/079017-3158. 60 Beyond Marginal ValidityConrad et al
2024 doi
-
[20]
Distribution-free prediction bands for non-parametric regression
Jing Lei and Larry Wasserman. Distribution-free prediction bands for non-parametric regression. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 76(1):71–96, 2014. doi: 10.1111/rssb.12021
2014 doi
-
[21]
Tibshirani, and Larry Wasserman
Jing Lei, Max G’Sell, Alessandro Rinaldo, Ryan J. Tibshirani, and Larry Wasserman. Distribution-free predictive inference for regression.Journal of the American Statistical Associ- ation, 113(523):1094–1111, 2018. doi: 10.1080/01621459.2017.1307116
2018
-
[22]
A unified theory of conditional coverage in conformal prediction with applications.arXiv preprint arXiv:2605.11602, 2026
Yinjie Min, Liuhua Peng, and Changliang Zou. A unified theory of conditional coverage in conformal prediction with applications.arXiv preprint arXiv:2605.11602, 2026. doi: 10.48550/ arXiv.2605.11602. URLhttps://arxiv.org/abs/2605.11602
-
[23]
Physicochemical properties of protein tertiary structure
Prashant Rana. Physicochemical properties of protein tertiary structure. UCI Machine Learning Repository, 2013. URLhttps://doi.org/10.24432/C5QW3H
2013 doi
-
[24]
Yaniv Romano, Evan Patterson, and Emmanuel J. Candès. Conformalized quantile regression. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32, pages 3543–3553. Curran Associa...
2019
-
[25]
Uniform confidence bands for local polynomial quantile estimators.ESAIM: Probability and Statistics, 18:265–276, 2014
Camille Sabbah. Uniform confidence bands for local polynomial quantile estimators.ESAIM: Probability and Statistics, 18:265–276, 2014. doi: 10.1051/ps/2013035
2014
-
[26]
A tutorial on conformal prediction.Journal of Machine Learning Research, 9:371–421, 2008
Glenn Shafer and Vladimir Vovk. A tutorial on conformal prediction.Journal of Machine Learning Research, 9:371–421, 2008. URLhttps://jmlr.org/papers/v9/shafer08a.html
2008
-
[27]
Springer, 2005
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer.Algorithmic Learning in a Random World. Springer, 2005. doi: 10.1007/b106715. URLhttps://doi.org/10.1007/b106715
2005 doi
-
[28]
Non-asymptotic analysis of efficiency in con- formalized regression
Yunzhen Yao, Lie He, and Michael Gastpar. Non-asymptotic analysis of efficiency in con- formalized regression. In C. Vondrick, B. Hariharan, C. Raffel, L. Pinto, D. Yang, and A. Faust, editors,International Conference on Learning Representations, volume 2026, pages 42546–42583...
2026
-
[29]
Concrete compressive strength
I-Cheng Yeh. Concrete compressive strength. UCI Machine Learning Repository, 1998. URL https://doi.org/10.24432/C5PK67. 61 Beyond Marginal ValidityConrad et al. 10¡1 100 set-length error j jC Rj ¡ jC orcj j (a) ¯ = 0:5 h ⋆ slope ¡0:50 [¡0:53; ¡ 0:49] balance ¡0:500 10¡1 100 (c...
1998 doi
-
[2012]
URLhttps://doi.org/10.1017/S0266466611000132
doi: 10.1017/S0266466611000132. URLhttps://doi.org/10.1017/S0266466611000132
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.