{"id":"bdbafbce-faa1-4fca-93c9-91aff0e97b6f","arxiv_id":"2608.06206","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For randomly localized conformal prediction, high-probability bounds, uniform over a realized neighborhood, control conditional coverage error and oracle-relative length at rate h^beta + sqrt(log(1/delta)/(n h^d)) plus a calibration term.","lead":"This paper proves finite-sample guarantees for randomly localized conformal prediction (RLCP): with high probability and uniformly over the realized localization neighborhood, the prediction set's conditional coverage error and its length error relative to the oracle are both bounded by an explicit bias-variance expression. The result clarifies when local calibration tracks the ideal conditional oracle and how to balance bandwidth against effective sample size.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Length-error guarantee is vacuous when the localized score-density floor is zero: κ>0 is not implied by A1–A3, so Theorem 2 (and Theorem 6 via Proposition 18) asserts only an infinite bound for such scores; the abstract's unqualified oracle-length claim overstates the result.","rationale":"I agree with the reader's weakest_assumption; the concern is real and load-bearing. Theorems 2 and 6 do state the 1/0=∞ convention, and the text after Lemma 1 admits vacuity, so the mathematical core is not internally inconsistent. However, the paper's advertised contribution includes finite-sample oracle-relative length control, and for arbitrary fixed scores satisfying A1–A3 that control is empty unless κ>0. The coverage theorem survives, and the worked examples verify κ for their specific scores, so a conditional acceptance—not rejection—is appropriate. The required revision is narrow: state κ>0 as an explicit assumption in the abstract and theorem statements, or clearly qualify the length claim. I did not find a separate flaw in the proof mechanism: the concentration step, quantile-stability lemma, and learned-score decomposition are coherent. The concern is about the scope of the claim, not the validity of the derivations under the stated (or intended) assumptions.","tokens_in":63375,"tokens_out":6840,"duration_ms":79819,"concrete_test":"Construct X∼Unif[0,1], Y independent of X with density 1.25 on [0,0.4]∪[0.6,1] and 0 on (0.4,0.6); take S(x,y)=y, α=0.5, α−=0.4, α+=0.6. This satisfies A1 (σ_up=1.25), A2 (the conditional CDF is x-independent), and A3 (the length map q↦q is 1-Lipschitz on Y=[0,1]). The localized quantile window of (11) is [q_{0.4}, q_{0.6}]=[0.32,0.6], on which the localized score density (12) is zero a.e. on (0.4,0.6), so κ(S)_{α,tilde x,h}=0 in (13). Plugging into Theorem 2 gives right-hand side ∞, confirming that the abstract's unqualified length guarantee fails for a fixed score satisfying A1–A3. This test settles whether κ>0 must be stated as an explicit assumption rather than left implicit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central advertised result has two components: conditional coverage and oracle-relative length. Theorem 3's coverage bound is genuinely uniform and depends only on A_cal + h^β. The length bound is different: Theorem 2 has right-hand side L_len(x)(A_cal + h^β)/κ(S)_{α,tilde x,h}, and Theorem 6 similarly has 1/κ(S⋆) multiplying calibration and training errors. With the stated convention 1/0=∞, the theorems are formally true for κ=0, but they provide no finite-sample length control. The abstract and introduction present 'length error relative to the oracle' as a proved finite-sample guarantee under 'Hölder regularity ... and standard density and kernel assumptions' (A1–A3), without the extra requirement κ>0. Assumptions A1–A3 do not imply κ>0: A1 is an upper bound on the score density, A2 is a CDF Hölder condition, and A3 concerns set length. Lemma 1's sufficient condition (17) requires a lower bound pS|X(s|z) ≥ σ_ψ(u,h)ψ(F(S|X)(s|z)) over every z in B(u,h)—an additional structural assumption verified for the residual/CQR/PIT examples but not for arbitrary fixed scores. The same caveat propagates to the learned-score theorems, since Proposition 18 needs κ(S⋆)>0 for a finite adaptive deviation rate. Thus the unqualified 'any fixed score' length claim in the abstract is not supported; the paper itself notes the vacuity only in the text after Lemma 1.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops finite-sample guarantees for randomly localized conformal prediction (RLCP) targeted at the realized prediction set rather than at marginal or averaged quantities. For a fixed score, under Hölder regularity of the conditional score CDF, kernel and design conditions, and a length-regularity assumption, Theorems 2 and 3 give, with probability at least 1−δ over the calibration sample and conditional on the realized localization center, uniform bounds over the localization ball for the conditional-coverage gap and for the oracle-relative length error. The bounds combine a calibration term A_cal_{n,h}(δ) with a localization bias h^β; the length bound also carries the inverse of the localized score-density floor κ. For a data-split learned score that targets a pivotal population score, Theorems 6 and 7 give analogous bounds in which calibration error and score-estimation error enter separately and linearly. The residual, CQR, and distributional scores are shown to satisfy the structural assumptions, and numerical experiments illustrate the predicted rates and the coverage-length trade-off in a disclosed real-data hybrid implementation.","tokens_in":63676,"tokens_out":8405,"duration_ms":93835,"significance":"If the results are taken with the required positivity of the localized score-density floor, the paper is a substantial contribution: it gives high-probability, realized-center oracle comparisons for RLCP, explicitly separates calibration error from localization bias, and clarifies when learned pivotal scores remove the localization bias. The proofs are detailed and appear internally consistent, with explicit lemmas for concentration, quantile inversion, and score perturbation, and the experimental sections are thorough and transparent, including exact seeds, bandwidth grids, and disclosure that the real-data procedure is a hybrid rather than formal RLCP. The main caveat is that the advertised length guarantee is non-vacuous only under an additional density-minorization condition that is not part of A1–A3; the paper itself notes this after Lemma 1, but the abstract and introduction do not.","major_comments":[{"comment":"The abstract's claim that, for any fixed score, the paper proves finite-sample bounds for the length error relative to the oracle is not supported as stated. The right-hand side of Theorem 2 contains the factor 1/κ(S)_{α,x-tilde,h}, and under the convention 1/0=∞ the bound is formally true but vacuous whenever κ=0. Assumptions A1–A3 do not imply κ>0: A1 is an upper bound on the score density, A2 is a Hölder condition on the score CDF, and A3 concerns length regularity. The only sufficient condition supplied, Lemma 1's inequality (17), requires a covariate-local lower bound on p_{S|X} through a shape function, and it is verified for the residual, CQR, and distributional scores but not for arbitrary fixed scores. The text after Lemma 1 acknowledges the vacuity, so the theorem itself is formally correct, but the abstract and the introduction should qualify the length guarantee by κ>0 or by an equivalent density-floor assumption.","section":"Abstract and Section 3.2, Theorem 2"},{"comment":"The same κ-positivity issue propagates to the learned-score results. The bounds in (28) and (29) are multiplied by 1/κ(S⋆)_{α,x-tilde,h}, yet the hypothesis list of Theorem 6 does not state κ(S⋆)>0. Assumptions A1(S⋆), A4(α), and A5 do not imply this positivity; Lemma 15's identification of the localized quantile already assumes κ(S⋆)>0, and Proposition 18 includes it as a hypothesis in the text. As written, Theorem 6 is non-vacuous only for target scores whose localized density floor is positive. The theorem statements should include this condition explicitly, or the theorems should be qualified as informative only when κ(S⋆)>0.","section":"Section 4.2, Theorems 6 and 7 and Proposition 18"}],"minor_comments":[{"comment":"The phrase 'for any fixed score' should be replaced by a formulation that explicitly conditions on the localized density minorization condition (13) for the length result.","section":"Abstract"},{"comment":"The description of the length bound as a 'constant multiple' is imprecise because the factor 1/κ(S)_{α,x-tilde,h} is not a universal structural constant; it depends on the realized center, the bandwidth, and the score, and it can be infinite under the stated convention.","section":"Section 1, Eq. (1) and following text"},{"comment":"For the fixed-score examples, positivity of κ is verified only for the symmetric level choice (15), whereas Theorems 2 and 3 are stated for arbitrary α−<α<α+. The paper should clarify whether the main fixed-score length theorem is intended for those symmetric levels or whether the user must verify (13) separately for other level choices.","section":"Section 3.3, Proposition 4 and Lemma 1"},{"comment":"The real-data procedure is carefully disclosed as a hybrid, but referring to it simply as 'RLCP' in the decile figures and tables may mislead readers; a label such as 'RLCP-hybrid' would make the distinction from formal RLCP clearer.","section":"Section 5 and Appendix M"},{"comment":"In the displayed expression for the decomposition proxy, the placeholder 'slow' appears where κ(S⋆) is intended; this should be corrected.","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The mathematical core appears sound and the coverage theorem is genuinely uniform and unaffected by the κ-vacuity issue. The main obstacle to acceptance is the mismatch between the abstract's unqualified length claim and the actual hypotheses of Theorems 2 and 6; this is fixable by adding the density-floor condition to the statements and revising the abstract and introduction accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution is the mode of control: a single high-probability calibration event that simultaneously bounds conditional coverage and oracle-relative length for every x in the realized localization ball. That is not in Min et al., and it is not just a repackaging—it changes what you can claim after seeing the auxiliary center. The bias-variance decomposition into h^beta plus the local calibration term is clean, and the learned-score theorems that propagate uniform score error linearly are a sensible extension. The proofs are detailed and internally consistent; I checked the main quantile-inversion lemmas and they hold together. The experiments are appropriately modest: rate diagnostics, honestly labeled, with the hybrid RLCP caveat spelled out.\n\nThe soft spot is the length bound when kappa(S) = 0. The theorems are formally true under the 1/0=infinity convention, but then they say nothing. The paper itself acknowledges this right after Lemma 1, so the math is not wrong—but the abstract's \"For any fixed score... length error relative to the oracle\" is too strong without adding \"when the localized density minorization holds.\" That is a presentation flaw, not a mathematical one, and it is easily fixed.\n\nA second, minor issue: the real-data experiments are a hybrid of RLCP and fallback rules, and the fallback frequency is nontrivial (up to 16% on concrete). The authors are transparent about this, but readers skimming the tables may miss that the headline real-data numbers are not formal RLCP.\n\nOverall, this is a serious theory paper. The assumptions are standard for local nonparametrics, the comparison to prior work is fair, and the proof quality is high. With a revised abstract and clearer labeling of the hybrid experiments, it is publishable. I would send it to a good statistical ML venue for full review.","headline":"Finite-sample uniform-local guarantees for RLCP with an honest fixed-score analysis; the length bound is vacuous when the localized score density vanishes, and the abstract should say so.","tokens_in":64232,"tokens_out":1173,"would_cite":true,"duration_ms":16201,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G15","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Conditional on the realized localization center, RLCP bounds the conditional-coverage gap and the oracle-relative length error uniformly over the localization ball with probability at least 1−δ over calibration, at the classical…","keywords":["conformal prediction","conditional coverage","distribution-free inference","finite-sample guarantees","localization","quantile regression","randomly localized conformal prediction"],"falsifier":"Run RLCP at the balanced bandwidth $h_\\star \\asymp (\\log(4n)/(n p_{\\min}))^{1/(2\\beta+d)}$ on a score that satisfies the paper's A1–A3 but whose conditional score density vanishes on a band around the $(1-\\alpha)$-quantile (for instance, a residual score with a gap in the noise density): if the realized oracle-relative length error does not shrink at the predicted $n^{-\\beta/(2\\beta+d)}$ rate, or stays bounded while the coverage gap shrinks, then the length claim fails precisely in the regime where $\\kappa^{(S)}_{\\alpha,\\tilde{x},h} = 0$ makes the theorem's bound vacuous.","tokens_in":63155,"feed_emoji":"📏","tokens_out":22640,"duration_ms":193976,"temperature":0.7,"pith_summary":"Conformal prediction guarantees marginal coverage on average over covariates, but that average can hide severe under-coverage in specific regions, and exact distribution-free conditional coverage is known to be unattainable in finite samples. Randomly localized conformal prediction (RLCP) attacks this gap by drawing an auxiliary localization center and reweighting calibration observations near it, while preserving the marginal guarantee. This paper proves finite-sample, high-probability control of the realized RLCP set: conditional on the drawn center, simultaneously for every covariate in the localization ball, the conditional-coverage gap and the interval-length error relative to the infeasible score oracle are bounded by explicit rates. Those rates decompose into a localization bias of order $h^\\beta$ and a calibration term driven by the local effective sample size $n p_{\\min} h^d$, which makes the bandwidth bias–variance trade-off explicit and, when balanced, yields the classical nonparametric rate $n^{-\\beta/(2\\beta+d)}$. For scores learned on an independent fold that target a pivotal score, the localization bias disappears and the learning error enters linearly.","feed_headline":"Localized conformal sets get finite-sample local coverage bounds","feed_subtitle":"Conditional on the realized center, coverage and length errors shrink at the classical nonparametric rate.","key_machinery":"Three devices carry the argument. The reverse-law representation: after the auxiliary center $\\tilde{x}$ is drawn from the kernel-smoothed law, the test covariate, conditional on $\\tilde{x}$, has exactly the localized distribution $\\Lambda_{\\tilde{x},h}(du) = (w_{\\tilde{x},h}(u)/Z_{\\tilde{x},h})P_X(du)$, so RLCP is weighted conformal prediction under covariate shift and inherits finite-sample marginal validity. The calibration event $\\Omega^{\\mathrm{cal}}_{n,\\tilde{x},h,\\delta}$: a single event of probability at least $1-\\delta$ on which the kernel-weighted empirical score CDF concentrates uniformly and the test-point weight is uniformly small, so every bound holds simultaneously for all $x \\in B(\\tilde{x}, h) \\cap \\mathcal{X}$. The inversion devices: the localized score-density floor $\\kappa^{(S)}_{\\alpha,\\tilde{x},h}$ (the essential infimum of the localized score density on the quantile window) converts level error into quantile error, the length map $q \\mapsto |C^{(S)}(x,q)|$ converts quantile error into length error, and Hölder continuity of the conditional score CDF converts the mismatch between the local mixture and the target covariate into the $h^\\beta$ bias.","core_discovery":"With probability at least $1-\\delta$ over the calibration sample, conditional on the realized localization center $\\tilde{x}$, and for calibration sizes and bandwidths above explicit thresholds, the RLCP threshold deviates from the conditional score-oracle quantile by $O\\big((A^{\\mathrm{cal}}_{n,h}(\\delta) + h^\\beta)/\\kappa^{(S)}_{\\alpha,\\tilde{x},h}\\big)$, uniformly over $B(\\tilde{x}, h) \\cap \\mathcal{X}$; from this, the conditional-coverage gap is bounded by $O(A^{\\mathrm{cal}}_{n,h}(\\delta) + h^\\beta)$ (Theorem 3) and the oracle-relative length error by $O\\big(L_{\\mathrm{len}}(x)\\,(A^{\\mathrm{cal}}_{n,h}(\\delta) + h^\\beta)/\\kappa^{(S)}_{\\alpha,\\tilde{x},h}\\big)$ (Theorem 2), where $A^{\\mathrm{cal}}_{n,h}(\\delta) = \\sqrt{\\log(4/\\delta)/(n p_{\\min} h^d)} + \\log(4/\\delta)/(n p_{\\min} h^d)$ captures the local effective sample size and $\\kappa^{(S)}_{\\alpha,\\tilde{x},h}$ is the localized score-density floor. The first term in the rates is the localization bias coming from Hölder regularity of the conditional score law; the others are calibration error, and balancing them at $h_\\star \\asymp (\\log(4n)/(n p_{\\min}))^{1/(2\\beta+d)}$ gives the classical pointwise nonparametric rate $n^{-\\beta/(2\\beta+d)}$ up to logarithmic factors. When the score is learned on an independent fold and targets a pivotal population score — one whose $(1-\\alpha)$-quantile is the same at every covariate, as with conformalized quantile regression and the probability-integral-transform score — the localization bias vanishes, and Theorems 6 and 7 deliver the same uniform local control with the error split into calibration error plus a linearly entering uniform score-estimation error.","pith_inferences":["A design lesson the paper leaves implicit: when a strong, near-pivotal score is available, the bandwidth should be pushed as large as the density-floor constraint allows rather than tuned to the fixed-score optimum, since localization then mainly buys effective sample size while the score itself carries covariate adaptivity.","The vacuous-kappa caveat points to a concrete stress test: scores whose conditional density is tiny at the target quantile (for example, extreme-quantile regions in heavy-tailed responses) should show realized length errors far above the predicted rate, and mapping where the degradation begins would delineate the exact class of scores for which the oracle-tracking claim holds.","The calibration-versus-localization decomposition is the same structure that governs locally weighted regression, suggesting the $n^{-\\beta/(2\\beta+d)}$ rate is the natural minimax benchmark for any locally weighted conformal procedure; the pivotal-score result further predicts that with a well-learned score, wide-bandwidth (near-marginal) calibration is nearly optimal, since localization adds no "],"forward_implications":["Balancing the two error sources pins the optimal bandwidth at $h_\\star \\asymp (\\log(4n)/(n p_{\\min}))^{1/(2\\beta+d)}$, where both the coverage gap and the length error attain the classical rate $n^{-\\beta/(2\\beta+d)}$ up to logarithmic factors (equations (19)–(20)).","Because the calibration event is sample-specific and common to the whole ball, the guarantee is simultaneous: no union bound over test covariates and no averaging over the auxiliary randomization is required for uniformity over $B(\\tilde{x}, h) \\cap \\mathcal{X}$.","Under a pivotal target score the localization bias disappears and the bandwidth enters only through the effective calibration size $n p_{\\min} h^d$, so the score-estimation error enters linearly: improving the learned score sharpens the localized oracle comparison at exactly the training rate.","The assumptions, including the density-floor condition via Lemma 1, are verified for residual, conformalized-quantile-regression, and distributional scores, so the bounds apply to the standard score constructions in practice."],"supporting_citations":[{"why":"the randomly localized conformal construction — auxiliary-center randomization, reverse conditional law, weighted-conformal marginal validity — whose realized-set behaviour this paper analyses.","marker":"[17]"},{"why":"the closest prior non-asymptotic conditional-miscoverage theory for weighted-quantile conformal methods; supplies the decomposition and pointwise rates that this paper's uniform-local oracle bounds complement and sharpen.","marker":"[22]"},{"why":"defines conformalized quantile regression, the canonical learned pivotal score whose covariate-independent target threshold makes Theorems 6–7 apply.","marker":"[24]"},{"why":"introduces the distributional (probability-integral-transform) score, the second pivotal example, with the sup-norm estimation framework used to instantiate the learning assumption.","marker":"[9]"},{"why":"the exchangeability-based marginal conformal guarantee that the RLCP construction preserves and that the paper's local guarantees are designed to complement.","marker":"[27]"},{"why":"supplies the covariate-support regularity condition behind the design-density floor and the uniform lower bound on the local effective calibration size.","marker":"[1]"},{"why":"provides the uniform Bahadur expansion for local-polynomial quantile estimators, which yields the score-estimation rate used in the conformalized-quantile-regression example.","marker":"[16]"}],"fun_headline_variants":["Finite-sample bounds on conditional coverage for localized conformal sets","Localized conformal prediction: finite-sample conditional guarantees","RLCP achieves finite-sample local coverage and length control","Explicit finite-sample bounds for localized conformal prediction","High-probability control of coverage and length for RLCP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the localized score distribution keeps strictly positive density across the quantile window; the paper's own passage before Theorem 2 concedes that this is not guaranteed by the main assumptions, and when that density floor is zero the stated length bound is formally true but vacuous, so the oracle-length guarantee rests on the separate shape-function condition of Lemma 1.","fun_headline_variants_meta":{"raw":{"variants":["Finite-sample bounds on conditional coverage for localized conformal sets","Localized conformal prediction: finite-sample conditional guarantees","RLCP achieves finite-sample local coverage and length control","Explicit finite-sample bounds for localized conformal prediction","High-probability control of coverage and length for RLCP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000621,"raw_usage":{"total_tokens":2998,"prompt_tokens":1186,"completion_tokens":1812,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":802,"completion_tokens_details":{"reasoning_tokens":1732}},"tokens_in":802,"tokens_out":1812,"duration_ms":11882,"temperature":1.0,"reasoning_tokens":1732,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:36:38.196393+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run RLCP at the balanced bandwidth $h_\\star \\asymp (\\log(4n)/(n p_{\\min}))^{1/(2\\beta+d)}$ on a score that satisfies the paper's A1–A3 but whose conditional score density vanishes on a band around the $(1-\\alpha)$-quantile (for instance, a residual score with a gap in the noise density): if the realized oracle-relative length error does not shrink at the predicted $n^{-\\beta/(2\\beta+d)}$ rate, or stays bounded while the coverage gap shrinks, then the length claim fails precisely in the regime where $\\kappa^{(S)}_{\\alpha,\\tilde{x},h} = 0$ makes the theorem's bound vacuous.","supporting_citations":[{"cited_title":"A Unified Theory of Conditional Coverage in Conformal Prediction with Applications","cited_arxiv_id":"2605.11602","evidence_quote":"the closest prior non-asymptotic conditional-miscoverage theory for weighted-quantile conformal methods; supplies the decomposition and pointwise rates that this paper's uniform-local oracle bounds complement and sharpen."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines conformalized quantile regression, the canonical learned pivotal score whose covariate-independent target threshold makes Theorems 6–7 apply."},{"cited_title":"Distributional conformal prediction","cited_arxiv_id":"1909.07889","evidence_quote":"introduces the distributional (probability-integral-transform) score, the second pivotal example, with the sup-norm estimation framework used to instantiate the learning assumption."},{"cited_title":"Uniform bias study and Bahadur representation for local polynomial estimators of the conditional quantile function.Econometric Theory, 28(1):87–129,","cited_arxiv_id":null,"evidence_quote":"provides the uniform Bahadur expansion for local-polynomial quantile estimators, which yields the score-estimation rate used in the conformalized-quantile-regression example."}],"review_version":1}