{"id":"a8c2cb13-9184-49a0-9d7b-5e38fb334682","arxiv_id":"2507.15741","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper proposes conformal and kNN prediction-region algorithms for regression with metric-space-valued responses, with consistency guarantees and a dimension-free radius result under metric homoscedasticity.","lead":"This paper introduces uncertainty quantification methods that produce prediction balls around regression estimates when the response is a random object in a metric space, such as a probability distribution or a graph. The methods include a split-conformal procedure with finite-sample coverage guarantees under a new homoscedasticity notion, and a kNN-based adaptive procedure for heteroscedastic data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is metric homoscedasticity: Algorithm 2's marginal coverage is standard split conformal, but its efficiency and rate claims rest entirely on Definition 1, and the rate bound (Proposition 4) is deferred and additionally requires a uniform Lipschitz condition on the…","rationale":"The paper's finite-sample marginal coverage guarantee (Proposition 2) is standard and correct, and I do not dispute it. The weakness is that the advertised efficiency and fast-rate results are not independently supported in the provided text, and their only stated engine is a homoscedasticity condition that is strong and can fail in the very metric-space applications the paper targets. The reader identified homoscedasticity as the weakest assumption; I agree and sharpen the issue by pointing to the deferred Proposition 4 and the unquantified Lipschitz dependence on G*, which is where a hidden assumption would enter. I therefore recommend keeping the conditional verdict: accept subject to verification of the appendix and clarification of the scope of the rate claim. No change to the reader's verdict is needed.","tokens_in":25110,"tokens_out":11044,"duration_ms":136119,"concrete_test":"Retrieve the appendix and verify the proof of Proposition 4 in the homoscedastic case: check that the constant C is bounded uniformly over \\hat m and n under Assumptions 1-2, and that the bound follows without an unstated condition such as uniform integrability of d2(Y,m(X)) or a uniform bound on the density of residuals around \\hat m. If the proof needs an extra assumption, the rate claim in the abstract should be weakened. As a complementary numerical check, in Settings 1-3 replace the global quantile by an oracle radius computed from an independent large sample and compare worst-slice coverage; if the oracle radius recovers 1-alpha while Algorithm 2 does not, the efficiency claim is tied to exact homoscedasticity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central non-standard claim is that Algorithm 2 converges to the oracle ball at rates driven by the Fréchet-mean error and a single scalar quantile. This is not a consequence of split conformal validity, which is distribution-free and holds even under heteroscedasticity; the efficiency claim depends entirely on Definition 1. If homoscedasticity fails, Remark 2 concedes inconsistency, and Table 1 (Settings 1-3) shows the homoscedastic method's worst-slice coverage falling well below nominal. So the paper's main efficiency result is conditional on a strong, hard-to-verify distributional assumption. The proof support is also incomplete: Theorem 3 proves only consistency, and the rate statement is Proposition 4, whose proof is in an appendix not included in the provided text. Proposition 4 further assumes a uniform Lipschitz constant for G(t,x) and for G*(t,x) = P(d2(Y,\\hat m(X)) ≤ t | X = x); no argument is given that G* inherits a finite L uniformly in \\hat m and n for non-Euclidean or discrete response spaces, where the paper itself acknowledges ties and atoms can occur. Without that, the claimed fast rates are not established. The finite-sample marginal coverage guarantee of Proposition 2 stands; the convergence-to-oracle and efficiency part is the insecure pillar.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops uncertainty quantification algorithms for regression with responses in a separable metric space. In the homoscedastic case, Algorithm 2 estimates the conditional Fréchet mean on a training split and a single global radius from calibration scores, yielding a split-conformal prediction ball; Proposition 2 gives the standard finite-sample marginal coverage guarantee, and Theorem 3/Proposition 4 claim consistency and a rate bound under a new metric notion of homoscedasticity (Definition 1). For heteroscedastic data, Algorithm 3 estimates local radii via k-nearest-neighbor calibration with a data-driven choice of k, with consistency stated in Theorem 5 and Corollary 6 and an error bound in Proposition 7. A sequential extension for metric-space-valued time series is sketched in Section 2.4. Simulations and applications to distribution-valued glucose data, graph Laplacians, and handwriting shapes illustrate the methodology.","tokens_in":25352,"tokens_out":6536,"duration_ms":73146,"significance":"If the stated results hold, the paper would provide a useful and computationally light alternative to depth-profile conformal methods for metric-space responses. The finite-sample marginal coverage guarantee of Proposition 2 is correct and standard, and the modular center-radius framework is attractive for large-scale non-Euclidean data. The proposed kNN local-radius procedure with a calibration-based selection rule is a reasonable practical contribution, and the empirical comparisons with Zhou and Müller (2025) are informative. The main theoretical novelty—fast convergence to the oracle ball under metric homoscedasticity—is conditional on a strong, hard-to-verify distributional assumption, and the rate proof is deferred to an appendix not available in the reviewed text. The paper is honest about some limitations (Remarks 2 and 3), which supports a constructive major revision rather than rejection.","major_comments":[{"comment":"The efficiency and oracle-convergence claims for Algorithm 2 are entirely conditional on metric homoscedasticity (Definition 1). This is acknowledged in Remark 2, where the authors note that under heteroscedasticity the algorithm may not be consistent, and Table 1 quantifies the consequence: in Setting 1 with n=1000 and 1-alpha=0.50, the homoscedastic method attains worst-slice coverage 0.176 against a nominal level of 0.50. Because Definition 1 is a property of the unknown joint distribution and no diagnostic or sensitivity check is provided, the paper's headline fast-rate claim is narrower than the abstract suggests. The authors should either provide a practical test for homoscedasticity or explicitly frame the convergence and rate results as conditional guarantees and discuss how a practitioner could assess the assumption.","section":"Section 2.1, Definition 1, Remark 2, Table 1"},{"comment":"Proposition 4, which is the basis for the claimed fast rates, has its proof deferred to an appendix that was not included in the version under review. More importantly, the bound requires G*(t,x)=P(d2(Y,\\hat m(X))<=t | X=x) to be uniformly Lipschitz in t for all x, uniformly in \\hat m. This is only justified in the Euclidean additive-noise setting of Remark 4; in general metric or discrete response spaces, ties and atoms (acknowledged in Remark 3) mean the Lipschitz condition need not hold. Without an argument that G* inherits a finite Lipschitz constant in the non-Euclidean cases covered by the paper, the fast-rate claim is not established as stated.","section":"Section 2.1.1, Proposition 4, Assumption 2, Remark 4"},{"comment":"The sequential extension for time series is presented only as a sketch: the algorithm, the full assumptions, the proof, and the multistep extension are all in Supplemental Material C, which is not available in the review text. The asymptotic time-average calibration claim (1/n) sum_t P(Y_{t+1} in \\hat C_t) -> 1-alpha is therefore not verifiable from the manuscript. The authors should include the supplement in the review version or substantially expand the main-text treatment if this is to count as a contribution.","section":"Section 2.4 and Supplemental Material C"}],"minor_comments":[{"comment":"The title contains a typo: 'Conformal andkNN' should read 'Conformal and kNN'.","section":"Title"},{"comment":"The text refers to 'see Section 3 for details' while already in Section 3; this should be a reference to Section 2.2.","section":"Section 3.1"},{"comment":"Algorithm 3 depends on a user-specified tolerance epsilon (set to 0.05 in the experiments); no guidance is given on how epsilon should scale with n or with the metric space, and the sensitivity of the selected k to epsilon is not reported.","section":"Algorithm 3, Step 6"},{"comment":"The abbreviation ISCE is used in the caption before it is defined in Section 3.1; consider defining it in the caption.","section":"Table 1"},{"comment":"Several references contain formatting artifacts (for example, 'V ovk' with stray spaces and incomplete arXiv identifiers); these should be cleaned.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core split-conformal guarantee (Proposition 2) is correct, and the applied comparisons are useful. The main risk is that the rate and oracle-convergence contributions rest on a strong assumption and on proofs not included in the reviewed text; the G* Lipschitz issue in particular needs a concrete argument for general metric and discrete response spaces. If the appendix supplies these details and the authors reframe the claims as conditional on Definition 1, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on Lugosi–Matabuena. The genuinely new thing is the metric homoscedasticity definition: P(Y in B(m(x),r)|X=x) = phi(r). Under that assumption, the prediction-radius problem collapses to estimating a single scalar quantile, which is the right kind of simplification. The kNN local-quantile procedure with a data-driven k is a sensible practical contribution, and the sequential extension is a nice extra. The paper is also honest: Remark 2 concedes inconsistency under heteroscedasticity, Remark 7 says the adaptive method has no finite-sample guarantee, and the shape-analysis section explicitly disclaims conformal validity.\n\nWhat holds up: Proposition 2 is textbook split-conformal and is correct. The empirical comparisons are fairly presented; the homoscedastic method's worst-slice coverage in heteroscedastic settings is poor, as expected, and the adaptive method's WSC is good in the nonlinear heteroscedastic settings. The authors do not oversell.\n\nThe soft spots are real but not concealed. The entire efficiency and fast-rate story for Algorithm 2 rests on Definition 1, which is strong and hard to verify in non-Euclidean or discrete response spaces. Proposition 4's rate bound is deferred to an appendix and additionally assumes a uniform Lipschitz condition on G* that is not argued to hold when the response space has atoms or ties; the paper itself acknowledges ties can occur. So the convergence-to-oracle part is conditional, while the marginal coverage guarantee stands. The k-selection criterion uses in-sample coverage, which is a mild optimism risk, though the simulations use a separate evaluation sample and look reasonable.\n\nI agree with the reader's conditional verdict. The central finite-sample claim is solid; the efficiency claims are conditional on a strong assumption, and the proofs of the main theorems are not fully checkable from the provided text. That said, the paper frames these limitations clearly, which is more than many papers do.\n\nWho should read it: people working on conformal prediction for object data, digital health applications with distributional responses, and metric-space regression. It is a serious contribution to that subfield. I'd send it to a referee, and I'd probably cite it if I work on metric-space UQ. Recommendation: engage with it; the homoscedasticity definition and the honest treatment of the adaptive method's limitations are worth the referee time.","headline":"A genuinely useful metric-space UQ framework whose headline efficiency claim rests on a strong homoscedasticity assumption, but the paper is honest about it and deserves refereeing.","tokens_in":25894,"tokens_out":2660,"would_cite":true,"duration_ms":23908,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G15","62G08","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under a metric-space notion of homoscedasticity, split conformal prediction on Fréchet regression gives finite-sample coverage and fast rates, with a kNN variant for heteroscedastic data.","keywords":["uncertainty quantification","conformal prediction","metric spaces","Fréchet mean","k-nearest neighbors","homoscedasticity","prediction regions","metric-space time series"],"falsifier":"Simulate a heteroscedastic metric-space regression such as $Y=m(X)+\\sigma(X)\\varepsilon$ with varying $\\sigma(X)$, then test whether $P(Y\\in B(m(x),r)\\mid X=x)$ differs across two covariate values; equivalently, run Algorithm 2 and measure worst-slice conditional coverage—if the high-variance slice remains below $1-\\alpha$ as $n$ grows, the homoscedasticity-driven consistency claim for the global-radius procedure is falsified.","tokens_in":24882,"feed_emoji":"📐","tokens_out":6709,"duration_ms":65697,"temperature":0.7,"pith_summary":"The paper claims that uncertainty quantification for regression with responses in a general metric space reduces to estimating two objects: the conditional Fréchet mean, which serves as the center of the prediction region, and a radius. Its central move is a notion of homoscedasticity adapted to metric spaces: the conditional probability that the response falls in a ball of radius $r$ around the Fréchet mean is the same function $\\varphi(r)$ for every covariate value. Under that assumption, a split-conformal algorithm using the distance to an estimated Fréchet mean as the conformity score yields finite-sample marginal coverage $P(Y\\in \\hat C_\\alpha(X))\\ge 1-\\alpha$ and converges to the oracle ball at rates controlled by the mean-estimation error and one scalar quantile. For heteroscedastic data, the paper proposes a kNN procedure with a data-driven neighborhood size that produces locally adaptive radii with consistency guarantees but without the finite-sample conformal guarantee. If the metric homoscedasticity premise holds, these methods provide distribution-free, computationally cheap prediction regions for objects such as probability distributions and graph Laplacians.","feed_headline":"One quantile gives conformal coverage on metric spaces","feed_subtitle":"A metric-space homoscedasticity definition yields finite-sample prediction balls, with a kNN variant for heteroscedastic data.","key_machinery":"The load-bearing object is the metric notion of homoscedasticity (Definition 1), paired with a two-step center–radius estimator. The center is the conditional Fréchet mean $m(x)=\\arg\\min_y E(d_1^2(Y,y)\\mid X=x)$; the radius is a quantile of the pseudo-residual $r=d_2(Y,\\hat m(X))$. In the homoscedastic case the radius is a single global empirical quantile of calibration distances, which is what makes the rate independent of the predictor dimension. In the heteroscedastic case the radius becomes a local empirical quantile over kNN neighborhoods in the predictor metric, with the neighborhood size $k$ selected to keep both global and local coverage deviations small. The two distances $d_1$ and $d_2$ may differ, so the geometry used for fitting the center need not match the geometry that defines the prediction balls.","core_discovery":"On the paper's own terms, the central discovery is that once the distribution of $(X,Y)$ is homoscedastic with respect to the conditional Fréchet mean, meaning $P(Y\\in B(m(x),r)\\mid X=x)=\\varphi(r)$ for all $x$, the oracle prediction region $C_\\alpha(x)=B(m(x),r(x))$ can be estimated by a single global radius: the $(1-\\alpha)(1+1/n_2)$-quantile of the calibration distances $d_2(Y_i,\\hat m(X_i))$. Algorithm 2 then inherits the split-conformal finite-sample guarantee $P(Y\\in \\hat C_\\alpha(X))\\ge 1-\\alpha$, and under mild consistency of the Fréchet mean estimator the expected symmetric-difference error to the oracle region converges to zero at a rate that splits into the mean-estimation error and the error of a single unconditional quantile. The same center–radius decomposition is carried into the heteroscedastic case: Algorithm 3 replaces the global quantile by a local empirical quantile over $k$ nearest neighbors in the predictor space, with $k$ chosen by a criterion that keeps both global and worst-local coverage deviations below a tolerance $\\varepsilon$, and the paper proves consistency for both deterministic and data-dependent $k$. These results deliberately avoid smoothness assumptions and work with any regression algorithm that estimates a conditional Fréchet mean. The paper also extends the kNN radius construction to stationary ergodic metric-space time series via nearest-neighbor expert aggregation, with asymptotic time-average calibration as the stated guarantee.","pith_inferences":["Editorial inference: the homoscedasticity definition is directly testable—if a practitioner estimates $P(Y\\in B(m(x),r)\\mid X=x)$ at several covariate values and the curves differ materially, the efficiency guarantee of Algorithm 2 should not be relied on, and the paper's own Remark 2 concedes the algorithm may then be inconsistent.","Editorial inference: the center–radius decomposition sketched in Remark 7 points toward a conformalized quantile-regression extension that would give the heteroscedastic kNN procedure a finite-sample coverage guarantee, not just consistency.","Editorial inference: the freedom to use different metrics for center and radius could be exploited to build prediction regions with interpretable shapes, such as functional bands under the supremum metric, even when the response geometry is more naturally Wasserstein or graph-based.","Editorial inference: the expert-aggregation time-series construction suggests a route to online or sequential conformal prediction for metric-space objects under dependence, a setting where finite-sample conformal guarantees are currently largely absent."],"forward_implications":["Under metric homoscedasticity, split conformal prediction with a Fréchet mean estimator yields non-asymptotic marginal coverage $\\ge 1-\\alpha$ for any regression algorithm and metric space, with no smoothness assumptions.","The prediction-radius estimation rate is governed by a single unconditional scalar quantile plus the Fréchet mean estimation error, so it does not suffer the usual curse of dimensionality in the predictor.","In heteroscedastic metric spaces, the kNN procedure adapts the radius locally and is consistent, with finite-sample bounds separating the center error and the radius error.","The time-series extension provides asymptotic time-average calibration for metric-space-valued stationary ergodic sequences without imposing mixing or smoothing assumptions.","After the center function is estimated, the pipeline needs only nearest-neighbor searches and empirical quantiles, so it scales to large datasets in practice."],"supporting_citations":[{"why":"Supplies the conformal prediction framework and the finite-sample exchangeability guarantee that Algorithm 2 builds on.","marker":"Vovk et al. [2005]"},{"why":"Defines the global Fréchet regression model used for center estimation and establishes its consistency conditions.","marker":"Petersen and Müller [2019]"},{"why":"Is the concurrent metric-space conformal method against which the center–radius approach is compared and which it generalizes.","marker":"Zhou and Müller [2025]"},{"why":"Provides conformalized quantile regression whose calibration strategy underpins the heteroscedastic extension discussed in Remark 7.","marker":"Romano et al. [2019]"},{"why":"Gives nearest-neighbor conformal prediction rates and serves as the kNN uncertainty-quantification baseline.","marker":"Györfi and Walk [2019]"},{"why":"Supplies the relaxed conditional-calibration perspective that motivates the admissible-$k$ selection criterion in Algorithm 3.","marker":"Gibbs et al. [2025]"},{"why":"Provides the sequential quantile expert-aggregation framework used for the metric-space time-series extension.","marker":"Biau and Patra [2011]"},{"why":"Establishes universal consistency of kNN regression on metric spaces, used as a center estimator for the proposed procedures.","marker":"Cohen and Kontorovich [2022a]"}],"fun_headline_variants":["A single radius yields conformal prediction balls","Conformal prediction with one global quantile on metric spaces","kNN adds local calibration to metric conformal prediction","Metric-space conformal UQ: one quantile, finite-sample guarantees"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is metric homoscedasticity: the probability $P(Y\\in B(m(x),r)\\mid X=x)$ is the same function of $r$ for every covariate value $x$, and the paper's own Remark 2 concedes that without it Algorithm 2 may not be consistent, leaving only marginal coverage.","fun_headline_variants_meta":{"raw":{"variants":["A single radius yields conformal prediction balls","Conformal prediction with one global quantile on metric spaces","kNN adds local calibration to metric conformal prediction","Metric-space conformal UQ: one quantile, finite-sample guarantees"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000485,"raw_usage":{"total_tokens":2444,"prompt_tokens":1049,"completion_tokens":1395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":1329}},"tokens_in":665,"tokens_out":1395,"duration_ms":11762,"temperature":1.0,"reasoning_tokens":1329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:25:02.967816+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a heteroscedastic metric-space regression such as $Y=m(X)+\\sigma(X)\\varepsilon$ with varying $\\sigma(X)$, then test whether $P(Y\\in B(m(x),r)\\mid X=x)$ differs across two covariate values; equivalently, run Algorithm 2 and measure worst-slice conditional coverage—if the high-variance slice remains below $1-\\alpha$ as $n$ grows, the homoscedasticity-driven consistency claim for the global-radius procedure is falsified.","supporting_citations":[],"review_version":1}