{"id":"c4eb2932-0371-4d75-9c28-be1be4eacb0b","arxiv_id":"2412.08533","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A nearest-neighbor control variates estimator gives faster integral approximation rates for random-design functional data and shorter prediction/confidence intervals in simulations, with a proved noisy-case CLT and a conjectured noiseless interval.","lead":"The paper applies a control-variates Monte Carlo method, using leave-one-out nearest neighbors, to estimate integrals of random functions in functional data analysis, and shows it converges faster than sample means or Riemann sums in the noiseless case. It also gives confidence and prediction intervals that are shorter in simulations, though the noiseless interval's validity is only conjectured.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Noiseless inference rests on an unproven subsampling conjecture, and the recommended half-sample rule M*=M/2 violates the standard M*/M -> 0 condition, leaving PI coverage unsupported even if a CLT held.","rationale":"The reader's weakest assumption identified the explicit conjecture on the asymptotic level of PI_{1-delta}; I agree that this is the most load-bearing gap. My stress-test sharpens the concern: the practical recommendation M*=floor(M/2) is not merely lacking a proof, it conflicts with the standard condition m/M -> 0 used to make the centering term negligible in subsampling theory. For the sample mean, this exact mechanism produces the finite-population correction factor (1 - m/M) in the variance, and Table 1's 'ms' column shows the resulting undercoverage. The same decomposition applies to the control-neighbor estimator if it has an asymptotic distribution at rate M^{-r}; the term M*^r(\\hat I(T)-I) is of the same order as the leading term unless M*/M -> 0. The paper does not derive a finite-population correction for the nearest-neighbor weights, so the half-sample prediction interval may be miscalibrated even under the conjectured CLT. This does not undermine the noisy-case CLT (Proposition 2), which is a solid contribution and does not use subsampling. The verdict should remain CONDITIONAL: either prove the subsampling level for a rule satisfying M*/M -> 0 (or with a correction for M*=M/2), or soften the noiseless inference claims to heuristics. The proposed simulation test would separate a genuine miscalibration from a benign finite-sample effect.","tokens_in":23908,"tokens_out":10153,"duration_ms":113038,"concrete_test":"Simulate Algorithm 1 for a known deterministic beta-Holder integrand, e.g. phi(t)=|t-1/2|^{1/2} on [0,1], with iid uniform design points, B=1000 subsamples, nominal level 95%, and at least 2000 replications. Compare empirical coverage for three subsample rules: M*=floor(M/2), M*=floor(M/log M), and M*=floor(M^{0.8}), across M=200, 400, 800. If the half-sample rule's coverage deviates systematically from nominal while the M/log M rule is on target, the recommended M* is the source of miscalibration; if all three rules give nominal coverage, this concern is refuted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central novel claim for the noiseless case is the prediction interval in Algorithm 1 (Section 3.3). Its validity is explicitly left as a conjecture: \"We conjecture that the prediction interval PI_{1-delta} has the asymptotic level 1-delta under the conditions of Proposition 1 and with a suitable rule for M*\". The only rule proposed and used in all experiments is M* = floor(M/2), justified only by counting subsamples. Standard subsampling consistency requires m -> infinity and m/M -> 0. With m = M/2 the centering term does not vanish: writing r = 1/2 + beta/d, (M*)^r [\\hat I(phi(T*)) - \\hat I(phi(T))] = (M*)^r[\\hat I(phi(T*)) - I(phi)] - (M*/M)^r M^r[\\hat I(phi(T)) - I(phi)], and (M*/M)^r = 2^{-r} is bounded away from 0 while M^r[\\hat I(phi(T)) - I(phi)] is O_P(1) under the rate of Proposition 1. The subsampled statistic therefore estimates the law of a difference of two dependent O_P(1) terms, not the law of M^r(\\hat I - I). No finite-population correction is supplied, and no argument shows that the nearest-neighbor functional has the same half-sample scaling as the full sample. Table 1's 'ms' column (sample mean with subsampling) shows coverage around 0.80 instead of 0.95, consistent with this finite-population effect. Hence the noiseless \"better coverage with shorter prediction intervals\" claim and Remark 1's interval-length comparison are not established by the theory in the paper; they are supported only by the specific simulation configurations reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a control-variate estimator based on leave-one-out nearest neighbors for computing integrals of random functions observed at random design points, building on the rate result of Leluc et al. (2024). In the noiseless case it proposes prediction intervals based on an M*-out-of-M subsampling algorithm scaled at the rate M^{-1/2-β/d}; in the noisy case it derives a CLT-based confidence interval with a plug-in variance estimator. The methodology is applied to functional linear regression, fPCA scores, and functional depth, with simulations and a swimmers data analysis. The central rate claim is Proposition 1; the central inferential claims are the subsampling prediction interval of Algorithm 1 and Proposition 2.","tokens_in":24337,"tokens_out":4117,"duration_ms":49689,"significance":"The manuscript addresses a practically important problem in functional data analysis: inference for integrals of random functions under random design. If the fast rate result of Proposition 1 is accepted, the estimation contribution is useful and the noisy-case CLT in Proposition 2 provides a simple, apparently new inferential tool with an easily computable variance. The paper also ships a package, gives reproducible simulation settings, and applies the method to a real sports data set. However, the headline noiseless inference claim is not established by the paper's own theory: Algorithm 1 is explicitly justified only by a conjecture, and the recommended half-sample rule conflicts with standard subsampling conditions. The significance of the paper in its current form is therefore conditional on resolving the subsampling gap.","major_comments":[{"comment":"The noiseless prediction interval PI_{1-δ} is load-bearing for the abstract's claim of 'better coverage with shorter prediction intervals' in the noiseless setup, but its validity is explicitly left as a conjecture: the text states 'We conjecture that the prediction interval PI_{1-δ} has the asymptotic level 1-δ' immediately after noting that the estimator's convergence in distribution remains an open question. The only implemented rule, M* = floor(M/2), violates the standard subsampling condition m/M → 0. Writing r = 1/2 + β/d, the quantity whose quantiles are used is (M*)^r [\\hat I(φ(T*)) - \\hat I(φ(T))] = (M*)^r [\\hat I(φ(T*)) - I(φ)] - (M*/M)^r M^r [\\hat I(φ(T)) - I(φ)]. With M*/M = 1/2, the second term does not vanish and M^r [\\hat I(φ(T)) - I(φ)] is O_P(1) under Proposition 1, so the subsampled statistic estimates the law of a difference of two dependent O_P(1) quantities rather than the law of M^r[\\hat I(φ(T)) - I(φ)]. No finite-population correction or alternative argument is supplied. This is not a purely theoretical nuance: Table 1 reports 'ms' coverage of 0.80–0.83 instead of the nominal 0.95, consistent with the finite-population centering effect. The paper should either prove a CLT for the full-sample estimator, prove subsampling consistency under a rule with M*/M → 0, or substantially restrict the noiseless inference claims.","section":"§3.3, Algorithm 1"},{"comment":"The noisy-case CLT is a central result, but its proof as printed delegates several essential steps to the Supplementary Material and leaves two steps as conjectures for the sphere case. In Step 4, condition (32) is said to follow from a supplementary-material result and the bound max_m |M w^(NN)_m|^2 = O_P(M^a) is used with 0 < a < 1/2; in Step 5, the required negligibility of the difference between the exact and computationally efficient weights is again said to be shown in the Supplementary Material, and the text explicitly states 'We conjecture that this holds also for the case where T is the unit sphere'. Since the introduction advertises the sphere as a domain of interest, the sphere extension should either be proved or explicitly excluded from the formal statements. More generally, the main text should state which parts of Proposition 2 are proved in the Appendix and which are deferred to the Supplementary Material, so that the reader can verify the theorem without accessing external documents.","section":"§3.4, Proposition 2 and Appendix 9.2"},{"comment":"The adaptive choice β = \\hat H - log^{-2}(M) is used for the noiseless subsampling procedure, but the paper states that 'a detailed theoretical analysis of the properties of the choice (24) is beyond the scope of this paper'. This matters for the coverage claim of Algorithm 1 because the scaling factor M^{-1/2-β/d} depends on β; if an estimated β overestimates the true Hölder exponent, the interval is rescaled by a smaller factor than the actual rate and can under-cover. The paper should at least provide a transparent statement of the conditions under which the plug-in β preserves the conjectured asymptotic level, or present a sensitivity analysis in the simulations.","section":"§5, Eq. (24)"}],"minor_comments":[{"comment":"The displayed identity I(φ) = E_M[φ] - E_M{φ̃ - E_M[φ̃]} is unconventional because the second E_M inside the braces is applied to the random variables φ̃(T_m), not to the function φ̃; please clarify the conditional-expectation notation.","section":"§2, equation before (3)"},{"comment":"There is a typo in 'the more irreggular the paths are'; also the sentence 'By suitable moment conditions ... it is the possible to check' should read 'it is possible to check'.","section":"§5, text around Eq. (23)"},{"comment":"The table reports coverage for the NN method exceeding the nominal level (up to 0.99) while lengths are much shorter than the mean-based intervals; the paper should comment on whether this over-coverage indicates that the half-sample subsample variance is conservative rather than asymptotically calibrated.","section":"Table 1"},{"comment":"The column (ℓ(NN) - ℓ(m))/ℓ(m) shows both negative and positive values, but the text summarizes the results only qualitatively; a sentence explaining the direction and magnitude of the length comparison for each smoothness level would improve readability.","section":"§6.2, Table 4"},{"comment":"The remark states that the conditional-expectation fPCA method is tailored to the noisy setup, but the swimmers data are treated as noiseless; please clarify whether the comparison to Yao et al. (2005) is intended as a conceptual statement or as a numerical comparison.","section":"§7, Remark 10"}],"recommendation":"major_revision","confidential_remarks":"The paper's strongest contribution is the adaptation of the fast control-neighbor rate and the noisy-case CLT to functional data problems. The main barrier to acceptance is the noiseless subsampling inference: the manuscript itself labels the key property a conjecture, and the half-sample rule is not covered by standard subsampling theory. This is fixable in principle, but the revision needs either a real proof or a substantial narrowing of the claims. I would also ask the editor to ensure the Supplementary Material contains the deferred steps of Proposition 2, since the printed proof is not self-contained."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper imports a fast Monte Carlo rate from Leluc et al. and adds a genuinely new CLT for noisy integrands. The noiseless prediction interval is the advertised inference tool, but its validity is explicitly left as a conjecture, and the half-sample rule used everywhere makes that conjecture unlikely to hold in the stated form. The paper deserves a serious referee, but the claims need trimming or the interval needs a real proof.\n\nWhat is new and good: Proposition 2, the noisy-case CLT with a simple weight-based variance estimator, is a real contribution. The proof has the right shape—conditional Lindeberg CLT for the weighted sums, then integration over the design points—and the main cube case is credible, even if steps are deferred to the Supplement and the sphere case is conjectured. The FDA adaptation (random integrands, fPCA scores, depths) is well motivated, the simulations are extensive, and the paper is unusually honest: it states the PI conjecture explicitly and gives Leluc et al. the credit for the rate.\n\nSoft spots: the noiseless PI is not supported by the theory in the paper. The stress-test decomposition is correct: with M* = M/2, the scaled subsample statistic is a difference of two dependent OP(1) terms, not an approximation to the full-sample pivot. Standard subsampling requires M*/M -> 0; half-sampling is a bagging heuristic. So the coverage in Table 1 is simulation evidence for specific configurations, not an asymptotically justified interval. The abstract's blanket claim of \"better coverage with shorter confidence and prediction intervals\" in the noiseless case overstates what has been shown. In addition, the plug-in beta from (24) is unanalyzed; the PI scaling depends on beta, and the log^{-2}(M) correction is plausible but unproven.\n\nWho should read this: anyone in FDA who computes integrals at random design points and wants an estimator faster than the sample mean with some form of inference. The noisy-case procedure is ready to use; the noiseless PI should be treated as a heuristic until a valid subsampling rule (e.g., M* = o(M)) is proved or the claims are softened. I would send this to peer review, asking the authors to either prove the PI or substantially scale back the noiseless inference claims.","headline":"Useful FDA adaptation of control-neighbor Monte Carlo with a solid noisy-case CLT, but the noiseless prediction interval rests on a conjecture that the paper's own half-sample rule undermines.","tokens_in":24830,"tokens_out":3109,"would_cite":true,"duration_ms":33264,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62R10","62G08","62M99","62-08"],"pacs":[],"model":"deepseek-v4-flash","headline":"A leave-one-out nearest-neighbor control variate estimates integrals of random functions at the optimal Monte Carlo rate.","keywords":["functional data analysis","Monte Carlo integration","control variates","nearest neighbor","random design","Hölder regularity","functional regression","fPCA scores"],"falsifier":"Simulate a noiseless, $\\beta$-Hölder integrand on $[0,1]^d$, compute Algorithm 1 intervals with a fixed subsample size $M^*=\\lfloor M/2\\rfloor$, and measure empirical coverage over many replications as $M$ grows; if coverage persistently falls below $1-\\delta$, the conjectured level fails. Alternatively, measure the empirical RMSE of $\\hat{I}(\\varphi)$ for increasing $M$ and check whether it obeys the claimed $M^{-1/2-\\beta/d}$ slope; a flatter slope would refute the core rate bound.","tokens_in":23706,"feed_emoji":"📊","tokens_out":5960,"duration_ms":58271,"temperature":0.7,"pith_summary":"This paper claims that a standard staple of functional data analysis—estimating integrals of random functions observed at random design points—can be done at the optimal Monte Carlo rate with a control-variate trick. The proposed estimator centers the sample average at a leave-one-out nearest-neighbor approximation of the integrand, whose integral is known exactly, making the estimate unbiased with root mean squared error bounded by $C M^{-1/2-\\beta/d}$ for $\\beta$-Hölder integrands. At that speed the integration error becomes negligible relative to sampling noise, so noisy-data confidence intervals follow from a central limit theorem with a simple variance estimator, and the authors report shorter prediction intervals with good coverage in simulations. The noiseless subsampling interval, however, has only a conjectured asymptotic coverage level, since the estimator's convergence in distribution remains an open question.","feed_headline":"Nearest-neighbor estimator hits optimal Monte Carlo rate","feed_subtitle":"Control variates give unbiased integral estimates with shorter intervals and a CLT for noisy functional data.","key_machinery":"The central object is the leave-one-out nearest-neighbor control variate: for each design point $T_m$, replace the integrand $\\varphi$ by its value at the nearest other point, so $\\tilde{\\varphi}^{(m)}(t)=\\varphi(\\widehat{N}^{(m)}(t))$. Because the integral of $\\tilde{\\varphi}^{(m)}$ equals a weighted sum of $\\varphi(T_\\ell)$ over leave-one-out Voronoi cells, the control-variate estimator becomes an explicit linear integration rule $\\hat{I}(\\varphi)=\\sum_{m=1}^M w_{M,m}\\varphi(T_m)$ with weights that make the estimator unbiased. The variance reduction comes from the fact that $|\\varphi-\\tilde{\\varphi}^{(m)}|_\\infty$ shrinks like the typical nearest-neighbor distance, $M^{-1/d}$, multiplying the $M^{-1/2}$ sample-average rate and producing the $M^{-1/2-\\beta/d}$ bound. In the noisy case the same mechanism makes the remainder negligible next to the noise term, so a CLT with an explicit plug-in variance follows.","core_discovery":"The paper's central claim is that control variates built from leave-one-out nearest neighbors provide unbiased, linear estimators for integrals $I(\\varphi)=\\mathbb{E}[\\varphi(T,X(T))\\mid X]$ whose root mean squared error is $O(M^{-1/2-\\beta/d})$ for $\\beta$-Hölder integrands, the optimal Monte Carlo rate. These estimators work on cubes and spheres, inherit the regularity of the sample paths, and adapt to unknown smoothness through recent regularity estimators. In the noisy observation model the approximation error is beaten by the noise contribution, so the estimator is asymptotically normal with a variance that can be estimated directly from the design points, yielding confidence intervals that do not require knowing $\\beta$. In the noiseless case the paper constructs subsampling prediction intervals and, in simulations, these achieve nominal coverage with much shorter lengths than sample-mean intervals, although their asymptotic validity is left as a conjecture.","pith_inferences":["The paper leaves implicit that the noiseless subsampling interval would become fully rigorous if convergence in distribution of the control-neighbor estimator were established; a natural next step is to prove, or disprove, the $M^*$-out-of-$M$ subsampling level under the conditions of Proposition 1.","If the conjectured asymptotic level holds, this would provide a practical way to obtain shorter-than-CLT intervals for individual-curve functionals, potentially changing default practice in sparse functional data analysis.","A testable extension is to apply the same control-neighbor rule to compact non-Euclidean domains beyond spheres, where the leave-one-out nearest-neighbor construction and Voronoi identities still make sense but the CLT step would need re-checking.","The variance bounds rely on known moment asymptotics for Voronoi cells; on domains with non-regular design densities the rate may degrade, since the paper assumes the design density is bounded above and below."],"forward_implications":["Integral estimates for random-design functional data—regression predictions, fPCA scores, and data depths—would converge at $M^{-1/2-\\beta/d}$, faster than the sample mean's $M^{-1/2}$ and than random Riemann sums' $M^{-\\beta/d}$.","For noisy observations, confidence intervals can be built from a central limit theorem and a plug-in variance estimator, with no need to estimate the Hölder exponent $\\beta$.","If the conjectured coverage holds, noiseless prediction intervals would be asymptotically shorter than sample-mean intervals by the factor $M^{-\\beta/d}$ while retaining nominal coverage.","An adaptive choice of $\\beta$ from estimated local regularity—for example $\\hat{\\beta}=\\hat{H}-\\log^{-2}(M)$—would make the method usable without prior knowledge of trajectory smoothness.","The same rule extends to spheres via geodesic distance, with the current CLT variance results proven for the cube and only conjectured on the sphere."],"supporting_citations":[{"why":"Supplies the control-neighbors method, the linear weight representation, and the variance bound used in Proposition 1.","marker":"Leluc et al. (2024)"},{"why":"Supplies the control-functionals principle on which the proposed estimator is based.","marker":"Oates et al. (2017)"},{"why":"Establishes the optimal Monte Carlo integration rate that the proposed estimator attains.","marker":"Novak (2016)"},{"why":"Establishes the optimal rate for Monte Carlo integration, cited as the benchmark the estimator matches.","marker":"Bakhvalov (2015)"},{"why":"Supplies the moment asymptotics for Voronoi cells used in the proof of the noisy-case CLT.","marker":"Devroye et al. (2017)"},{"why":"Supplies the bounded-degree property of nearest-neighbor graphs used to control the estimator weights.","marker":"Henze (1987)"},{"why":"Supplies the regularity estimators used to choose the Hölder exponent $\\beta$ adaptively from functional data.","marker":"Wang et al. (2024)"}],"fun_headline_variants":["Nearest-neighbor CVs hit optimal MC rate for integrals","Leave-one-out NN gives unbiased, faster integral estimates","NN control variates beat sample mean for function integrals","Optimal Monte Carlo rate via nearest-neighbor estimator","Shorter intervals and CLT from NN integral estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For the noiseless claim of shorter intervals with correct coverage to hold, the subsampling prediction interval must really have asymptotic level $1-\\delta$; the paper only conjectures this, because the estimator's convergence in distribution is an open question.","fun_headline_variants_meta":{"raw":{"variants":["Nearest-neighbor CVs hit optimal MC rate for integrals","Leave-one-out NN gives unbiased, faster integral estimates","NN control variates beat sample mean for function integrals","Optimal Monte Carlo rate via nearest-neighbor estimator","Shorter intervals and CLT from NN integral estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1176,"prompt_tokens":823,"completion_tokens":353,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":275}},"tokens_in":439,"tokens_out":353,"duration_ms":4738,"temperature":1.0,"reasoning_tokens":275,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:44:13.345759+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a noiseless, $\\beta$-Hölder integrand on $[0,1]^d$, compute Algorithm 1 intervals with a fixed subsample size $M^*=\\lfloor M/2\\rfloor$, and measure empirical coverage over many replications as $M$ grows; if coverage persistently falls below $1-\\delta$, the conjectured level fails. Alternatively, measure the empirical RMSE of $\\hat{I}(\\varphi)$ for increasing $M$ and check whether it obeys the claimed $M^{-1/2-\\beta/d}$ slope; a flatter slope would refute the core rate bound.","supporting_citations":[{"cited_title":"J., Girolami, M., and Chopin, N","cited_arxiv_id":null,"evidence_quote":"Supplies the control-functionals principle on which the proposed estimator is based."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the optimal Monte Carlo integration rate that the proposed estimator attains."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the optimal rate for Monte Carlo integration, cited as the benchmark the estimator matches."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the moment asymptotics for Voronoi cells used in the proof of the noisy-case CLT."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the bounded-degree property of nearest-neighbor graphs used to control the estimator weights."},{"cited_title":"W., Patilea, V., and Klutchnikoff, N","cited_arxiv_id":null,"evidence_quote":"Supplies the regularity estimators used to choose the Hölder exponent $\\beta$ adaptively from functional data."}],"review_version":1}