{"id":"bb9a1353-734e-4d5c-b254-3d754566811b","arxiv_id":"2508.18624","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A unified self-normalized testing framework for relevant (thresholded) hypotheses in functional time series, valid under arbitrary sparse-to-dense sampling with measurement error.","lead":"This paper builds statistical tests for deciding whether a functional time series deviates from a reference by more than a meaningful amount, using discretely sampled curves with noise. The tests work across sparse to dense sampling designs and avoid estimating nuisance parameters, with applications to financial implied volatility and traffic data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption (A6) permits the B-spline bias to be of the same order as the noise at the boundary; Theorem 2.1's 'α at ∫m²=Δ' conclusion is therefore not justified under (A1)-(A6).","rationale":"The reader's weakest assumption was the unavailable sequential Gaussian approximation in Lemma S.4.1. That is legitimate but unverifiable. The concern I raise is sharper: the main text's own Assumption (A6) is insufficient for the theorem's boundary size claim, independent of the supplement. The bias in the plug-in squared-norm statistic is of order J_n^{-q*} and is not made negligible by (A6) or (*). A concrete dense-design sequence J_n=c n^{1/8} satisfies the stated assumptions and produces a non-pivotal limit, so the 'α at the boundary' conclusion cannot hold as written. This is load-bearing because the paper's headline is asymptotic size/power; if the boundary case is wrong, the size control of the relevant hypothesis test is invalid in the tested regime. The paper's own simulation guidance avoids the issue by insisting J_n≫n^{1/8}, so the framework is likely salvageable by strengthening A6 to an undersmoothing condition (or changing the O(1) to o(1)). I therefore keep the reader's CONDITIONAL verdict: the contribution is plausible but requires a correction to the stated assumptions and proof.","tokens_in":22510,"tokens_out":26382,"duration_ms":236933,"concrete_test":"Take a dense design with E(N^{-1})=o(n^{-1/8}), q*=4, and m(x)=x. Set J_n=c n^{1/8} (so A6 holds) and symbolically compute the limit of P(Tn>Δ+Q_{1-α,ǫ}V_{n,ǫ}) at ||m||²_{L2}=Δ, keeping the O(J_n^{-4}) bias term in the asymptotic expansion. If the limit differs from α, Theorem 2.1's boundary statement fails. Alternatively, check the supplement's proof of Theorem 2.1 for whether the first line of A6 is used as o(1) or an extra undersmoothing condition is silently imposed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"At the boundary ∫m²=Δ, Theorem 2.1 requires (Tn−Δ)/V_{n,ǫ} to be asymptotically pivotal. The B-spline estimator \\hat m has L2 bias of order J_n^{-q*}, so the plug-in squared-norm statistic Tn carries a first-order bias of order J_n^{-q*} (projected on m). Scaled by √n, this contributes √n J_n^{-q*} to the numerator, while the stochastic scale is τ_n(m), which condition (*) makes of order sqrt{1+J_nE(N^{-1})}. Assumption (A6), first line, states only √n J_n^{-q*} {1+J_nE(N^{-1})}^{-1/2}=O(1), i.e., the bias-to-noise ratio need not tend to zero. Concrete failure: in the dense regime with E(N^{-1})=o(J_n^{-1}), choose J_n=c n^{1/(2q*)}, e.g., q*=4, J_n=c n^{1/8}. This satisfies A6, yet the limiting boundary rejection probability becomes P(W(1)+b > Q_{1-α,ǫ} sqrt{∫t²(W(t)-tW(1))²dt}) with b=O(1) (the signed bias projection), which is not α for generic m. Hence Theorem 2.1's boundary claim is false under (A1)-(A6) as written. The simulation's stricter choice J_n≫n^{1/8} would repair this, but that undersmoothing condition is not part of A6; condition (*) does not control the bias.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified framework for testing relevant hypotheses in functional time series with discretely observed, contaminated trajectories under arbitrary sampling schemes. It covers one-sample tests on the mean function, two-sample comparisons, and single and multiple change point tests. The test statistics are based on B-spline estimates of partial mean functions, and critical values are obtained by self-normalization, avoiding estimation of long-run covariance and noise variance functions. The main theorems claim asymptotic size alpha at the boundary of the relevant null hypothesis, consistency under fixed alternatives, and nontrivial power at local alternatives with a detection rate depending on sampling frequency, yielding a sparse-to-dense phase transition. Simulations and two real-data applications are reported. All proofs are deferred to a supplementary file that is not included in the submission.","tokens_in":22838,"tokens_out":16625,"duration_ms":164654,"significance":"If the theoretical claims are correct after the corrections below, the paper would be a useful contribution: it extends relevant-hypothesis testing from fully observed functional data to discretely observed data with measurement error, covers several testing problems in one framework, and provides a self-normalizing construction that avoids nuisance estimation. The unified treatment of sparse and dense sampling and the discussion of alternative self-normalizers are valuable. The paper also reports extensive simulations across distributions and sampling schemes and two real-data analyses. However, the main technical engine, a sequential Gaussian approximation for dependent random vectors, is only mentioned by reference to an unavailable supplement, and two load-bearing statements in the main text need correction: the boundary size claim in Theorem 2.1 is not justified under Assumption (A6), and the local-alternative scaling in Theorem 2.2 is internally inconsistent with the rates quoted in Corollary 2.1. The contribution is therefore contingent on these fixes.","major_comments":[{"comment":"At the boundary integral of m squared equals Delta, the proof requires sqrt(n)(T_n - Delta)/(2 tau_n(m)) to be asymptotically pivotal after cancellation of tau_n. The B-spline approximation error of order J_n^{-q*} in L2 produces a bias in the plug-in squared-norm statistic T_n of order J_n^{-2q*}; after scaling by sqrt(n)/tau_n this contributes sqrt(n) J_n^{-2q*}/tau_n. The first line of Assumption (A6) only requires sqrt(n) J_n^{-q*} (1 + J_n E(N^{-1}))^{-1/2} = O(1), which together with condition (*) does not imply that this T_n-bias term is o(1). For instance, in the dense regime E(N^{-1}) = o(J_n^{-1}) with q* = 4, choosing J_n = c n^{1/8} satisfies (A6), but the limiting boundary rejection probability becomes P(W(1) + b > Q_{1-alpha,epsilon} sqrt(integral t^2 (W(t) - t W(1))^2 dt)) for a generally nonzero constant b, not alpha. Condition (*) does not control this bias. The theorem should either impose a genuine undersmoothing condition, for example sqrt(n) J_n^{-2q*} (1 + J_n E(N^{-1}))^{-1/2} = o(1), or explicitly account for the bias. The analogous boundary claims in Theorems 2.3, 3.1, and 3.3 inherit this issue.","section":"Section 2.1, Assumption (A6) and Theorem 2.1"},{"comment":"Under the local alternative in Theorem 2.2, the noncentrality parameter of (T_n - Delta)/V_{n,epsilon} is sqrt(n)(||m||^2 - Delta)/(2 tau_n(m)) approximately c0/(2 sqrt(n)), because tau_n(m) is of order sqrt(1 + J_n E(N^{-1})) by condition (*). This tends to zero, so the stated alternative cannot yield nontrivial power. The scaling in the theorem appears to be off by a factor sqrt(n): with ||m||^2 - Delta = c0 sqrt(1 + J_n E(N^{-1}))/sqrt(n), the noncentrality is O(1), which matches the rates quoted in Corollary 2.1. The same scaling issue appears in Theorem 3.2. Moreover, the abstract's claim that the detection rate remains n^{-1/2} even in sparse regimes is contradicted by Corollary 2.1(1), where the squared-norm detection rate is sqrt(J_n E(N^{-1})/n), which is larger than n^{-1/2} whenever J_n E(N^{-1}) tends to infinity.","section":"Section 2.1, Theorem 2.2 and Corollary 2.1"},{"comment":"The central technical result is the joint weak convergence in (2.4), which rests on the sequential Gaussian approximation for dependent random vectors stated as Lemma S.4.1. This lemma is not stated in the main text, and all proofs are deferred to a supplement that is not part of the submitted manuscript. Since Theorems 2.1 through 3.3 all depend on this approximation, the central claims cannot be verified from the submitted text. The revision should include the supplement, or at minimum a complete statement of Lemma S.4.1 and the main proof steps needed to establish (2.4).","section":"Section 2.1, joint weak convergence (2.4) and supplementary materials"}],"minor_comments":[{"comment":"There is a typo in the first line: 'releva nt' should be 'relevant'.","section":"Abstract"},{"comment":"The moment condition on the measurement errors is written with E|epsilon_i|^{r1}; it should clearly be E|epsilon_{ij}|^{r1} with the absolute value displayed, and the notation should be aligned with the definitions of r1 and r.","section":"Section 2.1, display of Assumption (A6)"},{"comment":"The symbol rendered as 'greaterorsimilar' should be typeset as, for example, \\gtrsim or 'greater than or asymptotically similar to', to avoid confusion in conditions (*), (**), and (⋆).","section":"Throughout, notation"},{"comment":"'referred as' should be 'referred to as' in both statements.","section":"Corollary 2.1 and Corollary 3.1"},{"comment":"The text refers to the AU.SHF implied volatility data, while Figure 9 says 'AU.SHX'; the abbreviation should be made consistent.","section":"Section 5.1 and Figure 9"},{"comment":"The phrase 'under some regular conditions' before display (2.4) is vague; since Assumptions (A1)-(A6) are already stated, the display should refer to those assumptions directly.","section":"Just before Theorem 2.1"}],"recommendation":"major_revision","confidential_remarks":"The main text is not self-contained for verification because the supplement is absent, and the two statement errors identified above are load-bearing. Both are fixable, so I am not recommending rejection. I would ask the authors to include the supplement in the next revision and to correct the boundary bias condition and the local-alternative scaling before the paper is considered further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper extends relevant hypothesis testing to functional time series observed with discrete sampling and measurement error, covering one-sample, two-sample, and change-point problems under a unified self-normalized approach. That is a genuine gap: existing work by Dette et al. and Van Delft/Dette assumes fully observed trajectories. The construction is sensible, the nuisance-free property is appealing, and the discussion of alternative self-normalizers, including the range-adjusted version, is a useful addition.\n\nUnfortunately, the central asymptotic claim has a load-bearing problem. Theorem 2.1 states that at the boundary ∫m²=Δ the rejection probability goes to α under (A1)-(A6). The stress-test note is correct: the B-spline bias contributes a deterministic term of order √n J_n^{-q*} to the numerator, while the stochastic scale is roughly sqrt{1+J_n E(N^{-1})}. Assumption (A6)'s first line only requires that ratio to be O(1), not o(1). So with J_n = c n^{1/(2q*)} (e.g., cubic splines, J_n = c n^{1/8}), the limiting boundary rejection probability becomes P(W(1)+b > quantile), which is not α for generic m. Condition (*) prevents degeneracy of the self-normalizer but does not control bias. The theorem as stated is therefore not justified.\n\nThe paper also contradicts itself: the abstract claims the detection rate remains n^{-1/2} even in sparse regimes, while Corollary 2.1 gives √(J_n E(N^{-1})/n) in the sparse case, which is slower than n^{-1/2}. That needs a clear fix.\n\nTwo smaller issues. All proofs are in a supplement not included for verification, so the central Gaussian approximation lemma is uncheckable here. And the BIC-based knot selection in simulations is recommended without proving that the chosen J_n satisfies (A6) or the additional undersmoothing that would repair the boundary problem.\n\nNone of this kills the underlying idea. The framework is plausible and the simulations are supportive, but only under undersmoothing that is not part of the stated assumptions. A revision strengthening A6 to require vanishing bias, reconciling the detection-rate statements, and making the key lemmas accessible would turn this into a solid paper.\n\nI would send it to a serious referee with instructions to focus on the boundary behavior and the A6 conditions. I would not cite it in its current form.","headline":"A valuable unified framework that overreaches in its key boundary claim; the B-spline bias under (A6) breaks the stated size at the boundary, and the abstract contradicts the paper's own detection-rate corollary.","tokens_in":23375,"tokens_out":2988,"would_cite":false,"duration_ms":26786,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62M10","62G20","62R10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single self-normalized B-spline decision rule tests relevant hypotheses in functional time series under arbitrary sampling, with fixed Brownian-motion critical values and no nuisance-parameter estimation.","keywords":["relevant hypothesis testing","functional time series","self-normalization","B-spline estimation","phase transition","change point","Gaussian approximation","long-run covariance"],"falsifier":"Simulate a sparse functional time series at the boundary $\\int m^2 = \\Delta$ using a functional AR(1) with innovations near the edge of the assumed moment and dependence conditions, and check whether the empirical rejection probability approaches $\\alpha$ as $n$ grows; any systematic drift away from $\\alpha$, or a mismatch between the empirical quantiles of $(T_n-\\Delta)/V_{n,\\epsilon}$ and the Brownian-motion ratio quantile $Q_{1-\\alpha,\\epsilon}$, would indicate that the Gaussian approximation in Lemma S.4.1 fails.","tokens_in":22247,"feed_emoji":"📈","tokens_out":12124,"duration_ms":103686,"temperature":0.7,"pith_summary":"The paper aims to establish a unified theory for testing \"relevant\" hypotheses in functional time series—questions of whether a functional effect exceeds a pre-specified tolerance $\\Delta$, rather than merely whether it is nonzero—for one-sample, two-sample, and change-point problems, when curves are observed sparsely or densely with measurement error. It argues that a single recipe, B-spline estimation of partial-sample mean curves combined with self-normalization, produces test statistics whose limiting distributions are pivotal and independent of sampling frequency, so no estimation of long-run covariance or noise-variance functions is needed. The central theoretical result, Theorem 2.1, gives the asymptotic rejection probability as $0$, $\\alpha$, and $1$ depending on whether the squared $L^2$ norm of the mean function lies below, at, or above the tolerance $\\Delta$; a companion result shows that departures of order $n^{-1/2}$ from the tolerance are detectable at any sampling frequency. If correct, this gives practitioners one decision rule with fixed critical values across sparse and dense designs, along with a precise phase-transition boundary between the two regimes.","feed_headline":"One test rule detects relevant functional effects, sparse or dense","feed_subtitle":"No long-run covariance or noise variance needs estimation, and critical values are fixed.","key_machinery":"The machinery is the combination of a B-spline estimator computed on partial samples and a self-normalizer built from those same partial fits. With $\\hat m(t,\\cdot)$ denoting the mean estimated from the first $[nt]$ observations, the test rejects when $(T_n-\\Delta)/V_{n,\\epsilon}$ exceeds $Q_{1-\\alpha,\\epsilon}$, the quantile of the pivotal ratio $W(1)/[\\int_\\epsilon^1 t^2\\{W(t)-tW(1)\\}^2 dt]^{1/2}$ with $W$ a standard Brownian motion. The estimator and normalizer live on the same scale, so the nuisance factor $\\tau_n(m)$—the analogue of the long-run variance—cancels in the ratio. The technical engine that produces the joint weak convergence in (2.4) is a sequential Gaussian approximation for dependent random vectors of moderately high dimension (Lemma S.4.1 in the supplement), which turns the B-spline coefficient process into a Gaussian process whose self-normalized version is pivotal; the non-degeneracy condition $(*)$ ensures $\\tau_n^2(m)$ dominates the approximation error so the denominator does not degenerate.","core_discovery":"On the paper's own terms, the discovery is that relevant-hypothesis testing in functional time series need not be rebuilt for each problem or each sampling regime. For model (1.1)—discretely observed trajectories with mean $m$, stationary functional noise $\\xi_i$, and measurement error $\\sigma\\varepsilon$—the B-spline estimator $\\hat m(\\cdot)$ and the self-normalizer $V_{n,\\epsilon}$ built from partial-sample fits give a rejection rule (2.3) whose asymptotic rejection probability is $0$ when $\\int m^2 < \\Delta$, $\\alpha$ when $\\int m^2 = \\Delta$ under the non-degeneracy condition $(*)$, and $1$ when $\\int m^2 > \\Delta$ (Theorem 2.1). The same pivotal structure is proved for two-sample comparisons (Theorem 2.3), single relevant change points (Theorem 3.1), and multiple change points (Theorem 3.3), provided the change-point locations are estimated consistently. Under local alternatives $\\|m\\|_{L^2}^2 = \\Delta + c_0\\sqrt{(1+J_n E(N^{-1}))/n}$, the test retains nontrivial power (Theorem 2.2), so the detection rate is of order $n^{-1/2}$ in dense designs and $\\sqrt{J_n E(N^{-1})/n}$ in sparse designs, with the phase transition occurring where $\\{E(N^{-1})\\}^{-1}$ crosses $(n\\log n)^{1/(2q^*)}$.","pith_inferences":["Going beyond the paper: because the pivotal ratio depends only on a Brownian motion, the same self-normalized construction may extend to other functionals—covariance operators or eigen-structures—whenever a B-spline partial-sum representation and a sequential Gaussian approximation are available; the paper does not make this claim.","A practical reading of Theorem 2.2 that the paper leaves implicit is that in sparse designs, adding more observations per subject does not improve the detection rate; only increasing the number of subjects $n$ does, so sampling effort is best spent on subjects, not grid density, once the phase-transition boundary is passed.","The paper flags stationarity and low dimensionality as limitations; a natural extension, not pursued here, would adapt the Gaussian-approximation step to locally stationary functional time series, making the same recipe applicable to nonstationary panels.","Condition $(*)$ is a theoretical non-degeneracy condition, and the paper gives no data-driven check; an empirical researcher could estimate $\\tau_n^2(m)$ on the boundary case and verify that it dominates $1+J_n E(N^{-1})$ before trusting the nominal level."],"forward_implications":["For applied work, testing whether a mean difference is at least $\\Delta$ on discretely observed curves needs no separate procedure per sampling density: the same B-spline fit, self-normalizer, and critical value $Q_{1-\\alpha,\\epsilon}$ are used whether each subject is observed a few times or many times.","Because detection of $n^{-1/2}$-local alternatives holds at arbitrary sampling frequencies, the sparse-to-dense phase transition is fully characterized: the detection rate is $\\sqrt{J_n E(N^{-1})/n}$ in the sparse regime and $\\sqrt{1/n}$ in the dense regime, with the boundary at $\\{E(N^{-1})\\}^{-1}\\asymp(n\\log n)^{1/(2q^*)}$.","The same pivotal limiting distribution covers one-sample, two-sample, single and multiple change-point relevant tests, so a user substitutes the appropriate estimated difference and self-normalizer while keeping the critical value fixed.","At the exact boundary $\\|m\\|^2_{L^2}=\\Delta$, the test has asymptotic size exactly $\\alpha$ whenever the non-degeneracy condition $(*)$ (or its two-sample and change-point analogues) holds; without it, the boundary case may be mis-sized."],"supporting_citations":[{"why":"Supplies the self-normalized relevant-hypothesis testing template for fully observed functional time series that this paper extends to sparse and dense discretely observed data.","marker":"Dette et al. [2020b]"},{"why":"Introduces self-normalization for time series, the technique the paper uses to avoid estimating long-run covariance and noise variance.","marker":"Shao [2010]"},{"why":"Establishes the sparse/semi-dense/dense phase-transition boundaries for mean and covariance estimation that the paper re-derives for relevant inference.","marker":"Zhang and Wang [2016]"},{"why":"Provides the consistent change-point estimator and phase-transition analysis whose rates underlie Assumption (B2) and Theorem 3.1.","marker":"Cai and Hu [2024]"},{"why":"Supplies the physical dependence measure framework for functional time series used in Assumptions (A4) and (A5).","marker":"Zhou and Dette [2023]"},{"why":"Provides the consistent multiple change-point estimators required by Assumption (B2') for the multiple change-point extension.","marker":"Madrid Padilla et al. [2022]"},{"why":"Motivates the range-adjusted self-normalizer that the paper discusses as a power-enhancing alternative.","marker":"Hong et al. [2024]"}],"fun_headline_variants":["Unified tests for functional time series: sparse or dense","One framework tests relevant effects, sparse or dense","Self-normalized tests detect n^{-1/2} alternatives in sparse data","Sparse or dense: one test rule for relevant functional effects","Self-normalization eliminates covariance estimation in functional tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the sequential Gaussian approximation for dependent random vectors of moderately high dimension (Lemma S.4.1 in the supplementary material) holds under the paper's dependence and moment conditions; this lemma generates the pivotal joint convergence in (2.4), and it is stated only in the supplement, not proved in the main text.","fun_headline_variants_meta":{"raw":{"variants":["Unified tests for functional time series: sparse or dense","One framework tests relevant effects, sparse or dense","Self-normalized tests detect n^{-1/2} alternatives in sparse data","Sparse or dense: one test rule for relevant functional effects","Self-normalization eliminates covariance estimation in functional tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000774,"raw_usage":{"total_tokens":3497,"prompt_tokens":1091,"completion_tokens":2406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":2324}},"tokens_in":707,"tokens_out":2406,"duration_ms":17286,"temperature":1.0,"reasoning_tokens":2324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:55:01.780844+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a sparse functional time series at the boundary $\\int m^2 = \\Delta$ using a functional AR(1) with innovations near the edge of the assumed moment and dependence conditions, and check whether the empirical rejection probability approaches $\\alpha$ as $n$ grows; any systematic drift away from $\\alpha$, or a mismatch between the empirical quantiles of $(T_n-\\Delta)/V_{n,\\epsilon}$ and the Brownian-motion ratio quantile $Q_{1-\\alpha,\\epsilon}$, would indicate that the Gaussian approximation in Lemma S.4.1 fails.","supporting_citations":[],"review_version":2}