Pith. sign in

REVIEW 3 major objections 3 minor 5 references

A local cross-validated likelihood term is shown to converge to the asymptotic mean squared error of the zero-frequency spectral estimator, supporting automatic HAC bandwidth selection.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 20:39 UTC pith:LS23Y365

load-bearing objection A correct-looking local-to-zero theorem for the CVLL_c quadratic term, but the bandwidth-selection justification depends on an unproved negligibility conjecture that the authors themselves flag. the 3 major comments →

arxiv 2607.16535 v1 pith:LS23Y365 submitted 2026-07-17 stat.ME

On the Expectation of the Local-to-Zero Cross-Validated Log Likelihood Criterion for Bandwidth Selection in Kernel Spectral Estimation

classification stat.ME MSC 62M1562M1062G05
keywords kernel spectral estimationbandwidth selectioncross-validated log-likelihoodHAC standard errorslong-run variancezero frequencyasymptotic mean squared errorlocal-to-zero
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes a precise asymptotic connection between a 'local-to-zero' version of the cross-validated log-likelihood (CVLL_c) and the mean squared error (MSE) of a kernel spectral estimator at frequency zero. Specifically, it proves that for a shrinking set of Fourier frequencies (from 1 to n^c with 4/5 < c < 1), the scaled expectation of the quadratic term in the Taylor expansion of CVLL_c tends to the classical variance-plus-squared-bias formula c0(τ) = 1/2[τκ + τ^{-4}(h^{(2)} f^{(2)}(0)/f(0))^2]. Since HAC standard errors depend on an accurate estimate of f(0), this gives theoretical support for using CVLL_c to choose bandwidth automatically, though the full justification rests on an unproved conjecture that other terms in the expansion are negligible. The proof overcomes a technical obstacle: the usual Parseval identity that works for global CVLL fails for local sums, so the paper develops a Dirichlet-kernel bound that controls the local frequency sum.

Core claim

At the heart of the paper is Theorem 0.1: when the bandwidth is m = τ n^{1/5} and the number of local frequencies is n^c with 4/5 < c < 1, the scaled expectation (1/2)n^{4/5-c} Σ_{j=1}^{n^c} E[(fhat_j^{(j)}/f_j - 1)^2] converges to c0(τ). Here fhat_j^{(j)} is the leave-one-out kernel spectral estimate at Fourier frequency ω_j, f_j is the true spectrum, and c0(τ) is exactly the asymptotic MSE of the discrete periodogram average estimate of f(0): a term τκ for the variance and a term τ^{-4}(h^{(2)} f^{(2)}(0)/f(0))^2 for the squared bias. The theorem focuses on the quadratic term of the Taylor expansion of CVLL_c(τ) around the true spectrum; the paper shows that this term carries the τ-depende

What carries the argument

The central object is the Taylor expansion of the local cross-validated log-likelihood difference L~(τ) − L~, whose quadratic term (0.18) is a sum of squared relative errors of the leave-one-out estimates. The proof of Theorem 0.1 decomposes each relative error fhat_j^{(j)}/f_j − 1 into a negligible high-order term, an innovation-noise term g*_j (a centered periodogram-based spectral estimate of the innovations), and a bias term E[f~_j]/f_j − 1. Lemma 0.5 shows that the variance part converges to τκ by bounding the sum of Dirichlet kernels that appear because the frequency sum is local; this bound substitutes for the Parseval identity used in global CVLL proofs. Lemma 0.2 shows that the bias

Load-bearing premise

The load-bearing premise is that the expectations of the linear term (0.16) and the cross term (0.17) in the Taylor expansion of CVLL_c are negligible for general stationary processes; this is proven only for Gaussian white noise, and the authors state they believe it holds more generally but do not provide a proof.

What would settle it

A simulation or exact asymptotic computation for a non-Gaussian linear process (for example, an ARMA model with Student-t innovations) that evaluates E[CVLL_c(τ) − L~] and compares it with c0(τ) over τ would decide the matter. If the difference between the two does not vanish relative to n^{4/5−c} for c in (4/5, 1), then the unproved terms are not negligible and the bandwidth selector would be biased. One could also directly compute the expectations of (0.16) and (0.17) at low frequencies; if they are not o(n^{c−4/5}), the theorem's premise fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the authors' conjecture about the negligible linear and cross terms holds, then E[CVLL_c(τ) − L~] equals c0(τ) up to a constant, so the minimizer of CVLL_c(τ) asymptotically matches the minimizer of the zero-frequency ASE, giving a fully data-driven bandwidth rule for HAC standard errors.
  • The theorem pins down the convergence rate: the scaling n^{4/5−c} is exactly what is needed for the quadratic sum to have a finite limit, and the condition c > 4/5 is necessary for the bias term to survive in the limit; c < 1 keeps the frequency window local to zero.
  • Because the result holds for each fixed τ, it can be applied to compare different bandwidth constants τ in the same local frequency window, turning CVLL_c into a practical selector that does not require knowledge of f, its derivatives, or the noise variance.
  • The proof's technical core—a Dirichlet-kernel bound for sums over low frequencies—is a reusable tool for other local-to-zero model selection problems where global Parseval arguments are unavailable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A concrete next step would be to prove the negligibility of the linear and cross terms using mixing or cumulant conditions; if those terms have non-negligible expected limits, c0(τ) would gain an additive shift that changes the optimal bandwidth constant τ*, a possibility the paper does not quantify.
  • The Dirichlet-kernel technique suggests that other global criteria, such as AIC or FPE, might also have local-to-zero versions whose expectations match the MSE of the spectral density at a single frequency; one could test this by deriving analogous expansions for those criteria.
  • The paper's assumption of a finite fourth innovation moment and the appearance of the kurtosis term E[ε^4]−1 in the variance lemma (though it vanishes asymptotically) hints that finite-sample CVLL_c could be sensitive to heavy tails; a simulation study of an AR(1) with t-distributed innovations would reveal whether the bandwidth selection deteriorates in moderate samples.
  • Since the theorem requires only that f be twice continuously differentiable near zero and that the underlying process be short-memory with finite fourth moment, the result may carry over to fractionally integrated processes at frequency zero, though the rates would need re-derivation because the spectral density may be unbounded.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper studies the local-to-zero cross-validated log-likelihood criterion CVLL_c(τ) for bandwidth selection in kernel spectral estimation at frequency zero. The authors define CVLL_c(τ) as a leave-one-out average over Fourier frequencies j=1,...,n^c with 0<c<1, and consider a Taylor expansion of the criterion difference L~(τ)-L~. Their main theorem (Theorem 0.1) states that, for m=τ n^{1/5}, ñ=n^c, and 4/5<c<1, the scaled expectation of the quadratic term in the expansion, (1/2)n^{4/5-c} ∑ E( fhat_j^j/f_j - 1)^2, converges to the asymptotic mean squared percentage error c0(τ) of the zero-frequency spectral estimator, where c0(τ)=1/2[τκ+τ^{-4}(h^{(2)}f^{(2)}(0)/f(0))^2]. The proof decomposes the leave-one-out error and proves two lemmas: one on the squared bias of the lag-window estimate (Lemma 0.2) and one on the variance contribution of the innovation periodogram (Lemma 0.5). The paper explicitly acknowledges that the expectations of two other Taylor terms, (0.16) and (0.17), are only shown to vanish in the Gaussian white-noise case, and conjectures but does not prove their negligibility for general linear processes.

Significance. If the central claim were fully established, the paper would provide an important bridge between the local CVLL criterion and the classical asymptotic MSE of the Parzen spectral window estimator, thereby giving a formal justification for HAC bandwidth selection via CVLL_c. The paper correctly identifies that existing global CVLL results cannot be transplanted to the local-to-zero case because Parseval's formula fails for local sums. The authors are also commendably transparent about the unproved negligibility of the (0.16) and (0.17) terms. However, as presented, the result justifies only the quadratic term in the Taylor expansion, not the expectation of the full CVLL_c difference; the missing control of (0.16) and (0.17) is load-bearing for the advertised conclusion. The proof of Lemma 0.5 also contains an incomplete step. Thus the paper is of moderate significance pending a fix of these gaps.

major comments (3)
  1. [After Eq. (0.19)] The abstract claims that the results provide 'some justification' for CVLL_c-based HAC bandwidth selection. But the paper explicitly states after Eq. (0.19) that the expectations of (0.16) and (0.17) are only proved negligible in the extremely special Gaussian white-noise case, and that 'we believe that it could be proved... but we do not pursue this here.' For general non-Gaussian linear processes, leave-one-out estimates are not independent of I_j, so E[(fhat_j^j/f_j-1)(I_j/f_j-1)] and E[(I_j/f_j-1)(fhat_j^j/f_j-1)^2] need not vanish. If these terms have the same order as the quadratic term after n^{4/5-c} scaling, then E[CVLL_c(τ)-L~] would be c0(τ) plus additional unknown terms, so Theorem 0.1 alone does not justify CVLL_c-based bandwidth selection. This is the central load-bearing gap and should be addressed, either by proving the negligibility under explicit regularity conditions o
  2. [Lemma 0.5, around Eq. (0.75)] The proof of Lemma 0.5 contains a missing line: 'and . Thus' at Eq. (0.75) obscures the simplification of E[γ^ϵ_r γ^ϵ_s]. In particular, the derivation leading to Eq. (0.77) requires a careful accounting of the nonzero fourth-moment terms when |r|=|s|≠0 and when r or s equals zero. The sentence is broken and the step is not verifiable as written. Since Lemma 0.5 is essential for the variance contribution in Theorem 0.1, this missing line and the underlying algebra should be completed and checked.
  3. [Eq. (0.36) and surrounding notation] The expression in Eq. (0.36) is not written consistently: it should be τ^{-4}(h^{(2)} f^{(2)}(ω_j)/f(ω_j))^2, comparing with the definition of b0(τ) in Eq. (0.29). Also, Theorem 0.1 uses c0(τ) while Eq. (0.11) writes 'co(τ)'. These are presentation errors, but they make it difficult to verify the proof's algebra and should be fixed.
minor comments (3)
  1. [References] Reference [1] has typos ('Bandwith', 'Specturm'), and Lemma 0.3 cites 'Chen and Hurvich (1998)' while the reference list gives 'Chen, W. & Hurvich, C. (2000)'. Please standardize.
  2. [Abstract] The abstract appears twice in the manuscript text (once at the top and again after the arXiv line). This duplication should be removed.
  3. [Eq. (0.15)] The Op term in Eq. (0.15) is written as Op( n^~ (fhat_j^j/f_j-1)^3 ) with the tilde over n in a nonstandard way; please clarify the notation and the order of the remainder term.

Circularity Check

0 steps flagged

No circularity: Theorem 0.1 is derived rather than assumed; the unproved negligibility of terms (0.16)/(0.17) is a stated gap, not a circular reduction.

full rationale

The central Theorem 0.1 is not circular. It proves that a normalized local average of squared leave-one-out relative estimation errors converges to c0(τ), which is Parzen's known zero-frequency asymptotic MSE. Those are distinct objects: Parzen's formula describes the full spectral estimator, while Theorem 0.1 concerns the leave-one-out estimates summed over frequencies 1,...,ñ. The proof derives, rather than assumes, both the bias and variance components: Lemma 0.2 computes the squared-bias limit by expanding E[tilde f(ω_j)] - f(ω_j), and Lemma 0.5 computes E[g_j*^2] explicitly. τ is a free variable, not a fitted parameter. The self-citations (Chen-Hurvich for a Dirichlet kernel bound, Xu-Hurvich for the CVLL_c proposal) are auxiliary: the Dirichlet bound is an elementary, parameter-free inequality used to control a local average, and the Xu-Hurvich citation supplies context, not the theorem. The manuscript transparently identifies the main limitation: after Eq. (0.19) it states 'We believe that it could be proved that the expectations of (0.16) and (0.17) are negligible more generally under suitable regularity conditions, but we do not pursue this here.' This is an acknowledged unproved assumption about the rest of the Taylor expansion, not a circular step: the paper does not claim that this negligibility follows from the theorem or from the fitted values. Thus, while the full justification of CVLL_c-based HAC bandwidth selection remains incomplete outside Gaussian white noise, the derivation chain in Theorem 0.1 is self-contained and non-circular.

Axiom & Free-Parameter Ledger

2 free parameters · 6 axioms · 0 invented entities

The paper introduces no new entities. It relies on standard time-series assumptions (linear process, finite fourth moment, smoothness at zero), a specific kernel choice, and a cited technical bound from Robinson (1991). The two user-chosen constants tau and c are not fitted; they are arguments of the asymptotic criterion.

free parameters (2)
  • tau
    Bandwidth scale in m = tau n^{1/5}; it is the argument of the asymptotic AMSE c0(tau), not fitted to data. The theorem holds uniformly for tau in compact intervals.
  • c
    Number of Fourier frequencies used, n^c; restricted to 4/5 < c < 1. A modeling choice from the CVLL_c definition, not estimated.
axioms (6)
  • domain assumption X_t is a linear process with iid innovations having finite fourth moment and sum_j j^{1/2}|beta_j| < infinity (Eq 0.1).
    Defines the class of processes studied; used throughout.
  • domain assumption Spectral density f is twice continuously differentiable at zero, f(0)>0, f^{(2)}(0) finite.
    Used for the bias expansion and continuity arguments in Lemma 0.2.
  • domain assumption Autocovariances satisfy sum_r r^2 |gamma_r| < infinity.
    Needed for f^{(2)} and dominated convergence steps.
  • domain assumption Parzen lag window with k(0)=1, |k|<=1, finite support, h^{(2)}=6.
    Specific kernel choice; the theorem may be kernel-specific.
  • standard math integral of |x| k^2(x) over R is finite.
    Used in Lemma 0.5 for the vanishing term (0.80); holds for Parzen kernel.
  • domain assumption Robinson (1991) bounds (C.7)-(C.9), including max_{tau,j}|fhat_j^j - f_j| = O_p(n^{-2/5}).
    Used to discard the first two terms in (0.22) and the O_p remainder.

pith-pipeline@v1.3.0-alltime-deepseek · 13127 in / 12947 out tokens · 130769 ms · 2026-08-01T20:39:42.045638+00:00 · methodology

0 comments
read the original abstract

We consider data-driven bandwidth selection for a kernel spectral estimator at zero frequency based on a local-to-zero version of the cross validated log-likelihood (CVLL) criterion. The modified version is $\mbox{CVLL}_c$, based on a sum over Fourier frequencies from $1$ to $n^c$ with $0<c<1$, where $n$ is the sample size. We focus on the expectation of a key term in a Taylor series expansion for $\mbox{CVLL}_c$ and show that in the case $4/5 < c < 1$ it converges to the corresponding asymptotic mean squared error of the spectral estimator at zero frequency. This provides some justification for the use of the local CVLL criterion for Heteroskedasticity and Autocorrelation Consistent (HAC) standard error estimation. Our theoretical results do not follow from existing literature on CVLL because those results exploit the fact that CVLL is global, summing over all frequencies in $(0,\pi)$ rather than local-to-zero frequency, as is the case for $\mbox{CVLL}_c$.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

5 extracted references

  1. [1]

    Beltr˜ ao, K. I. & Bloomfield, P. (1987) Determing the Bandwith of A Kernel Specturm Estimate.Journal of Time Series Analysis8(1), 21-38

  2. [2]

    & Hurvich, C

    Chen, W. & Hurvich, C. (2000) An Efficient Taper for Potentially Overdifferenced Long- memory Time Series.Journal of Time Series Analysis8(1), 21-38

  3. [3]

    Robinson, P. M. (1991) Automatic Frequency Domain Inference on Semiparametric and Nonparametric Models.Econometrica59(5), 1329-1363. 17

  4. [4]

    (1957) On Consistent Estimates of the Spectrum of a Stationary Time Series

    Parzen, E. (1957) On Consistent Estimates of the Spectrum of a Stationary Time Series. The Annals of Mathematical Statistics21(2), 155-180

  5. [5]

    & Hurvich, C

    Xu, Z. & Hurvich, C. (2026) A Unified Frequency Domain Cross-Validatory Approach to HAC Standard Error Estimation.Econometrics and Statistics37, 214-229. 18