Pith. sign in

REVIEW 7 cited by

Online Statistical Inference for Nonlinear Stochastic Approximation with Markovian Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07690 v2 pith:25ARY5IQ submitted 2023-02-15 math.ST stat.MEstat.MLstat.TH

classification math.STstat.MEstat.MLstat.TH
keywords stochasticapproximationdatafunctionalinferenceboldsymbolboundcentral
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study the statistical inference of nonlinear stochastic approximation algorithms utilizing a single trajectory of Markovian data. Our methodology has practical applications in various scenarios, such as Stochastic Gradient Descent (SGD) on autoregressive data and asynchronous Q-Learning. By utilizing the standard stochastic approximation (SA) framework to estimate the target parameter, we establish a functional central limit theorem for its partial-sum process, $\boldsymbol{\phi}_T$. To further support this theory, we provide a matching semiparametric efficient lower bound and a non-asymptotic upper bound on its weak convergence, measured in the L\'evy-Prokhorov metric. This functional central limit theorem forms the basis for our inference method. By selecting any continuous scale-invariant functional $f$, the asymptotic pivotal statistic $f(\boldsymbol{\phi}_T)$ becomes accessible, allowing us to construct an asymptotically valid confidence interval. We analyze the rejection probability of a family of functionals $f_m$, indexed by $m \in \mathbb{N}$, through theoretical and numerical means. The simulation results demonstrate the validity and efficiency of our method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

    stat.ML 2026-07 accept novelty 7.0 of 10

    Quantile fixed-point estimators of return distributions attain the parametric √n rate and the semiparametric efficiency bound for fixed and diverging numbers of quantiles, with a Berry–Esseen guarantee for smooth functionals.

  2. Optimism Stabilizes Thompson Sampling for Adaptive Inference

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Optimistic modifications of Thompson sampling make each arm's pull count concentrate around a deterministic scale, yielding asymptotically valid Wald inference in K-armed Gaussian bandits with multiple optimal arms.

  3. Statistical inference for Linear Stochastic Approximation with Markovian Noise

    stat.ML 2025-05 conditional novelty 7.0 of 10

    Polyak-Ruppert averaged linear stochastic approximation with Markovian noise achieves Berry-Esseen rate O(n^{-1/4}) in Kolmogorov distance, and a multiplier subsample bootstrap achieves coverage error O(n^{-1/10}).

  4. Beyond Self-Repellent Kernels: History-Driven Target Towards Efficient Nonlinear MCMC on General Graphs

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Replacing the target μ by a history-adjusted target μ(x/μ)^{-α} in any graph MCMC sampler gives O(1/α) variance reduction at constant per-step cost, and extends to non-reversible chains.

  5. Online Policy Evaluation for MDPs with Dynamic UBSR Measures

    cs.LG 2026-07 conditional novelty 6.0 of 10

    UBSR-TD, a temporal-difference algorithm with a loss function applied to the TD error, evaluates policies under dynamic utility-based shortfall risk with linear function approximation and converges almost surely when ...

  6. Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization

    stat.ML 2026-04 unverdicted novelty 6.0 of 10

    Q-learning with PD2Z/LD2Z step sizes admits sharp non-asymptotic bounds, a tail Polyak–Ruppert CLT, and a time-uniform Gaussian approximation, establishing a best-of-both-worlds rate-and-bias tradeoff.

  7. Online Statistical Inference of Constrained Stochastic Optimization via Random Scaling

    stat.ML 2025-05 conditional novelty 6.0 of 10

    A random scaling statistic based on averaged AI-SSQP iterates is asymptotically pivotal for constrained stochastic optimization, enabling matrix-free online confidence intervals.

Pith tools