REVIEW 7 cited by
Online Statistical Inference for Nonlinear Stochastic Approximation with Markovian Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We study the statistical inference of nonlinear stochastic approximation algorithms utilizing a single trajectory of Markovian data. Our methodology has practical applications in various scenarios, such as Stochastic Gradient Descent (SGD) on autoregressive data and asynchronous Q-Learning. By utilizing the standard stochastic approximation (SA) framework to estimate the target parameter, we establish a functional central limit theorem for its partial-sum process, $\boldsymbol{\phi}_T$. To further support this theory, we provide a matching semiparametric efficient lower bound and a non-asymptotic upper bound on its weak convergence, measured in the L\'evy-Prokhorov metric. This functional central limit theorem forms the basis for our inference method. By selecting any continuous scale-invariant functional $f$, the asymptotic pivotal statistic $f(\boldsymbol{\phi}_T)$ becomes accessible, allowing us to construct an asymptotically valid confidence interval. We analyze the rejection probability of a family of functionals $f_m$, indexed by $m \in \mathbb{N}$, through theoretical and numerical means. The simulation results demonstrate the validity and efficiency of our method.
Forward citations
Cited by 7 Pith papers
-
Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning
Quantile fixed-point estimators of return distributions attain the parametric √n rate and the semiparametric efficiency bound for fixed and diverging numbers of quantiles, with a Berry–Esseen guarantee for smooth functionals.
-
Optimism Stabilizes Thompson Sampling for Adaptive Inference
Optimistic modifications of Thompson sampling make each arm's pull count concentrate around a deterministic scale, yielding asymptotically valid Wald inference in K-armed Gaussian bandits with multiple optimal arms.
-
Statistical inference for Linear Stochastic Approximation with Markovian Noise
Polyak-Ruppert averaged linear stochastic approximation with Markovian noise achieves Berry-Esseen rate O(n^{-1/4}) in Kolmogorov distance, and a multiplier subsample bootstrap achieves coverage error O(n^{-1/10}).
-
Beyond Self-Repellent Kernels: History-Driven Target Towards Efficient Nonlinear MCMC on General Graphs
Replacing the target μ by a history-adjusted target μ(x/μ)^{-α} in any graph MCMC sampler gives O(1/α) variance reduction at constant per-step cost, and extends to non-reversible chains.
-
Online Policy Evaluation for MDPs with Dynamic UBSR Measures
UBSR-TD, a temporal-difference algorithm with a loss function applied to the TD error, evaluates policies under dynamic utility-based shortfall risk with linear function approximation and converges almost surely when ...
-
Sharp asymptotic theory for Q-learning with LDTZ learning rate and its generalization
Q-learning with PD2Z/LD2Z step sizes admits sharp non-asymptotic bounds, a tail Polyak–Ruppert CLT, and a time-uniform Gaussian approximation, establishing a best-of-both-worlds rate-and-bias tradeoff.
-
Online Statistical Inference of Constrained Stochastic Optimization via Random Scaling
A random scaling statistic based on averaged AI-SSQP iterates is asymptotically pivotal for constrained stochastic optimization, enabling matrix-free online confidence intervals.
Discussion (0). Continue with ORCID to comment.