REVIEW 8 cited by
Regret of exploratory policy improvement and $q$-learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J. Mach. Learn. Res., 24 (2023), 161) for controlled diffusion processes. Under suitable conditions on the growth and regularity of the model parameters, we provide a quantitative error and regret analysis of both the exploratory policy improvement algorithm and the $q$-learning algorithm.
Forward citations
Cited by 8 Pith papers
-
A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models
A continuous-time RL framework for fine-tuning discrete diffusion models is proposed, but the key objective equivalence in the paper is flawed.
-
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
ART-RL learns adaptive diffusion sampling timesteps via continuous-time control and Gaussian actor–critic RL, improving and transferring over hand-designed grids at matched budgets.
-
Upper and lower bounds for local Lipschitz stability of Bayesian posteriors
Lower bounds on posterior sensitivity are derived, but the advertised "sensitivity must increase with concentration" conclusion is not supported.
-
Beyond separability: convergence rate of vanishing viscosity approximations to mean field games via FBSDE stability
The vanishing viscosity approximation to nonlocal, possibly non-separable mean field games converges at rate O(β) in L∞ on compact sets, matching the classical Hamilton-Jacobi rate.
-
Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.
-
Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach
By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.
-
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
Data-driven adaptive exploration achieves O(N^{3/4}) regret in continuous-time linear-quadratic reinforcement learning, matching fixed-schedule methods and extending them to zero initial states.
-
Continuous-time reinforcement learning for optimal switching over multiple regimes
An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.
Discussion (0). Sign in to comment.