Pith. sign in

REVIEW 8 cited by

Regret of exploratory policy improvement and $q$-learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01302 v1 pith:BZMS5VXV submitted 2024-11-02 cs.LG math.OCmath.PR

classification cs.LGmath.OCmath.PR
keywords learningalgorithmexploratoryimprovementpolicyregretalgorithmsanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J. Mach. Learn. Res., 24 (2023), 161) for controlled diffusion processes. Under suitable conditions on the growth and regularity of the model parameters, we provide a quantitative error and regret analysis of both the exploratory policy improvement algorithm and the $q$-learning algorithm.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

    cs.LG 2026-07 reject novelty 7.0 of 10

    A continuous-time RL framework for fine-tuning discrete diffusion models is proposed, but the key objective equivalence in the paper is flawed.

  2. ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    cs.LG 2026-07 unverdicted novelty 7.0 of 10

    ART-RL learns adaptive diffusion sampling timesteps via continuous-time control and Gaussian actor–critic RL, improving and transferring over hand-designed grids at matched budgets.

  3. Upper and lower bounds for local Lipschitz stability of Bayesian posteriors

    math.ST 2025-05 reject novelty 7.0 of 10

    Lower bounds on posterior sensitivity are derived, but the advertised "sensitivity must increase with concentration" conclusion is not supported.

  4. Beyond separability: convergence rate of vanishing viscosity approximations to mean field games via FBSDE stability

    math.OC 2025-05 conditional novelty 7.0 of 10

    The vanishing viscosity approximation to nonlocal, possibly non-separable mean field games converges at rate O(β) in L∞ on compact sets, matching the classical Hamilton-Jacobi rate.

  5. Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies

    math.OC 2026-07 conditional novelty 6.0 of 10

    For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.

  6. Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

    cs.AI 2026-02 conditional novelty 6.0 of 10

    By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.

  7. Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Data-driven adaptive exploration achieves O(N^{3/4}) regret in continuous-time linear-quadratic reinforcement learning, matching fixed-schedule methods and extending them to zero initial states.

  8. Continuous-time reinforcement learning for optimal switching over multiple regimes

    math.OC 2025-12 conditional novelty 5.0 of 10

    An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.

Pith tools