Pith. sign in

REVIEW 1 cited by

Quasi-optimal Reinforcement Learning with Continuous Actions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.08940 v2 pith:X47VNZJ6 submitted 2023-01-21 stat.ML cs.LGstat.ME

classification stat.MLcs.LGstat.ME
keywords algorithmlearningactionsapplicationscontinuousconvergencedosemedical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many real-world applications of reinforcement learning (RL) require making decisions in continuous action environments. In particular, determining the optimal dose level plays a vital role in developing medical treatment regimes. One challenge in adapting existing RL algorithms to medical applications, however, is that the popular infinite support stochastic policies, e.g., Gaussian policy, may assign riskily high dosages and harm patients seriously. Hence, it is important to induce a policy class whose support only contains near-optimal actions, and shrink the action-searching area for effectiveness and reliability. To achieve this, we develop a novel \emph{quasi-optimal learning algorithm}, which can be easily optimized in off-policy settings with guaranteed convergence under general function approximations. Theoretically, we analyze the consistency, sample complexity, adaptability, and convergence of the proposed algorithm. We evaluate our algorithm with comprehensive simulated experiments and a dose suggestion real application to Ohio Type 1 diabetes dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where to Intervene: Action Selection in Deep Reinforcement Learning

    stat.ML 2025-07 conditional novelty 5.0 of 10

    Knockoff sampling selects the minimal sufficient action set during online deep reinforcement learning with false discovery rate control.

Pith tools