Pith. sign in

REVIEW 2 cited by

CAQL: Continuous Action Q-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12397 v3 pith:MOYWHB2H submitted 2019-09-26 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords caqlq-learningactionmax-qmethodsapproximatecontinuouscontinuous-action
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Value-based reinforcement learning (RL) methods like Q-learning have shown success in a variety of domains. One challenge in applying Q-learning to continuous-action RL problems, however, is the continuous action maximization (max-Q) required for optimal Bellman backup. In this work, we develop CAQL, a (class of) algorithm(s) for continuous-action Q-learning that can use several plug-and-play optimizers for the max-Q problem. Leveraging recent optimization results for deep neural networks, we show that max-Q can be solved optimally using mixed-integer programming (MIP). When the Q-function representation has sufficient power, MIP-based optimization gives rise to better policies and is more robust than approximate methods (e.g., gradient ascent, cross-entropy search). We further develop several techniques to accelerate inference in CAQL, which despite their approximate nature, perform well. We compare CAQL with state-of-the-art RL algorithms on benchmark continuous-control problems that have different degrees of action constraints and show that CAQL outperforms policy-based methods in heavily constrained environments, often dramatically.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Combinatorial Reinforcement Learning with Preference Feedback

    stat.ML 2025-02 conditional novelty 7.0 of 10

    MNL-VQL is the first algorithm with regret bounds for combinatorial reinforcement learning with multinomial-logit preference feedback, and it is nearly minimax-optimal in linear MDPs.

  2. Conformal Mixed-Integer Constraint Learning with Feasibility Guarantees

    cs.LG 2025-06 reject novelty 6.0 of 10

    C-MICL embeds conformal prediction sets into mixed-integer constraint learning, claiming a 1-alpha probability that optimized solutions are feasible for the true unknown constraint.

Pith tools