Pith. sign in

REVIEW 2 cited by

State Augmented Constrained Reinforcement Learning: Overcoming the Limitations of Learning with Rewards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.11941 v2 pith:JE3LL2II submitted 2021-02-23 cs.LG cs.ROmath.OC

classification cs.LGcs.ROmath.OC
keywords learningoptimalreinforcementconstrainedmethodspolicyproblemsrewards
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A common formulation of constrained reinforcement learning involves multiple rewards that must individually accumulate to given thresholds. In this class of problems, we show a simple example in which the desired optimal policy cannot be induced by any weighted linear combination of rewards. Hence, there exist constrained reinforcement learning problems for which neither regularized nor classical primal-dual methods yield optimal policies. This work addresses this shortcoming by augmenting the state with Lagrange multipliers and reinterpreting primal-dual methods as the portion of the dynamics that drives the multipliers evolution. This approach provides a systematic state augmentation procedure that is guaranteed to solve reinforcement learning problems with constraints. Thus, as we illustrate by an example, while previous methods can fail at finding optimal policies, running the dual dynamics while executing the augmented policy yields an algorithm that provably samples actions from the optimal policy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Operator Splitting for Convex Constrained Markov Decision Processes

    math.OC 2024-12 conditional novelty 6.0 of 10

    OS-CMDP uses Douglas-Rachford splitting to solve convex-constrained MDPs by alternating between a quadratically regularized MDP update and a projection onto the constraint set, with convergence and infeasibility-detec...

  2. Long-Horizon Wireless Link Scheduling with State-Augmented Graph Neural Networks

    eess.SP 2026-07 conditional novelty 5.0 of 10

    A state-augmented GNN that imitates dual subgradient descent produces near-optimal, constraint-satisfying long-horizon wireless link schedules.

Pith tools