The paper proposes a constraint-aware Bellman operator formed by composing the optimal Bellman operator with a proximal projection, and claims it stays contractive while enforcing convex domain constraints exactly.
Off-policy deep reinforcement learning without exploration,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.SY 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Implicit Constraint-Aware Off-Policy Correction for Offline Reinforcement Learning
The paper proposes a constraint-aware Bellman operator formed by composing the optimal Bellman operator with a proximal projection, and claims it stays contractive while enforcing convex domain constraints exactly.