CAPS trains multiple policies with shared features and switches among them at test time to satisfy arbitrary cost constraints, beating prior offline safe RL baselines on 38 benchmark tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning
CAPS trains multiple policies with shared features and switches among them at test time to satisfy arbitrary cost constraints, beating prior offline safe RL baselines on 38 benchmark tasks.