Symmetric behavior regularization for offline RL becomes tractable by expanding any f-divergence into a truncated Pearson-Vajda series, yielding a closed-form policy and bounded approximation error.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Symmetric Behavior Regularized Policy Optimization
Symmetric behavior regularization for offline RL becomes tractable by expanding any f-divergence into a truncated Pearson-Vajda series, yielding a closed-form policy and bounded approximation error.