Bidirectional SAC combines an explicit forward-KL policy projection with reverse-KL policy refinement and reports up to 30% higher episodic rewards on MuJoCo and Box2D continuous control tasks.
Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
Bidirectional SAC combines an explicit forward-KL policy projection with reverse-KL policy refinement and reports up to 30% higher episodic rewards on MuJoCo and Box2D continuous control tasks.