GACL adds domain grounding via alternating reference and synthetic task sampling to a VAE-based regret-driven curriculum teacher, reporting higher success rates than CLUTR on BARN navigation and quadruped locomotion in simulation.
Automatic Curriculum Learning with Gradient Reward Signals
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper investigates the impact of using gradient norm reward signals in the context of Automatic Curriculum Learning (ACL) for deep reinforcement learning (DRL). We introduce a framework where the teacher model, utilizing the gradient norm information of a student model, dynamically adapts the learning curriculum. This approach is based on the hypothesis that gradient norms can provide a nuanced and effective measure of learning progress. Our experimental setup involves several reinforcement learning environments (PointMaze, AntMaze, and AdroitHandRelocate), to assess the efficacy of our method. We analyze how gradient norm rewards influence the teacher's ability to craft challenging yet achievable learning sequences, ultimately enhancing the student's performance. Our results show that this approach not only accelerates the learning process but also leads to improved generalization and adaptability in complex tasks. The findings underscore the potential of gradient norm signals in creating more efficient and robust ACL systems, opening new avenues for research in curriculum learning and reinforcement learning.
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
GACL: Grounded Adaptive Curriculum Learning with Active Task and Performance Monitoring
GACL adds domain grounding via alternating reference and synthetic task sampling to a VAE-based regret-driven curriculum teacher, reporting higher success rates than CLUTR on BARN navigation and quadruped locomotion in simulation.