ROSKA co-evolves reward functions and policies, using Bayesian optimization to fuse the previous best policy with random weights, reporting 95.3% average normalized improvement over Eureka on six Isaac Gym tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
ROSKA co-evolves reward functions and policies, using Bayesian optimization to fuse the previous best policy with random weights, reporting 95.3% average normalized improvement over Eureka on six Isaac Gym tasks.