An LLM-generated and LLM-selected rule set dynamically reweights reward components in a multi-branch value network, yielding modest average gains over PPO in simulated robot tasks.
Reinforcement learning for reduced-order models of legged robots,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning
An LLM-generated and LLM-selected rule set dynamically reweights reward components in a multi-branch value network, yielding modest average gains over PPO in simulated robot tasks.