An LLM-generated and LLM-selected rule set dynamically reweights reward components in a multi-branch value network, yielding modest average gains over PPO in simulated robot tasks.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning
An LLM-generated and LLM-selected rule set dynamically reweights reward components in a multi-branch value network, yielding modest average gains over PPO in simulated robot tasks.