Combining supervised fine-tuning with closed-loop RL lets small Qwen LLMs tune an MPC controller, and the 3B model scores 63.3% versus 58.5% for GPT-4o on the paper's custom control adaptability metric.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning
Combining supervised fine-tuning with closed-loop RL lets small Qwen LLMs tune an MPC controller, and the 3B model scores 63.3% versus 58.5% for GPT-4o on the paper's custom control adaptability metric.