Reward distillation from a 317M TD-MPC2 teacher yields a 1M-parameter MT30 agent that scores 28.12 to 28.45, narrowly above a from-scratch baseline, with FP16 halving size.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Knowledge Transfer in Model-Based Reinforcement Learning Agents for Efficient Multi-Task Learning
Reward distillation from a 317M TD-MPC2 teacher yields a 1M-parameter MT30 agent that scores 28.12 to 28.45, narrowly above a from-scratch baseline, with FP16 halving size.