Showing that an LVLM teacher's soft action probabilities, blended into the RL loss, speed up MiniGrid agents by about 2.5x in sample efficiency.
Mastering atari, go, chess and shogi by planning with a learned model,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
Showing that an LVLM teacher's soft action probabilities, blended into the RL loss, speed up MiniGrid agents by about 2.5x in sample efficiency.