OPD-V uses the gap between a zoom-in teacher and a mask teacher, the Modality-Balance Logits Margin, to pick which on-policy tokens to distill, improving MLLM visual reasoning and cutting training time.
On-policy distillation of language models: Learning from self-generated mistakes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
OPD-V uses the gap between a zoom-in teacher and a mask teacher, the Modality-Balance Logits Margin, to pick which on-policy tokens to distill, improving MLLM visual reasoning and cutting training time.