OPOD trains one omni-modal model by routing each generated response to a matching modality teacher, and it reports the best twelve-benchmark average at 3B, 7B, and 30B scale.
On-policy distillation of language models: Learning from self-generated mistakes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
OPOD: On-Policy Omni Distillation
OPOD trains one omni-modal model by routing each generated response to a matching modality teacher, and it reports the best twelve-benchmark average at 3B, 7B, and 30B scale.