Pith. sign in

On-policy distillation of language models: Learning from self-generated mistakes

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

method 1

citation-polarity summary

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

roles

method 1

polarities

use method 1

representative citing papers

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

cs.CV · 2026-08-05 · conditional · novelty 6.0

OPD-V uses the gap between a zoom-in teacher and a mask teacher, the Modality-Balance Logits Margin, to pick which on-policy tokens to distill, improving MLLM visual reasoning and cutting training time.

citing papers explorer

Showing 1 of 1 citing paper.

  • OPD-V: Visual On-Policy Self-Distillation with Modality Balance cs.CV · 2026-08-05 · conditional · none · ref 1

    OPD-V uses the gap between a zoom-in teacher and a mask teacher, the Modality-Balance Logits Margin, to pick which on-policy tokens to distill, improving MLLM visual reasoning and cutting training time.