Pith. sign in

On-policy distillation of language models: Learning from self-generated mistakes

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

OPOD: On-Policy Omni Distillation

cs.AI · 2026-07-23 · conditional · novelty 5.0

OPOD trains one omni-modal model by routing each generated response to a matching modality teacher, and it reports the best twelve-benchmark average at 3B, 7B, and 30B scale.

citing papers explorer

Showing 1 of 1 citing paper.

  • OPOD: On-Policy Omni Distillation cs.AI · 2026-07-23 · conditional · none · ref 2

    OPOD trains one omni-modal model by routing each generated response to a matching modality teacher, and it reports the best twelve-benchmark average at 3B, 7B, and 30B scale.