DAPD uses a self-conditioned bridge and bidirectional anchoring to match information between teacher and student during on-policy self-distillation, improving reasoning, coding, and instruction-following benchmarks over OPSD.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DAPD: Dual-Anchored Policy Distillation
DAPD uses a self-conditioned bridge and bidirectional anchoring to match information between teacher and student during on-policy self-distillation, improving reasoning, coding, and instruction-following benchmarks over OPSD.