← back to paper
arxiv: 2608.09447 · 2 revisions
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training