TRB introduces a KL-trust-region warmup for on-policy distillation that blends toward teacher behavior early in training and anneals to zero, reporting the highest average performance across two math-reasoning distillation experiments.
Micota: Bridgingthelearnabilitygapwithintermediatecotandteacherassistants.ArXiv, abs/2507.01887
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
A trajectory for student-LLM distillation is better when its tokens are surprising but still high-ranked, and the ratio of average rank to average surprisal (RSR) selects such trajectories better than existing metrics.
citing papers explorer
-
Trust-Region Behavior Blending for On-Policy Distillation
TRB introduces a KL-trust-region warmup for on-policy distillation that blends toward teacher behavior early in training and anneals to zero, reporting the highest average performance across two math-reasoning distillation experiments.
-
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
A trajectory for student-LLM distillation is better when its tokens are surprising but still high-ranked, and the ratio of average rank to average surprisal (RSR) selects such trajectories better than existing metrics.