Teacher top-K distillation drops the low-probability <tool call> token from the response teacher's support, creating a one-sided gradient that causally raises tool over-calling.
EMNLP 2025 (Oral)
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
ADWIN adaptively selects training horizons in on-policy distillation via prefix alignment checks, cutting end-to-end cost by up to 4.1x while matching or exceeding full-rollout accuracy on math and code benchmarks.
Position-Weighted On-Policy Self-Distillation (PW-OPSD) weights later tokens more heavily after a diagnostic shows position predicts teacher reliability better than entropy, yielding +1.0 and +1.1 Avg@12 gains on AIME 2024/2025.
citing papers explorer
-
When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
Teacher top-K distillation drops the low-probability <tool call> token from the response teacher's support, creating a one-sided gradient that causally raises tool over-calling.
-
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
ADWIN adaptively selects training horizons in on-policy distillation via prefix alignment checks, cutting end-to-end cost by up to 4.1x while matching or exceeding full-rollout accuracy on math and code benchmarks.
-
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
Position-Weighted On-Policy Self-Distillation (PW-OPSD) weights later tokens more heavily after a diagnostic shows position predicts teacher reliability better than entropy, yielding +1.0 and +1.1 Avg@12 gains on AIME 2024/2025.