Feedback-augmented self-distillation, training a search agent on its own successful rollouts, fails to improve retrieval-interleaved search agents: it collapses into generic, question-agnostic reasoning templates, and an EMA teacher only partially stabilizes it.
Aligning Sizes of Intermediate Layers by
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?
Feedback-augmented self-distillation, training a search agent on its own successful rollouts, fails to improve retrieval-interleaved search agents: it collapses into generic, question-agnostic reasoning templates, and an EMA teacher only partially stabilizes it.