A debate-and-reflect distillation pipeline with tree-structured DPO improves small-model accuracy on MMLU Pro and MATH compared with standard distillation baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
A debate-and-reflect distillation pipeline with tree-structured DPO improves small-model accuracy on MMLU Pro and MATH compared with standard distillation baselines.