COLLATE improves small LLMs' reasoning by training multiple clones to generate diverse rationales, selecting the rationale that most increases ground-truth answer likelihood, and tuning via DPO.
15 samples ran- domly from the test sets of each of the 5 task datasets
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
COLLATE improves small LLMs' reasoning by training multiple clones to generate diverse rationales, selecting the rationale that most increases ground-truth answer likelihood, and tuning via DPO.