Instruction-tuning reduces LLM output diversity, DPO causes the biggest drop, and conformative decoding, a log-probability mixture of instruct and base models, partly restores diversity while keeping quality.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
Instruction-tuning reduces LLM output diversity, DPO causes the biggest drop, and conformative decoding, a log-probability mixture of instruct and base models, partly restores diversity while keeping quality.