Preference tuning on weak-versus-weaker response pairs can improve a strong LLM as much as tuning on strong supervision, via the relative quality delta.
math, code), we find pairs where Qwen 3B responds correctly but Qwen 1.5B does not (Figure A2)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
Preference tuning on weak-versus-weaker response pairs can improve a strong LLM as much as tuning on strong supervision, via the relative quality delta.