DPO fine-tuning of the summary and recommendation writers improves conversational recommendation ranking on two Japanese datasets, but the evaluation shares the scorer that generated the training signal.
Amatriain and Justin D
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Refining Text Generation for Realistic Conversational Recommendation via Direct Preference Optimization
DPO fine-tuning of the summary and recommendation writers improves conversational recommendation ranking on two Japanese datasets, but the evaluation shares the scorer that generated the training signal.