SMTPO uses multi-task SFT to improve simulator feedback quality and RL with fine-grained rewards to optimize multi-turn preference reasoning in LLM-based conversational recommendation.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.IR 2years
2026 2roles
background 1polarities
background 1representative citing papers
Optimal preference elicitation in conversational recommenders is stage-dependent (attributes early, items later), and a MoE model trained on a new annotated dataset improves offline recommendation and response quality.
citing papers explorer
-
User Simulator-Guided Multi-Turn Preference Optimization for Reasoning LLM-based Conversational Recommendation
SMTPO uses multi-task SFT to improve simulator feedback quality and RL with fine-grained rewards to optimize multi-turn preference reasoning in LLM-based conversational recommendation.
-
When and How to Ask: Dynamic Preference Elicitation Strategies for Conversational Recommendation
Optimal preference elicitation in conversational recommenders is stage-dependent (attributes early, items later), and a MoE model trained on a new annotated dataset improves offline recommendation and response quality.