PrefEval, a benchmark of 3,000 preference-query pairs in multi-session conversations up to 100k tokens, finds that most LLMs follow user preferences poorly beyond a few turns, though fine-tuning helps.
User: The H ˆotel des Deux ˆIles sounds perfect for my needs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
PrefEval, a benchmark of 3,000 preference-query pairs in multi-session conversations up to 100k tokens, finds that most LLMs follow user preferences poorly beyond a few turns, though fine-tuning helps.