REVIEW 4 cited by
Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Synthetic users are cost-effective proxies for real users in the evaluation of conversational recommender systems. Large language models show promise in simulating human-like behavior, raising the question of their ability to represent a diverse population of users. We introduce a new protocol to measure the degree to which language models can accurately emulate human behavior in conversational recommendation. This protocol is comprised of five tasks, each designed to evaluate a key property that a synthetic user should exhibit: choosing which items to talk about, expressing binary preferences, expressing open-ended preferences, requesting recommendations, and giving feedback. Through evaluation of baseline simulators, we demonstrate these tasks effectively reveal deviations of language models from human behavior, and offer insights on how to reduce the deviations with model selection and prompting strategies.
Forward citations
Cited by 4 Pith papers
-
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
On τ-bench, LLM user simulators are more cooperative, more verbose, and more lenient than real human users, so agent benchmarks that rely on them overstate real-world performance.
-
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
A design proposal, not a tested system: a dual-view simulation loop in which an LLM-simulated user issues reflective instructions when about to disengage, and an instruction-following recommender adapts to maximize si...
-
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
A new 109-conversation, 86-API benchmark for LLM function-calling in multi-turn dialogue shows top models at about 40% accuracy and near-zero performance on chains of 4+ calls.
-
A Survey on LLM-powered Agents for Recommender Systems
The paper categorizes LLM-powered agents for recommender systems into recommender-oriented, interaction-oriented, and simulation-oriented paradigms and describes a common four-module agent architecture.
Discussion (0). Continue with ORCID to comment.