Pith. sign in

REVIEW 3 cited by

Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12813 v1 pith:BWOSH3PM submitted 2025-02-18 cs.CL

Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models

classification cs.CL
keywords userdialoguellmsprofilestask-orientedanalysisattributesconversational
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this study, we explore the application of Large Language Models (LLMs) for generating synthetic users and simulating user conversations with a task-oriented dialogue system and present detailed results and their analysis. We propose a comprehensive novel approach to user simulation technique that uses LLMs to create diverse user profiles, set goals, engage in multi-turn dialogues, and evaluate the conversation success. We employ two proprietary LLMs, namely GPT-4o and GPT-o1 (Achiam et al., 2023), to generate a heterogeneous base of user profiles, characterized by varied demographics, multiple user goals, different conversational styles, initial knowledge levels, interests, and conversational objectives. We perform a detailed analysis of the user profiles generated by LLMs to assess the diversity, consistency, and potential biases inherent in these LLM-generated user simulations. We find that GPT-o1 generates more heterogeneous user distribution across most user attributes, while GPT-4o generates more skewed user attributes. The generated set of user profiles are then utilized to simulate dialogue sessions by interacting with a task-oriented dialogue system.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

    cs.CL 2026-06 unverdicted novelty 6.0

    RUT-Bench evaluates 19 LLMs on realistic user tool-calling scenarios and finds success rates below 40% with further drops on non-ideal inputs.

  2. WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

    cs.CL 2026-06 unverdicted novelty 6.0

    WRIT is a synthesis pipeline that generates write-read intensive trajectories along axes of write-decision count and per-decision evidence burden, enabling a 4B model to outperform GPT-5.1 on τ²-bench with reduced inf...

  3. MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

    cs.AI 2026-01 unverdicted novelty 6.0

    MirrorBench defines a reproducible benchmark combining lexical metrics (MATTR, Yule's K, HD-D) and LLM-judge metrics with calibration controls to measure human-likeness of user-proxy agents across four datasets.