Pith. sign in

REVIEW 1 cited by

Neural User Simulation for Corpus-based Policy Optimisation for Spoken Dialogue Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.06966 v1 pith:7FINZ67S submitted 2018-05-17 cs.CL cs.AIstat.ML

classification cs.CLcs.AIstat.ML
keywords userabussimulatorsdialoguesimulatorcorpus-basedevaluationneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems. For this task the Agenda-Based User Simulator (ABUS) is often used. The ABUS is based on hand-crafted rules and its output is in semantic form. Issues arise from both properties such as limited diversity and the inability to interface a text-level belief tracker. This paper introduces the Neural User Simulator (NUS) whose behaviour is learned from a corpus and which generates natural language, hence needing a less labelled dataset than simulators generating a semantic output. In comparison to much of the past work on this topic, which evaluates user simulators on corpus-based metrics, we use the NUS to train the policy of a reinforcement learning based Spoken Dialogue System. The NUS is compared to the ABUS by evaluating the policies that were trained using the simulators. Cross-model evaluation is performed i.e. training on one simulator and testing on the other. Furthermore, the trained policies are tested on real users. In both evaluation tasks the NUS outperformed the ABUS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MotiveBench shows that large language models still lag human consensus on motivational reasoning, with GPT-4o scoring 80.89 percent and chain-of-thought prompting usually reducing accuracy.

Pith tools