Pith. sign in

REVIEW 12 cited by

MultiWOZ -- A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.00278 v3 pith:NRHUVB26 submitted 2018-09-29 cs.CL

classification cs.CL
keywords dialoguedatadatasetbeliefcollectionmulti-domainmultiwoztask-oriented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Even though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available. To address this fundamental obstacle, we introduce the Multi-Domain Wizard-of-Oz dataset (MultiWOZ), a fully-labeled collection of human-human written conversations spanning over multiple domains and topics. At a size of $10$k dialogues, it is at least one order of magnitude larger than all previous annotated task-oriented corpora. The contribution of this work apart from the open-sourced dataset labelled with dialogue belief states and dialogue actions is two-fold: firstly, a detailed description of the data collection procedure along with a summary of data structure and analysis is provided. The proposed data-collection pipeline is entirely based on crowd-sourcing without the need of hiring professional annotators; secondly, a set of benchmark results of belief tracking, dialogue act and response generation is reported, which shows the usability of the data and sets a baseline for future studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Detecting Inconsistencies in End-to-end Generated TODs

    cs.CL 2026-07 conditional novelty 7.0 of 10

    TOD consistency can be cast as a CSP over slot-value variables and six domain-independent constraints, detecting inconsistencies at 75.9% accuracy with GPT-4o variable extraction.

  2. After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions

    cs.HC 2026-02 conditional novelty 6.0 of 10

    A two-stage framework — category-structured fine-tuning on LLM-simulated personas plus on-device activation steering — improves proactive-assistant timing and perceived quality, though the biggest gains are measured w...

  3. Real-World En Call Center Transcripts Dataset with PII Redaction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CallCenterEN offers 91,706 PII-redacted English call center transcripts, spanning 10,448 hours of withheld audio, for conversational AI research and benchmarking.

  4. AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A cascaded open-weight speech agent using ReAct reasoning and external tools reaches 92.75% on VoiceBench OpenBookQA and 90% success on 30 multi-turn voice tasks.

  5. Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OKCV is a new human-annotated video dialogue dataset where answering questions requires both visual grounding in the video and external knowledge.

  6. SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors

    cs.AI 2025-05 conditional novelty 6.0 of 10

    SEAR is the first cited multimodal dataset of AR-plus-LLM social engineering, showing high self-reported trust and click intentions in a 60-person lab study.

  7. PATHFinder Agent for Tailored Prenatal Care

    cs.AI 2026-06 conditional novelty 5.0 of 10

    PATHFinder Agent drafts tailored prenatal care plans from patient dialogue and Michigan 211 resource lookups; GPT-5.2 scored 77.6% on expert rubrics, but no human validation is reported.

  8. Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

    cs.MA 2026-01 reject novelty 5.0 of 10

    AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...

  9. MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    MATRIX combines a structured safety taxonomy, an LLM hazard judge, and a patient simulator to benchmark clinical dialogue agents, claiming expert-level hazard detection and revealing weak emergency handling in current LLMs.

  10. Agent WARPP: Workflow Adherence via Runtime Parallel Personalization

    cs.AI 2025-07 conditional novelty 5.0 of 10

    WARPP prunes task-oriented dialogue workflows at runtime using user attributes and a parallel Personalizer agent, improving tool, parameter, and exact-match accuracy while cutting token use versus ReAct.

  11. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  12. Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.

Pith tools