REVIEW 12 cited by
MultiWOZ -- A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Even though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available. To address this fundamental obstacle, we introduce the Multi-Domain Wizard-of-Oz dataset (MultiWOZ), a fully-labeled collection of human-human written conversations spanning over multiple domains and topics. At a size of $10$k dialogues, it is at least one order of magnitude larger than all previous annotated task-oriented corpora. The contribution of this work apart from the open-sourced dataset labelled with dialogue belief states and dialogue actions is two-fold: firstly, a detailed description of the data collection procedure along with a summary of data structure and analysis is provided. The proposed data-collection pipeline is entirely based on crowd-sourcing without the need of hiring professional annotators; secondly, a set of benchmark results of belief tracking, dialogue act and response generation is reported, which shows the usability of the data and sets a baseline for future studies.
Forward citations
Cited by 12 Pith papers
-
Towards Detecting Inconsistencies in End-to-end Generated TODs
TOD consistency can be cast as a CSP over slot-value variables and six domain-independent constraints, detecting inconsistencies at 75.9% accuracy with GPT-4o variable extraction.
-
After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
A two-stage framework — category-structured fine-tuning on LLM-simulated personas plus on-device activation steering — improves proactive-assistant timing and perceived quality, though the biggest gains are measured w...
-
Real-World En Call Center Transcripts Dataset with PII Redaction
CallCenterEN offers 91,706 PII-redacted English call center transcripts, spanning 10,448 hours of withheld audio, for conversational AI research and benchmarking.
-
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
A cascaded open-weight speech agent using ReAct reasoning and external tools reaches 92.75% on VoiceBench OpenBookQA and 90% success on 30 multi-turn voice tasks.
-
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
OKCV is a new human-annotated video dialogue dataset where answering questions requires both visual grounding in the video and external knowledge.
-
SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors
SEAR is the first cited multimodal dataset of AR-plus-LLM social engineering, showing high self-reported trust and click intentions in a 60-person lab study.
-
PATHFinder Agent for Tailored Prenatal Care
PATHFinder Agent drafts tailored prenatal care plans from patient dialogue and Michigan 211 resource lookups; GPT-5.2 scored 77.6% on expert rubrics, but no human validation is reported.
-
Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)
AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...
-
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
MATRIX combines a structured safety taxonomy, an LLM hazard judge, and a patient simulator to benchmark clinical dialogue agents, claiming expert-level hazard detection and revealing weak emergency handling in current LLMs.
-
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
WARPP prunes task-oriented dialogue workflows at runtime using user attributes and a parallel Personalizer agent, improving tool, parameter, and exact-match accuracy while cutting token use versus ReAct.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
-
Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality
A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.
Discussion (0). Sign in to comment.