Pith. sign in

REVIEW 1 cited by

MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.00773 v2 pith:OTG62HDZ submitted 2021-04-01 cs.CL

classification cs.CL
keywords multiwozannotationsdialoguestatedatasetevaluationmodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The MultiWOZ 2.0 dataset has greatly stimulated the research of task-oriented dialogue systems. However, its state annotations contain substantial noise, which hinders a proper evaluation of model performance. To address this issue, massive efforts were devoted to correcting the annotations. Three improved versions (i.e., MultiWOZ 2.1-2.3) have then been released. Nonetheless, there are still plenty of incorrect and inconsistent annotations. This work introduces MultiWOZ 2.4, which refines the annotations in the validation set and test set of MultiWOZ 2.1. The annotations in the training set remain unchanged (same as MultiWOZ 2.1) to elicit robust and noise-resilient model training. We benchmark eight state-of-the-art dialogue state tracking models on MultiWOZ 2.4. All of them demonstrate much higher performance than on MultiWOZ 2.1.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs

    cs.CL 2025-09 conditional novelty 4.0 of 10

    On MultiWOZ 2.1 multi-intent classification, Mistral-7B-v0.1 beats Llama-2-7B and Yi-6B in few-shot prompting (weighted F1 0.50), while supervised BERT remains far stronger (F1 0.92).

Pith tools