REVIEW 2 cited by
ConvLab-3: A Flexible Dialogue System Toolkit Based on a Unified Data Format
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Task-oriented dialogue (TOD) systems function as digital assistants, guiding users through various tasks such as booking flights or finding restaurants. Existing toolkits for building TOD systems often fall short of in delivering comprehensive arrays of data, models, and experimental environments with a user-friendly experience. We introduce ConvLab-3: a multifaceted dialogue system toolkit crafted to bridge this gap. Our unified data format simplifies the integration of diverse datasets and models, significantly reducing complexity and cost for studying generalization and transfer. Enhanced with robust reinforcement learning (RL) tools, featuring a streamlined training process, in-depth evaluation tools, and a selection of user simulators, ConvLab-3 supports the rapid development and evaluation of robust dialogue policies. Through an extensive study, we demonstrate the efficacy of transfer learning and RL and showcase that ConvLab-3 is not only a powerful tool for seasoned researchers but also an accessible platform for newcomers.
Forward citations
Cited by 2 Pith papers
-
CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues
An LLM-based pipeline adds synthetic misunderstandings, non-understandings, and vaguely related questions plus repair responses to MultiWOZ dialogues, and the released CoPrUS-MultiWOZ dataset preserves task performanc...
-
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
ReAct-prompted GPT-3.5 and GPT-4 underperform classical task-oriented dialogue systems on task success, but humans rate them as more satisfying despite lower success.
Discussion (0). Continue with ORCID to comment.