Pith. sign in

REVIEW 2 cited by

The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.09969 v4 pith:2BMLPVV3 submitted 2019-11-22 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords dialoguejddccorpusmillionchineseconversationconversationsdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MERCI: Multimodal Emotional and peRsonal Conversational Interactions Dataset

    cs.HC 2024-12 conditional novelty 6.0 of 10

    MERCI is a 30-participant multimodal human-robot conversation dataset that pairs personal profiles and emotion labels with video, audio, and text records.

  2. From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification

    cs.CL 2024-11 conditional novelty 5.0 of 10

    An LLM-enhanced HMM generates intent-aware multilingual e-commerce dialogues, and a contrastive multi-task classifier (MINT-CL) improves multi-turn intent classification accuracy by about 0.5 percent on average.

Pith tools