Introduces CoMuMDR, a code-mixed, multi-modal, multi-domain discourse corpus for conversations, with benchmarks showing poor relation classification by current models.
An Annotation Scheme of A Large-scale Multi-party Dialogues Dataset for Discourse Parsing and Machine Comprehension
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we propose the scheme for annotating large-scale multi-party chat dialogues for discourse parsing and machine comprehension. The main goal of this project is to help understand multi-party dialogues. Our dataset is based on the Ubuntu Chat Corpus. For each multi-party dialogue, we annotate the discourse structure and question-answer pairs for dialogues. As we know, this is the first large scale corpus for multi-party dialogues discourse parsing, and we firstly propose the task for multi-party dialogues machine reading comprehension.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
Introduces CoMuMDR, a code-mixed, multi-modal, multi-domain discourse corpus for conversations, with benchmarks showing poor relation classification by current models.