Pith. sign in

REVIEW

When Crowd Meets Persona: Creating a Large-Scale Open-Domain Persona Dialogue Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00350 v1 pith:NCQKC2GM submitted 2023-04-01 cs.CL

classification cs.CL
keywords personacorpusdialogueopen-domaincreatingcrowdcrowdworkersdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building a natural language dataset requires caution since word semantics is vulnerable to subtle text change or the definition of the annotated concept. Such a tendency can be seen in generative tasks like question-answering and dialogue generation and also in tasks that create a categorization-based corpus, like topic classification or sentiment analysis. Open-domain conversations involve two or more crowdworkers freely conversing about any topic, and collecting such data is particularly difficult for two reasons: 1) the dataset should be ``crafted" rather than ``obtained" due to privacy concerns, and 2) paid creation of such dialogues may differ from how crowdworkers behave in real-world settings. In this study, we tackle these issues when creating a large-scale open-domain persona dialogue corpus, where persona implies that the conversation is performed by several actors with a fixed persona and user-side workers from an unspecified crowd.

Discussion (0). Continue with ORCID to comment.

Pith tools