A literature review that catalogs sarcasm detection datasets, word-embedding strategies, and neural models, but adds no new experimental results.
A Large Self-Annotated Corpus for Sarcasm
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous dataset -- and many times more instances of non-sarcastic statements, allowing for learning in both balanced and unbalanced label regimes. Each statement is furthermore self-annotated -- sarcasm is labeled by the author, not an independent annotator -- and provided with user, topic, and conversation context. We evaluate the corpus for accuracy, construct benchmarks for sarcasm detection, and evaluate baseline methods.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Was that Sarcasm?: A Literature Survey on Sarcasm Detection
A literature review that catalogs sarcasm detection datasets, word-embedding strategies, and neural models, but adds no new experimental results.