Pith. sign in

REVIEW 1 cited by

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.11869 v1 pith:5KWYPIVM submitted 2018-05-30 cs.CL

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

classification cs.CL
keywords sarcasmcode-mixeddatasetdetectionenglish-hindilikemediasocial
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Social media platforms like twitter and facebook have be- come two of the largest mediums used by people to express their views to- wards different topics. Generation of such large user data has made NLP tasks like sentiment analysis and opinion mining much more important. Using sarcasm in texts on social media has become a popular trend lately. Using sarcasm reverses the meaning and polarity of what is implied by the text which poses challenge for many NLP tasks. The task of sarcasm detection in text is gaining more and more importance for both commer- cial and security services. We present the first English-Hindi code-mixed dataset of tweets marked for presence of sarcasm and irony where each token is also annotated with a language tag. We present a baseline su- pervised classification system developed using the same dataset which achieves an average F-score of 78.4 after using random forest classifier and performing 10-fold cross validation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

    cs.CL 2026-02 conditional novelty 6.0

    MixSarc is a new public Bangla–English code-mixed corpus of 9,087 sentences annotated for humor, sarcasm, offensiveness, and vulgarity, with benchmark results showing sarcasm and minority classes remain difficult.