Pith. sign in

REVIEW 1 cited by

BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08964 v3 pith:WQ5TXLTW submitted 2024-08-16 cs.CL

BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis

classification cs.CL
keywords code-mixedsentimentanalysisbengalidatadatasetacrossbengali-english
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fundamental text classification task across several languages for code-mixed data. However, there has yet to be a large-scale and diverse sentiment analysis dataset on code-mixed Bengali. We address this limitation by introducing BnSentMix, a sentiment analysis dataset on code-mixed Bengali consisting of 20,000 samples with 4 sentiment labels from Facebook, YouTube, and e-commerce sites. We ensure diversity in data sources to replicate realistic code-mixed scenarios. Additionally, we propose 14 baseline methods including novel transformer encoders further pre-trained on code-mixed Bengali-English, achieving an overall accuracy of 69.8% and an F1 score of 69.1% on sentiment classification tasks. Detailed analyses reveal variations in performance across different sentiment labels and text types, highlighting areas for future improvement.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

    cs.CL 2026-02 conditional novelty 6.0

    MixSarc is a new public Bangla–English code-mixed corpus of 9,087 sentences annotated for humor, sarcasm, offensiveness, and vulgarity, with benchmark results showing sarcasm and minority classes remain difficult.