Pith. sign in

REVIEW 1 cited by

Harnessing Pre-Trained Sentence Transformers for Offensive Language Detection in Indian Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02249 v1 pith:CJOU5MKS submitted 2023-10-03 cs.CL cs.LG

classification cs.CLcs.LG
keywords hatespeechlanguagesoffensiveassamesebengalicontentdetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In our increasingly interconnected digital world, social media platforms have emerged as powerful channels for the dissemination of hate speech and offensive content. This work delves into the domain of hate speech detection, placing specific emphasis on three low-resource Indian languages: Bengali, Assamese, and Gujarati. The challenge is framed as a text classification task, aimed at discerning whether a tweet contains offensive or non-offensive content. Leveraging the HASOC 2023 datasets, we fine-tuned pre-trained BERT and SBERT models to evaluate their effectiveness in identifying hate speech. Our findings underscore the superiority of monolingual sentence-BERT models, particularly in the Bengali language, where we achieved the highest ranking. However, the performance in Assamese and Gujarati languages signifies ongoing opportunities for enhancement. Our goal is to foster inclusive online spaces by countering hate speech proliferation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages

    cs.CL 2024-11 conditional novelty 3.0 of 10

    A survey cataloging datasets, features, and machine-learning methods for automatic hate speech detection in low-resource languages, organized by world region, with an overview of open challenges.

Pith tools