REVIEW 4 cited by
Data Bootstrapping Approaches to Improve Low Resource Abusive Language Detection for Indic Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Abusive language is a growing concern in many social media platforms. Repeated exposure to abusive speech has created physiological effects on the target users. Thus, the problem of abusive language should be addressed in all forms for online peace and safety. While extensive research exists in abusive speech detection, most studies focus on English. Recently, many smearing incidents have occurred in India, which provoked diverse forms of abusive speech in online space in various languages based on the geographic location. Therefore it is essential to deal with such malicious content. In this paper, to bridge the gap, we demonstrate a large-scale analysis of multilingual abusive speech in Indic languages. We examine different interlingual transfer mechanisms and observe the performance of various multilingual models for abusive speech detection for eight different Indic languages. We also experiment to show how robust these models are on adversarial attacks. Finally, we conduct an in-depth error analysis by looking into the models' misclassified posts across various settings. We have made our code and models public for other researchers.
Forward citations
Cited by 4 Pith papers
-
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
HateDay provides a representative global sample of one day of Twitter and shows that hate speech detection models achieve much lower average precision on real-world data than on academic datasets.
-
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
A gated fusion head that conditions English toxicity, Indic abuse, and rule-based severity scores on the text context improves code-mixed abuse detection in 10/12 in-domain and 7/8 transfer comparisons.
-
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
The best overall F1 scores come from larger models such as Qwen2.5-VL and DeBERTa-v3, but performance is dataset-dependent and victim detection remains poor across all model families.
-
NLPineers@ NLU of Devanagari Script Languages 2025: Hate Speech Detection using Ensembling of BERT-based models
An ensemble of XLM-RoBERTa, MuRIL, and an abusive-tuned MuRIL achieved 0.7762 recall and 0.6914 F1 on Devanagari hate speech detection.
Discussion (0). Continue with ORCID to comment.