Pith. sign in

REVIEW 3 cited by

Hate and Offensive Speech Detection in Hindi and Marathi

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.12200 v2 pith:GVVXZH7T submitted 2021-10-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords hindimarathimodelsbasichatespeechtextdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sentiment analysis is the most basic NLP task to determine the polarity of text data. There has been a significant amount of work in the area of multilingual text as well. Still hate and offensive speech detection faces a challenge due to inadequate availability of data, especially for Indian languages like Hindi and Marathi. In this work, we consider hate and offensive speech detection in Hindi and Marathi texts. The problem is formulated as a text classification task using the state of the art deep learning approaches. We explore different deep learning architectures like CNN, LSTM, and variations of BERT like multilingual BERT, IndicBERT, and monolingual RoBERTa. The basic models based on CNN and LSTM are augmented with fast text word embeddings. We use the HASOC 2021 Hindi and Marathi hate speech datasets to compare these algorithms. The Marathi dataset consists of binary labels and the Hindi dataset consists of binary as well as more-fine grained labels. We show that the transformer-based models perform the best and even the basic models along with FastText embeddings give a competitive performance. Moreover, with normal hyper-parameter tuning, the basic models perform better than BERT-based models on the fine-grained Hindi dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Per-language architecture selection among generalists, specialists, and ensembles achieves 0.796 macro F1 across 22 languages in SemEval-2026 Task 9.

  2. LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LoRA fine-tuning of Nemo-Instruct achieves the best F1 among four LLMs on Devanagari hate speech detection (90.05%) and target identification (71.47%), but without baseline comparisons the approach's efficacy is unproven.

  3. A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages

    cs.CL 2024-11 conditional novelty 3.0 of 10

    A survey cataloging datasets, features, and machine-learning methods for automatic hate speech detection in low-resource languages, organized by world region, with an overview of open challenges.

Pith tools