Pith. sign in

REVIEW 4 cited by

L3Cube-MahaNLP: Marathi Natural Language Processing Datasets, Models, and Library

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14728 v2 pith:TX7TB5GP submitted 2022-05-29 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagemarathidatasetsresourcesmodelsl3cube-mahanlplibrarynatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite being the third most popular language in India, the Marathi language lacks useful NLP resources. Moreover, popular NLP libraries do not have support for the Marathi language. With L3Cube-MahaNLP, we aim to build resources and a library for Marathi natural language processing. We present datasets and transformer models for supervised tasks like sentiment analysis, named entity recognition, and hate speech detection. We have also published a monolingual Marathi corpus for unsupervised language modeling tasks. Overall we present MahaCorpus, MahaSent, MahaNER, and MahaHate datasets and their corresponding MahaBERT models fine-tuned on these datasets. We aim to move ahead of benchmark datasets and prepare useful resources for Marathi. The resources are available at https://github.com/l3cube-pune/MarathiNLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new human-corrected Marathi paraphrase detection corpus with 8,000 pairs in five difficulty buckets, benchmarked with BERT models, with MahaBERT reaching 88.7% F1.

  2. L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 15,000-sentence Marathi emotion benchmark shows GPT-4 and Llama3-405B outperform fine-tuned Marathi BERT and MuRIL, while BERT trained on GPT-4-generated labels still trails GPT-4.

  3. SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The authors release sense-annotated WSD/WiC datasets for ten low-resource languages and report that English-based zero-shot transfer often beats small in-language fine-tuning, while mixed training usually helps.

  4. BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Fine-tuned MahaBERT models reach 0.88–0.91 F1 on Marathi NER, beating zero-shot LLMs (0.57–0.69) by more than 20 points.

Pith tools