REVIEW 4 cited by
L3Cube-MahaNLP: Marathi Natural Language Processing Datasets, Models, and Library
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite being the third most popular language in India, the Marathi language lacks useful NLP resources. Moreover, popular NLP libraries do not have support for the Marathi language. With L3Cube-MahaNLP, we aim to build resources and a library for Marathi natural language processing. We present datasets and transformer models for supervised tasks like sentiment analysis, named entity recognition, and hate speech detection. We have also published a monolingual Marathi corpus for unsupervised language modeling tasks. Overall we present MahaCorpus, MahaSent, MahaNER, and MahaHate datasets and their corresponding MahaBERT models fine-tuned on these datasets. We aim to move ahead of benchmark datasets and prepare useful resources for Marathi. The resources are available at https://github.com/l3cube-pune/MarathiNLP.
Forward citations
Cited by 4 Pith papers
-
MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
A new human-corrected Marathi paraphrase detection corpus with 8,000 pairs in five difficulty buckets, benchmarked with BERT models, with MahaBERT reaching 88.7% F1.
-
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
A new 15,000-sentence Marathi emotion benchmark shows GPT-4 and Llama3-405B outperform fine-tuned Marathi BERT and MuRIL, while BERT trained on GPT-4-generated labels still trails GPT-4.
-
SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods
The authors release sense-annotated WSD/WiC datasets for ten low-resource languages and report that English-based zero-shot transfer often beats small in-language fine-tuning, while mixed training usually helps.
-
BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
Fine-tuned MahaBERT models reach 0.88–0.91 F1 on Marathi NER, beating zero-shot LLMs (0.57–0.69) by more than 20 points.
Discussion (0). Continue with ORCID to comment.