Pith. sign in

REVIEW 2 cited by

L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.01159 v2 pith:F6X6E2YZ submitted 2022-02-02 cs.CL cs.LG

classification cs.CLcs.LG
keywords marathicorpuslanguageresourcesmodelsmonolingualdatal3cube-mahacorpus
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources. We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We further present, MahaBERT, MahaAlBERT, and MahaRoBerta all BERT-based masked language models, and MahaFT, the fast text word embeddings both trained on full Marathi corpus with 752M tokens. We show the effectiveness of these resources on downstream Marathi sentiment analysis, text classification, and named entity recognition (NER) tasks. We also release MahaGPT, a generative Marathi GPT model trained on Marathi corpus. Marathi is a popular language in India but still lacks these resources. This work is a step forward in building open resources for the Marathi language. The data and models are available at https://github.com/l3cube-pune/MarathiNLP .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new human-corrected Marathi paraphrase detection corpus with 8,000 pairs in five difficulty buckets, benchmarked with BERT models, with MahaBERT reaching 88.7% F1.

  2. Topic Modeling in Marathi

    cs.CL 2025-02 conditional novelty 4.0 of 10

    BERTopic with Indic BERT embeddings produces higher topic coherence scores than LDA on Marathi datasets of long, medium, and short documents.

Pith tools