Pith. sign in

REVIEW 4 cited by

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00085 v1 pith:R3ZPIC54 submitted 2020-04-30 cs.CL

classification cs.CL
keywords corpusembeddingsindicnlplanguagesai4bharat-indicnlpavailablecorporaindic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article category classification datasets for 9 languages to evaluate the embeddings. We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks. We hope that the availability of the corpus will accelerate Indic NLP research. The resources are available at https://github.com/ai4bharat-indicnlp/indicnlp_corpus.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Inspect India Evals provides six India-specific LLM benchmarks and preliminary scores for five open-weight models, with most performance gaps not statistically significant at n=5.

  2. MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new human-corrected Marathi paraphrase detection corpus with 8,000 pairs in five difficulty buckets, benchmarked with BERT models, with MahaBERT reaching 88.7% F1.

  3. SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The authors release sense-annotated WSD/WiC datasets for ten low-resource languages and report that English-based zero-shot transfer often beats small in-language fine-tuning, while mixed training usually helps.

  4. BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

    cs.LG 2026-08 reject novelty 4.0 of 10

    A 90%-pruned few-shot Bengali model is reported to rival larger baselines on some tasks, but the reported F1 scores contradict the paper's own precision and recall values.

Pith tools