Naamah is a silver-standard Sanskrit NER dataset of 102,942 sentences generated by seeding DBpedia entities into a 24B-parameter LLM to produce grammatically natural training data, then used to benchmark XLM-RoBERTa and IndicBERTv2.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation
Naamah is a silver-standard Sanskrit NER dataset of 102,942 sentences generated by seeding DBpedia entities into a 24B-parameter LLM to produce grammatically natural training data, then used to benchmark XLM-RoBERTa and IndicBERTv2.