{"id":"acc5ace7-f072-489a-b2d0-c33b79386c15","arxiv_id":"2412.15471","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review of Marathi NLP resources, models, and evaluation metrics, with no new experiments or data.","lead":"Marathi is one of India's major languages, but its NLP tooling has lagged behind English. This paper collects the datasets, tokenization methods, neural models, and evaluation metrics currently used in Marathi NLP, serving mainly as a starting point for researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's 'state-of-the-art' claim is unsupported because resource selection has no stated protocol or cutoff, and included descriptions contain checkable inaccuracies such as the XLM-R parameter count.","rationale":"I read this paper as a survey whose usefulness is exactly its reliability as a map of Marathi NLP. The historical narrative and the sections on pipeline, tokenization, models, and metrics are generally coherent, and the cited primary resources are real. I give credit for covering both corpora and models and for noting the limitations of string-based metrics for morphologically rich languages. The weak point is not the absence of new science; it is the uncheckable boundary of the inventory and the concrete metadata error in Section 7.3. The reader's weakest assumption identified the representativeness of the selection; I agree with that and add that even the described resources are not always accurately characterized. Since the reader's verdict was already UNVERDICTED, my concern does not move the verdict; it sharpens the condition under which the survey could be trusted. If the concrete test passes, the survey would be a useful orientation aid; if it fails, the 'SOTA' label should be replaced by a dated, scope-limited statement.","tokens_in":14324,"tokens_out":5517,"duration_ms":44280,"concrete_test":"Compile an independent validation set by searching ACL Anthology and arXiv for 'Marathi' combined with 'NLP', 'corpus', 'language model', or 'dataset' over 2018-2024, and add all resources cited in Lahoti et al. (2022) plus any Marathi resource with over 100 citations found by the search. Then check whether the survey covers each item and whether the reported statistics match the primary source, starting with the XLM-R parameter count in Section 7.3. If a major current resource is missing or the parameter count is not 279M, the 'state-of-the-art' and 'broad overview' claims should be downgraded to a dated, scope-limited survey.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this paper presents a broad overview and 'state-of-the-art' Marathi NLP resources and tools. For that claim to hold, two conditions must be met: (i) the selected resources must be representative and current, and (ii) their descriptions must be accurate. Neither condition is currently verifiable. Section 2 compares the paper with Lahoti et al. (2022) but never states how this survey's inventory was assembled: no search protocol, no inclusion or exclusion criteria, and no cutoff date. The phrase 'state-of-the-art' and the line in Section 7.6, 'Current SOTA model is NLLB 54B MOE,' presuppose a temporal boundary that is never given. The accuracy problem is not hypothetical: Section 7.3 reports 'XLM-RoBERTa (279M parameters),' whereas Conneau et al. (2020) report XLM-R Base at 270M and XLM-R Large at 559M parameters. A reader relying on this survey for resource selection would be misled about a widely used model. Because the paper's only product is a trustworthy summary of the field, these unverifiable scope and accuracy gaps are load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of Marathi NLP covering the evolution of Indic NLP research, training corpora and benchmarks, tokenization methods, neural models (BERT, BART, mT5, MahaBERT, XLM-R, IndicBERT, IndicBART, IndicTrans2), and evaluation metrics (BLEU, chrF++, BERTScore, COMET, ROUGE, BLEURT). The abstract claims to provide a broad overview of the field and of state-of-the-art resources and tools for Marathi. The paper cites many primary sources and positions itself relative to the earlier survey by Lahoti et al. (2022), but it does not describe any search protocol, inclusion criteria, or cutoff date for the resources surveyed, and it contains several checkable technical inaccuracies in model descriptions.","tokens_in":14532,"tokens_out":3996,"duration_ms":28786,"significance":"If made accurate, this survey would be a useful entry point for researchers new to Marathi NLP: it assembles in one place the main corpora (EMILLE, IndicNLP, Samanantar, FLORES, OPUS/NLLB), the main models (MahaBERT family, mT5, XLM-R, IndicBERT, IndicBART, IndicTrans2), and the main evaluation metrics, with references to the primary literature. Its strengths are breadth of coverage and the inclusion of very recent resources such as IndicTrans2 and mahaNLP. However, because the paper's only product is a trustworthy summary of the field, its reliability is load-bearing: the lack of a documented selection methodology and the presence of concrete factual errors in resource descriptions currently undercut the 'state-of-the-art' claim. The paper has no machine-checked proofs or code, but that is not expected for a survey; the obligation is instead accuracy and transparency of coverage.","major_comments":[{"comment":"The central claim of a 'broad overview' and 'state-of-the-art resources and tools' is not verifiable because the manuscript never states how the resource inventory was assembled. Section 2 compares the paper with Lahoti et al. (2022) but provides no search protocol, no inclusion or exclusion criteria, and no cutoff date. Please add a methodology paragraph specifying the databases searched, the time window, and the criteria for including or excluding resources; this is essential for a survey whose value depends on representativeness and currency.","section":"Abstract, Section 2"},{"comment":"The reported parameter count for XLM-RoBERTa is incorrect. The text states 'Their best model XLM-RoBERTa (279M parameters)', but Conneau et al. (2020) report XLM-R Base at approximately 270M parameters and XLM-R Large at approximately 559M parameters. Since readers may rely on this survey to choose models, this factual error should be corrected with the specific variant and its actual parameter count.","section":"Section 7.3"},{"comment":"The description of BERT conflates BERT with mBERT. The sentence 'It is also known as mBERT or multilingual-BERT' is inaccurate: BERT is the monolingual English model, while mBERT is the multilingual variant trained on 104 languages including Marathi. The distinction matters because Section 7.3 and Section 7.4 compare XLM-R and IndicBERT against mBERT. Please revise the text to distinguish the two models explicitly.","section":"Section 6.1"},{"comment":"The BERT hyperparameter definitions are mislabeled. The manuscript says 'A (the number of Attention Layers), L (the number of Encoder Layers), and H (the number of Hidden Layers)', but in the BERT paper A is the number of attention heads, L is the number of transformer layers (encoder blocks), and H is the hidden size. This makes the subsequent values BER T BASE (L=12, H=768, A=12) confusing. Please correct these definitions.","section":"Section 6.1"},{"comment":"The assertion 'Current SOTA (State-Of-The-Art) model is NLLB 54B MOE' is unsupported and time-dependent. No benchmark, task, or leaderboard is cited for this claim, and 'current' is meaningless without a cutoff date. Even if intended as 'state of the art for supervised machine translation at the time of writing', the claim needs a citation, a date, and a scope restriction; otherwise it should be removed or qualified.","section":"Section 7.6"}],"minor_comments":[{"comment":"The final paragraph contains typographical and grammatical errors: 'its too large to be to be deployed' should be 'it is too large to be deployed', and 'However, its too large' should be 'However, it is too large'.","section":"Section 7.6"},{"comment":"The section heading 'BER T' has an unintended space and should be 'BERT'.","section":"Section 6.1"},{"comment":"The sentence 'It uses sacreBLEU to compute the scores' is imprecise: chrF++ is an evaluation metric, and sacreBLEU is a software implementation that can compute it. Please rephrase to clarify that the authors refer to the reference implementation available through sacreBLEU.","section":"Section 8.2"},{"comment":"Several performance claims are reported from the original papers, but the text sometimes says 'Authors claim' and sometimes states results as fact. Please consistently mark third-party claims as claims, since this survey does not independently evaluate them.","section":"Section 7.2"},{"comment":"There are minor reference formatting issues: 'Tomá Mikolov' should be 'Tomáš Mikolov', 'A vik Bhattacharyya' should be 'Avik Bhattacharyya', and 'Gokul N.C.' is inconsistently spaced. These should be cleaned up.","section":"References"},{"comment":"The conclusion states 'the morphological richness make it harder for the models to deal with dialect variation' and 'Cross-lingual Information Retrieval (IR) and Question Answering (QA) is quite limited'; these subject-verb agreement errors should be corrected.","section":"Section 9"}],"recommendation":"major_revision","confidential_remarks":"The manuscript overlaps substantially with the Lahoti et al. (2022) survey it cites, and its incremental contribution currently rests on the inclusion of more recent models and resources. The absence of a documented methodology and the presence of several factual errors in model descriptions make the survey unreliable as a reference in its present form, but these issues are fixable within the scope of a revision. I would encourage the editor to require a methodology statement, a cutoff date, and a careful fact-check of all parameter counts and model names before any archival publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: if you need a quick map of Marathi NLP resources, this is a reasonable starting point. It has no new science, and the survey's own framing overreaches. But it does collect a solid set of references.\n\nWhat's good: the paper touches the major corpora (EMILLE, IndicNLP, Samanantar, FLORES, OPUS), the main models (mBERT, XLM-R, IndicBERT, IndicBART, IndicTrans2, MahaBERT), and evaluation metrics. For a newcomer, that list alone is useful. The pipeline figure and the narrative from rule-based to neural are fine.\n\nWhere it gets shaky: the authors call it 'state-of-the-art' but give no search protocol, inclusion criteria, or cutoff date. You can't verify that anything is missing or that the SOTA claim is current. In a survey that's a real gap.\n\nAlso, there are checkable mistakes. Section 6.1 says BERT 'is also known as mBERT' — that's wrong; mBERT is the multilingual variant. The hyperparameter definitions are flipped: A is attention heads, not attention layers; H is hidden size, not hidden layers. Section 7.3 reports XLM-RoBERTa as '279M parameters' — the XLM-R paper lists 270M for base and 559M for large. These are the kind of errors that a reader will hit immediately. They are not fatal, but for a survey whose sole product is accurate summary, they matter.\n\nI'm being asked whether this deserves serious refereeing. I'd say yes, if the venue publishes surveys. The topic is real and the consolidation has value. But it needs a real revision: state the scope and cutoff, fix the technical inaccuracies, and probably tone down the 'state-of-the-art' language. If you want to use it as a citation anchor, wait until it's cleaned up.\n\nBottom line: useful as a working document, not as a citable reference yet.","headline":"A useful but uneven survey of Marathi NLP; the resource coverage is broad, but the 'state-of-the-art' claim is under-supported and several technical details are wrong.","tokens_in":14964,"tokens_out":2442,"would_cite":false,"duration_ms":17821,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps the evolution of Marathi NLP from early rule-based tools to modern multilingual transformer models and identifies the resources and gaps that define the field.","keywords":["Marathi NLP","Indic languages","language resources","neural machine translation","multilingual language models","evaluation metrics","code-mixing","NLP survey"],"falsifier":"Compare the survey's resource inventory against a systematic, time-bounded search of Marathi NLP publications and public model hubs; if major recent datasets or models are absent, or if reported corpus sizes and benchmark numbers do not match the cited releases, the 'state-of-the-art' framing fails.","tokens_in":14142,"feed_emoji":"📚","tokens_out":7847,"duration_ms":59564,"temperature":0.7,"pith_summary":"This paper is a survey of Marathi natural language processing. It argues that NLP advances did not reach Marathi quickly because of script diversity, scarce public resources, and the language's rich morphology, but that the past decade of Indic-language initiatives has changed the situation. The survey's central claim is that Marathi NLP now has a broad, if uneven, set of corpora, pretrained models, benchmarks, and evaluation metrics, and it maps those resources onto the stages of a neural NLP pipeline. A reader interested in building or evaluating Marathi NLP systems would use this as a starting map of what exists and where the gaps are.","feed_headline":"From rule-based tools to 752M-token models: Marathi NLP surveyed","feed_subtitle":"A review of corpora, models, and metrics shows Marathi now has large datasets and strong baselines, and names the gaps.","key_machinery":"The organizing device is the nine-step neural NLP processing pipeline—data collection and preprocessing, tokenization, embedding creation, model training, evaluation, fine-tuning, inference, post-processing, and deployment—used as a checklist against which every Marathi resource and tool is placed. The paper pairs this pipeline with a taxonomy of corpora (monolingual versus parallel) and of evaluation metrics (string-based versus model-based). This is what lets the survey convert a list of resources into a statement about where Marathi NLP stands.","core_discovery":"On the paper's own terms, the discovery is a synthesis: Marathi has moved from being a low-resource language with early rule-based morphological analyzers and small corpora to one with large monolingual and parallel corpora (the paper cites, among others, a 142-million-word IndicNLP corpus and a 752-million-token MahaCorpus), multilingual and Marathi-specific transformer models (mBERT, XLM-R, mT5, IndicBERT, IndicBART, MahaBERT, and IndicTrans2, which covers all 22 scheduled Indian languages), and a set of evaluation practices that favor character-level and model-based metrics such as chrF++, BERTScore, and COMET over raw BLEU. It also identifies what remains missing: code-mixed Marathi-English and Marathi-Hindi data, annotated machine-reading-comprehension and summarization datasets, and robust handling of dialect variation in tasks like text-to-speech.","pith_inferences":["If the survey's inventory is representative, the practical bottleneck in Marathi NLP has shifted from raw data scarcity to scarcity of high-quality annotated task data, which suggests annotation efforts may now have higher marginal value than further corpus crawling.","The same resource pattern likely holds for other moderately resourced Indo-Aryan languages, so the pipeline-and-corpus map in this survey could serve as a template for surveys of Hindi, Gujarati, or Bengali NLP.","A testable extension would be to track the growth of Marathi corpora by year and correlate corpus size with benchmark gains; the paper's data (from 2.2 million words in EMILLE to 142 million in IndicNLP to 752 million tokens in MahaCorpus) suggests a steep recent takeoff."],"forward_implications":["Marathi NLP no longer lacks basic building blocks: large monolingual corpora and English-Marathi parallel corpora are publicly available, so new work can start from pretrained models rather than from data collection.","Because the paper reports that monolingual Marathi models (MahaBERT and related models) claim state-of-the-art results on sentiment analysis, NER, and text classification, a practitioner would reasonably start from these rather than from multilingual models for Marathi tasks.","Evaluation of Marathi generation should move beyond BLEU toward character-level metrics like chrF++ and model-based metrics like COMET, since the paper presents these as better correlated with human judgment for morphologically rich languages.","IndicTrans2, as presented, provides an accessible translation baseline for all 22 scheduled languages, with compact distilled variants that are deployable in limited-resource settings.","The gaps the paper names—code-mixed text, MRC, summarization, and TTS dialect variation—are the likely next bottlenecks for Marathi NLP."],"supporting_citations":[{"why":"Provides the EMILLE corpus, the early Marathi monolingual and English-Marathi parallel text resource the survey cites as a starting point.","marker":"Baker et al., 2004"},{"why":"Supplies the large web-crawled IndicNLP corpus with 142 million Marathi words that anchors the survey's corpus inventory.","marker":"Kunchukuttan et al., 2020"},{"why":"Gives Samanantar, the large English-Indic parallel corpus including Marathi that the survey presents as a key translation resource.","marker":"Ramesh et al., 2022"},{"why":"Provides IndicNLPSuite, IndicBERT, and the IndicGLUE benchmark used to characterize Marathi NLU evaluation.","marker":"Kakwani et al., 2020"},{"why":"Supplies IndicBART, the compact multilingual NLG model for Marathi and 10 other Indic languages that the survey compares with mBART50.","marker":"Dabre et al., 2022"},{"why":"Provides IndicTrans2 and the BPCC parallel corpus covering all 22 scheduled languages, including Marathi, as the survey's translation state of the art.","marker":"Gala et al., 2023"},{"why":"Gives MahaCorpus and the MahaBERT, MahaRoBERTa, and MahaGPT Marathi monolingual models that the survey cites for sentiment analysis, NER, and classification.","marker":"Joshi, 2022"},{"why":"Supplies the mahaNLP Marathi NLP library that the survey describes as covering preprocessing and advanced tasks.","marker":"Magdum et al., 2023"},{"why":"Supplies NLLB and FLORES-200, the largest aligned English-Marathi corpora and the state-of-the-art translation baseline referenced in the survey.","marker":"Costa-jussà et al., 2022"},{"why":"Is the prior Marathi NLP resource survey that this paper positions itself as extending with deep-learning models and evaluation metrics.","marker":"Lahoti et al., 2022"}],"fun_headline_variants":["Marathi NLP: From scarce resources to 752M-token models","Low-resource to large-scale: Marathi NLP's evolution","Marathi NLP survey: Big corpora, strong baselines, still gaps","752M tokens later: Marathi NLP comes of age","Marathi NLP: A review of models, data, and metrics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's overall picture of Marathi NLP depends on the assumption that the corpora, models, and tools it chose to include are representative of the field, since it does not state a search protocol, inclusion criterion, or cutoff date.","fun_headline_variants_meta":{"raw":{"variants":["Marathi NLP: From scarce resources to 752M-token models","Low-resource to large-scale: Marathi NLP's evolution","Marathi NLP survey: Big corpora, strong baselines, still gaps","752M tokens later: Marathi NLP comes of age","Marathi NLP: A review of models, data, and metrics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000665,"raw_usage":{"total_tokens":3026,"prompt_tokens":923,"completion_tokens":2103,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":2022}},"tokens_in":539,"tokens_out":2103,"duration_ms":11877,"temperature":1.0,"reasoning_tokens":2022,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:23:29.484955+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the survey's resource inventory against a systematic, time-bounded search of Marathi NLP publications and public model hubs; if major recent datasets or models are absent, or if reported corpus sizes and benchmark numbers do not match the cited releases, the 'state-of-the-art' framing fails.","supporting_citations":[],"review_version":1}