{"id":"f5505c46-9b05-4d62-8af9-ba46ca963066","arxiv_id":"2507.13390","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A technical report on a 2.9B English-Hindi model whose headline evaluation numbers are internally inconsistent and whose promoted tokenizer was not used to train the final model.","lead":"This paper describes PARAM-1, a 2.9 billion parameter language model trained from scratch on English and Hindi data. The authors claim it is a strong India-centric baseline, but the report's own tables and prose contradict each other on several key results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 8.2's SOTA claim is contradicted by Table 4: PARAM-1 MILU Hi/En are 30.17/36.3 vs Qwen-3B 33.6/49.84, and MMLU-Hindi few-shot 36.1 vs Sarvam-1 41.4; prose values do not appear in any table.","rationale":"The reader correctly identified the reported benchmark numbers as the weakest assumption. My independent reading finds the same point is the central failure: the empirical basis for the abstract's claim is self-contradictory in the exact metrics that define 'India-centric SOTA'. This is not a stylistic or minor reporting issue; the load-bearing comparison numbers are irreconcilable. I agree with the reader's REJECT verdict. No outside-consensus attack is needed, and the critique is based on internal inconsistency rather than any external expectation about model capabilities.","tokens_in":22398,"tokens_out":3763,"duration_ms":37158,"concrete_test":"Re-run the exact MILU (en/hi) and MMLU-Hindi evaluations with the released PARAM-1 checkpoint and the same harness/prompts used for Table 4, then check whether the resulting numbers match Table 4 or Section 8.2. Because no checkpoint is released, an equally decisive analytical check is to reproduce Table 4 for Qwen-3B and Sarvam-1 on the identical prompts; if Qwen-3B MILU-En approximates 49.8 and Sarvam-1 MMLU-Hindi approximates 41.4, then PARAM-1's prose claim of 49.7/36.1 is not SOTA and the table/prose contradiction is resolved in favor of the table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires the benchmark numbers in Sections 8.1-8.2 to be accurate. They are not internally consistent, and the contradiction is load-bearing because the India-centric SOTA claim rests on prose values that no table contains. In Table 4, MILU (Hi) is 30.17 and MILU (En) is 36.3 for PARAM-1, while Qwen-3B scores 33.6 and 49.84; MMLU-Hindi few-shot is 36.1 for PARAM-1, while Sarvam-1 scores 41.4 and Llama-3.2 scores 37.5. Section 8.2 instead states MILU 48.3 (Hi) and 49.7 (En) and MMLU-Hindi 36.1 'improving over SARVAM-1 (33.4)', a Sarvam number that appears nowhere. Section 8.1 similarly claims ARC-Challenge few-shot 52.9 outperforms Sarvam-1 44.8 and Qwen-3B 50.4, but Table 3 lists Qwen-3B 57.08, Sarvam-1 54.4, and Granite-2B 64.16. If Table 4 is authoritative, PARAM-1 is not state-of-the-art on MILU or MMLU-Hindi; if the prose is authoritative, the tables are wrong and no reliable evidence is left. The additional unsupported claims (Nemotron tokenizer used instead of the advertised tokenizer, 5T corpus vs 'tens of billions' in Section 6.1) reinforce that the empirical record cannot be trusted, but the benchmark contradiction alone defeats the claimed SOTA result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PARAM-1, a 2.9B-parameter English-Hindi decoder-only transformer pretrained from scratch on a claimed 5T-token corpus with a 25% Hindi allocation, a custom tokenizer (Section 3), three progressive pretraining phases, and bilingual SFT. It reports English and Indic benchmark results in Tables 3 and 4, and Section 8.2 claims state-of-the-art performance among open models on MILU, MMLU-Hindi, and SANSKRITI. The paper also reports tokenizer fertility comparisons, Prometheus-based generation quality, and toxicity evaluations.","tokens_in":22813,"tokens_out":10978,"duration_ms":100329,"significance":"If the reported results were reliable, the paper would offer a useful design point for India-centric bilingual modeling: a relatively small dense model with a deliberate 25% Hindi data allocation and fertility-aware tokenization, compared directly against several open baselines. The tokenizer fertility comparison in Table 1 is a useful artifact. However, the manuscript is dominated by internal contradictions between prose claims and tables, and no code, model weights, or evaluation harness are released. As submitted, the central claims of competence and SOTA performance are not supported by a consistent empirical record.","major_comments":[{"comment":"Section 8.2 claims PARAM-1 scores 48.3% (Hindi) and 49.7% (English) on MILU, surpassing SARVAM-1 by approximately 6 points in both languages, but Table 4 lists PARAM-1 MILU scores of 30.17 (Hi) and 36.3 (En), with SARVAM-1 at 28.48 (Hi) and 32.12 (En) and QWEN-3B at 33.6 (Hi) and 49.84 (En). According to the only tabulated numbers, PARAM-1 is not state-of-the-art on MILU, and the claimed ~6-point advantage is not supported. This directly contradicts the central 'state-of-the-art among open models' statement.","section":"§8.2, Table 4"},{"comment":"Section 8.1 states that PARAM-1 achieves 52.9% (few-shot) on ARC-Challenge, outperforming SARVAM-1 (44.8%) and QWEN-2.5 3B (50.4%), but Table 3 reports SARVAM-1 at 54.4% and QWEN-3B at 57.08% for this column. The same paragraph claims MMLU few-shot 35.2% versus SARVAM-1 28.7%, while Table 3 gives 46.0% for PARAM-1 and 47.7% for SARVAM-1. The prose comparisons and table entries disagree on both the model's own score and the baseline scores, so the reader cannot determine which set of numbers constitutes the reported result.","section":"§8.1, Table 3"},{"comment":"Section 8.3 says that in English, PARAM-1 achieves an overall score of 3.259, 'slightly ahead of' SARVAM-1 (3.318). The numeric value 3.259 is lower than 3.318, so this sentence is internally inconsistent and contradicts the surrounding claim of consistent gains over SARVAM-1. If the intended ordering was the opposite, the text should state it explicitly and use the correct scores.","section":"§8.3"},{"comment":"Section 3 describes a custom SentencePiece BPE tokenizer (BharatGen-128K v1) and presents its fertility results in Table 1, but the final paragraph of that section states that 'PARAM-1 was trained using the Nemotron tokenizer [31].' Section 6.2, in contrast, says the NeMo tokenizer customization 'helped us to use or inhouse multilingual tokenizer.' These statements are mutually exclusive. If the Nemotron tokenizer was used, Section 3 and Table 1 do not describe the model's tokenizer and the tokenizer-fairness claim is unfounded; if the in-house tokenizer was used, the note in Section 3 is false. This is a load-bearing inconsistency for the paper's tokenization-fairness contribution.","section":"§3, §6.2"},{"comment":"Section 2.1 and Section 5.1.1 describe a 5 trillion-token multilingual corpus (3.48T English + 1.52T Hindi) for pretraining, while Section 6.1 says the training ran 'over tens of billions of tokens using hundreds of H100 GPUs.' These statements differ by roughly two orders of magnitude and cannot both be true. The additional Phase-2 (2T tokens) and Phase-3 (500B tokens) corpora further compound the discrepancy. Without a consistent statement of the actual training data volume, the training narrative is not reproducible.","section":"§2.1, §5.1.1, §6.1"},{"comment":"In Table 4, the HellaSwag-Hindi scores for PARAM-1 are 71.4 (zero-shot) and 73.4 (few-shot), which are exactly identical to the English HellaSwag scores reported for PARAM-1 in Table 3. In contrast, every other baseline shows a substantial drop when moving from English to Hindi HellaSwag (e.g., QWEN-3B from 73.6 to 32.9; SARVAM-1 from 66.9 to 42.9). The identical values strongly suggest the Hindi HellaSwag evaluation was not performed separately or was misreported, which materially weakens the claimed cross-lingual evaluation evidence.","section":"Table 4"}],"minor_comments":[{"comment":"The text says 'LogiQA [16]' but reference [16] is the TriviaQA paper; a correct citation for LogiQA is needed.","section":"Section 7"},{"comment":"The description of LLAMA-3.2-3B as 'a pruned and distilled variant of the LLaMA 3.1 70B variant' is inaccurate for the dense 3B model; this characterization should be corrected or removed.","section":"Section 6.3"},{"comment":"Tables 3 and 4 report single evaluation runs without standard errors, confidence intervals, or any statement of the number of repetitions; for claimed improvements of 1–3 points (e.g., SANSKRITI in Table 4), this is insufficient to establish that the differences are meaningful.","section":"Tables 3 and 4"},{"comment":"The introduction promises evaluations on IndicQA, code-mixed reasoning, and socio-linguistic robustness tasks, but no results for these tasks are reported in Section 8; either add the results or adjust the introduction to describe what is actually evaluated.","section":"Introduction and Section 8"},{"comment":"Section 8.2 claims that PARAM-1 'outperforms SARVAM-1 in all four question types' on SANSKRITI, but Table 4 reports only an overall SANSKRITI score and no per-type breakdown, so this claim is unverifiable from the presented results.","section":"Table 4 and Section 8.2"},{"comment":"The Hindi HellaSwag evaluation is referenced as 'its Hindi adaptation' without citing a dataset or describing the adaptation procedure, making the results difficult to reproduce.","section":"Section 7"}],"recommendation":"reject","confidential_remarks":"The number and severity of internal inconsistencies go beyond ordinary presentation flaws: the MILU and MMLU-Hindi prose values directly contradict Table 4, the HellaSwag-Hindi scores in Table 4 are identical to the English HellaSwag scores, and Section 3 explicitly disclaims the tokenizer that the rest of the paper describes as the model's own. The claimed 5T-token corpus for a 2.9B model is also computationally implausible given the described 512-GPU cluster and the 'tens of billions of tokens' statement in Section 6.1. I recommend that the editor treat the reported benchmark tables as unverified until the authors provide a single consistent, versioned set of experimental results and release or specify the evaluation harness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Step aside from the marketing: the paper's central empirical claim does not survive contact with its own tables. Section 8.2 says PARAM-1 scores 48.3/49.7 on MILU Hi/En and beats Sarvam-1 on MMLU-Hindi. Table 4 gives 30.17/36.3 for MILU, with Qwen-3B at 33.6/49.84, and MMLU-Hindi 36.1 for PARAM-1 versus Sarvam-1's 41.4. The ARC-Challenge prose (52.9, 'outperforming' Sarvam's 44.8 and Qwen's 50.4) is similarly at odds with Table 3, where Sarvam sits at 54.4 and Qwen at 57.08. You cannot pick one version and still have a SOTA claim.\n\nWhat is genuinely there: a sensible design question—how much Indic data in pretraining actually buys you at the 2.9B scale—and a fertility-aware tokenizer mixture algorithm that is clearly described, even if the final model does not use that tokenizer (Section 3's last paragraph says the Nemotron tokenizer was used instead). The choice of MILU, MMLU-Hindi, and SANSKRITI as evaluation targets is appropriate, and the training recipe is unusually concrete about data mixtures and phases.\n\nThe soft spots are load-bearing, not cosmetic. Beyond the benchmark contradictions, Section 5.1.1 claims a 5-trillion-token corpus while Section 6.1 says training ran on 'tens of billions of tokens'—a five-order discrepancy that a typo does not explain. The Prometheus prose has PARAM-1's English score (3.259) 'ahead' of Sarvam-1 (3.318) when it is behind. No model weights, tokenizer, or data are released, so nothing can be checked independently. The absence of error bars is minor by comparison.\n\nWhere does this leave us? If Table 4 is authoritative, the paper's main conclusion is false. If the prose is authoritative, the tables are wrong and no reliable evidence remains. Either way, the manuscript as submitted is not a reliable report. The idea is worth exploring, but this version doesn't give us a verified artifact or consistent results.\n\nRecommendation: desk reject, with an invitation to resubmit if the authors can release the actual model and tokenizer, audit every number against its table, and reconcile the corpus-scale claim. I would not spend referee time on the current version.","headline":"A 2.9B Hindi-English model whose headline SOTA claims are contradicted by its own tables, and whose advertised tokenizer was not used; not ready for review.","tokens_in":23446,"tokens_out":4709,"would_cite":false,"duration_ms":49238,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a 2.9-billion-parameter model pretrained on a corpus that is 25% Hindi plus English can serve both as a competent general-purpose model and as a state-of-the-art open baseline for India-centric benchmarks…","keywords":["PARAM-1","Hindi-English bilingual model","India-centric LLM","25% Hindi pretraining allocation","MILU benchmark","SANSKRITI benchmark","tokenizer fertility","bilingual instruction tuning"],"falsifier":"Run the released PARAM-1 checkpoint on the MILU Hindi and English subsets and on MMLU-Hindi under the paper's stated prompts and decoding settings. The central claim is falsified if the reproduced scores land near Table 4's values (MILU 30.17/36.3) instead of the Section 8.2 prose values (48.3/49.7), or if the margin over Sarvam-1 disappears. A second check is to audit the released data recipe and training logs to see whether the corpus was 5 trillion tokens as in Section 5.1.1 or tens of billions as in Section 6.1.","tokens_in":22191,"feed_emoji":"🇮🇳","tokens_out":10003,"duration_ms":94332,"temperature":0.7,"pith_summary":"PARAM-1 is a 2.9-billion-parameter, decoder-only, text-only language model trained from scratch on a bilingual corpus of Hindi and English, with a deliberate 25% Hindi allocation, and then instruction-tuned on a curated bilingual dataset. The paper's central claim is that this design yields a model that is both a competent general-purpose model and a robust baseline for India-centric applications, achieving state-of-the-art results among open models on MILU, MMLU-Hindi, and SANSKRITI and surpassing the Indic-focused Sarvam-1 on several Indian benchmarks. The authors argue this shows linguistic diversity should be embedded at pretraining level through data, tokenization, and evaluation, rather than deferred to post-hoc fine-tuning. A sympathetic reader would care because, if the claim holds, it offers a concrete recipe for small models to serve linguistically diverse regions efficiently.","feed_headline":"2.9B Hindi-English model tops Indic benchmarks","feed_subtitle":"A 25% Hindi pretraining share plus bilingual instruction tuning is claimed to beat open rivals on MILU, MMLU-Hindi, SANSKRITI.","key_machinery":"The load-bearing mechanism is the data mixture and training curriculum rather than a new architecture. The paper specifies three design levers: a 25% Hindi corpus allocation (1.52 trillion of 5 trillion tokens), a three-phase pretraining schedule—bootstrap on the full corpus, factual-preservation on a 2-trillion-token bilingual corpus with 30% parallel data, and long-context adaptation on a 500-billion-token corpus with many documents above 2,048 tokens—and bilingual instruction tuning on two SFT sets (roughly 1 million and 473,000 pairs) filtered by a strict scoring rubric. The tokenizer described for tokenization fairness is a SentencePiece BPE with a 128K vocabulary and byte fallback, although a footnote in Section 3 states that PARAM-1 itself was trained with the Nemotron tokenizer rather than the in-house one described there.","core_discovery":"PARAM-1 is presented as a 2.9-billion-parameter, decoder-only, text-only transformer trained from scratch on a bilingual corpus of Hindi and English, with 1.52 trillion Hindi tokens out of a total 5 trillion, and then instruction-tuned on curated bilingual datasets. The paper's central claim is that this recipe yields a model that is simultaneously a competent general-purpose model and a robust baseline for India-centric applications, achieving state-of-the-art results among open models on MILU, MMLU-Hindi, and SANSKRITI and outperforming the Indic-focused Sarvam-1 by roughly 6 points on MILU in both Hindi and English. The authors present this as evidence that linguistic diversity should be built into pretraining through corpus allocation, tokenization, and culturally aligned evaluation, rather than added later through fine-tuning.","pith_inferences":["The paper reports MILU prose scores of 48.3/49.7 and table scores of 30.17/36.3; I infer that the headline 'state-of-the-art' claim is provisional until the released checkpoint reproduces one set of numbers.","The comparison against Sarvam-1, which was trained on multiple Indic languages, hints that depth in one major language may transfer better than breadth across many, but the paper runs no ablation that isolates this factor.","A testable extension is to train a sibling model with 12.5% Hindi and 12.5% of a second Indic language, to separate the effect of Hindi depth from the effect of the overall 25% Indic allocation.","The discrepancy between the 5-trillion-token pretraining claim and the 'tens of billions' figure in the infrastructure section could be settled by auditing the released data recipe and training logs."],"forward_implications":["A 2.9B model with deep Hindi representation could serve as an efficient open baseline for Indian knowledge and cultural tasks, reducing the compute needed for India-centric deployment.","The 25% Hindi allocation becomes a concrete design target for future Indic models, implying that broad multilinguality is not required for strong India-focused performance.","The bilingual instruction-tuning pipeline, filtered to a perfect rubric score, would be a reusable recipe for culturally aligned alignment.","If the reported results hold, they suggest that pretraining-stage representation, rather than post-hoc fine-tuning, is the place to secure performance in underrepresented languages."],"supporting_citations":[{"why":"Supplies the primary India-centric benchmark (MILU) on which the paper claims state-of-the-art results and defines the Hindi and English subsets used for comparison.","marker":"[43]"},{"why":"The cultural knowledge benchmark where PARAM-1 claims to outperform Sarvam-1 across all four question types.","marker":"[29]"},{"why":"Source of MMLU-Hindi, the Hindi-translated multitask benchmark used to claim improvement over Sarvam-1 and LLaMA-3.2.","marker":"[15]"},{"why":"Provides the Sangraha Hindi web-scale corpus that makes up part of the 1.52 trillion Hindi token allocation.","marker":"[19]"},{"why":"One of the high-quality English corpora forming the 3.48 trillion English portion of the pretraining mixture.","marker":"[39]"},{"why":"Source of roughly 207,000 English-Hindi instruction pairs used in the instruction-tuning pool.","marker":"[22]"},{"why":"Baseline model (LLaMA-3.2 3B) and the cited evidence that mainstream models allocate about 0.01% of training data to Indic languages.","marker":"[13]"},{"why":"Baseline model (Qwen-3B) compared across English and Indic benchmarks.","marker":"[33]"}],"fun_headline_variants":["PARAM-1: 2.9B bilingual Hindi-English model beats open rivals on Indic benchmarks","25% Hindi pretraining share yields SOTA for India-centric AI","PARAM-1 beats Sarvam-1 by 6 points on MILU","Bilingual from scratch: PARAM-1 tops MILU, MMLU-Hindi"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the accuracy of the paper's reported benchmark numbers and training-data description; the paper itself contains discrepancies—MILU is given as 48.3/49.7 in the prose and 30.17/36.3 in Table 4, ARC-Challenge few-shot is 52.9% in prose but 54.4% in the table, and the pretraining corpus is described in one section as 5 trillion tokens and in another as tens of billions—so if those numbers do not reproduce, the claimed state-of-the-art India-centric performance collapses.","fun_headline_variants_meta":{"raw":{"variants":["PARAM-1: 2.9B bilingual Hindi-English model beats open rivals on Indic benchmarks","25% Hindi pretraining share yields SOTA for India-centric AI","PARAM-1 beats Sarvam-1 by 6 points on MILU","Bilingual from scratch: PARAM-1 tops MILU, MMLU-Hindi"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000925,"raw_usage":{"total_tokens":3964,"prompt_tokens":947,"completion_tokens":3017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2926}},"tokens_in":563,"tokens_out":3017,"duration_ms":25945,"temperature":1.0,"reasoning_tokens":2926,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:58:03.286761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the released PARAM-1 checkpoint on the MILU Hindi and English subsets and on MMLU-Hindi under the paper's stated prompts and decoding settings. The central claim is falsified if the reproduced scores land near Table 4's values (MILU 30.17/36.3) instead of the Section 8.2 prose values (48.3/49.7), or if the margin over Sarvam-1 disappears. A second check is to audit the released data recipe and training logs to see whether the corpus was 5 trillion tokens as in Section 5.1.1 or tens of billions as in Section 6.1.","supporting_citations":[{"cited_title":"Sanskriti: A comprehensive benchmark for evaluating language models’ knowledge of indian culture, 2025","cited_arxiv_id":null,"evidence_quote":"The cultural knowledge benchmark where PARAM-1 claims to outperform Sarvam-1 across all four question types."},{"cited_title":"Mi- randa, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D","cited_arxiv_id":null,"evidence_quote":"Source of roughly 207,000 English-Hindi instruction pairs used in the instruction-tuning pool."}],"review_version":1}