{"id":"e34a9d9d-12e5-41e7-835f-7dd5cd798e87","arxiv_id":"2607.25375","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Inspect India Evals provides six India-specific LLM benchmarks and preliminary scores for five open-weight models, with most performance gaps not statistically significant at n=5.","lead":"This paper introduces an open benchmark suite that tests AI chatbots on Indian languages, social bias, digital-payment safety, and cultural knowledge. It reports that Indian-tuned and top general models score highest, but cautions that only five test questions per task were used, so the rankings are preliminary.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cultural-knowledge ranking ('Sarvam beats 32B') rests on Llama-3.1-8B-as-judge scores with no validation and only n=5 (p=0.222); the abstract states this as fact despite the paper's own caution.","rationale":"The paper's contribution—an open framework and pilot results—is real and reproducible, and the authors explicitly call the numbers indicative. But the abstract and conclusion present a definitive ranking, and the least-secure condition for that ranking is the measurement chain: translated items, two-layer scorer, and Llama-3.1-8B rubric judge. Since cultural knowledge is excluded from the IFI (§3.9), the judge bias does not affect the composite index, but it does affect the separate 'beats larger 32B on cultural knowledge' claim. I agree with the reader's item-validity concern, but I narrow it to the judge-based module because that is the most concrete, testable weak link. The correct verdict remains CONDITIONAL: the framework is acceptable as a pilot, but the headline should not be used for deployment decisions until the judge is validated and samples are larger.","tokens_in":13189,"tokens_out":6225,"duration_ms":64132,"concrete_test":"Run the cultural-knowledge module on the full N=300 (or at least n≥50) and independently score the same model outputs with (a) two human annotators using the published rubric and (b) a second LLM judge (e.g., GPT-4o or Qwen 2.5 72B). Compute judge–human agreement (e.g., Cohen's κ on binarized scores) and compare the Sarvam-M 24B vs Qwen 2.5 32B ranking. Also spot-check 50 random Multilingual MMLU/BharatBBQ items by back-translation and native-speaker ratings. If the cultural gap disappears or agreement is low, the abstract's ranking claim must be downgraded to preliminary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central superiority claim for cultural knowledge depends on Module 6 (§3.7), where model answers are graded by Llama 3.1 8B as LLM-as-judge from a four-criterion rubric. This is not validated: the paper reports no human-annotator agreement, no second-judge comparison, and no example items; §5.8 explicitly concedes 'a single judge model can also have its own blind spots baked in.' With limit=5 samples per model (§4.2), Table 2 gives Sarvam-M 24B vs DeepSeek-R1 14B as 3/5 vs 0/5, p=0.222; the reported Sarvam-vs-Qwen 60% vs 30% gap is not tested, and 30% is not representable as a count out of 5, suggesting a scoring/threshold inconsistency. Because the same Llama-3.1-8B judge scores 20% on this benchmark and is itself one of the evaluated models, the observed advantage could reflect judge preferences rather than cultural knowledge. A similar unvalidated translation pipeline (§3.2-3.3) threatens the MMLU and BharatBBQ scores, but the cultural-knowledge result is the specific load-bearing support for the 'beats larger 32B' half of the headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Inspect India Evals, an open-source evaluation framework built on the UK AISI Inspect AI platform, comprising six benchmarks: Multilingual MMLU across 16 languages, BharatBBQ bias probes, DPI safety, multilingual safety refusal, jailbreak resistance, and Indian cultural knowledge scored by LLM-as-judge. Five open-weight models (Llama 3.1 8B, DeepSeek-R1 14B, Sarvam-M 24B, Gemma 2 27B, Qwen 2.5 32B) are evaluated with 5 samples per task-model, and results feed an India Fairness Index (IFI). The paper claims Sarvam-M and Gemma 2 are top, with Sarvam-M beating larger 32B models on cultural knowledge and DPI safety.","tokens_in":13601,"tokens_out":7747,"duration_ms":66541,"significance":"If the evaluation items are valid, the framework addresses a genuine gap: India-specific multilingual safety and fairness measurement, with a reproducible, open pipeline that does not require commercial API access. The IFI is transparently defined. However, the empirical findings are not yet reliable: sample sizes are too small, the cultural-knowledge judge is an evaluated model with no human validation, translation quality is not demonstrated, and several reporting inconsistencies appear. The paper's contribution is best read as a framework plus a preliminary pilot, not as a definitive model ranking.","major_comments":[{"comment":"§4.2 states 'limit=5 samples per dataset module per model,' giving 150 total runs. Yet Table 2 reports per-language percentages for 16 languages for each of five models; e.g., DeepSeek-R1 14B shows 66.7% on Assamese and 100% on Bengali and several other languages, which requires far more than 5 total MMLU samples per model. Either the sampling protocol is misdescribed or the table is not derived from the stated protocol. This makes the headline Multilingual MMLU ranking (DeepSeek 80.6%) impossible to interpret and undermines the reproducibility claim.","section":"§4.2 and Table 2"},{"comment":"Indian Cultural Knowledge is scored with Llama 3.1 8B as LLM-as-judge; §5.8 concedes 'a single judge model can also have its own blind spots baked in.' Since Llama 3.1 8B is itself one of the five evaluated models (scoring 20% on the same benchmark), the relative ranking, especially Sarvam-M 24B at 60% vs DeepSeek-R1 14B at 10%, may reflect judge bias or rubric interpretation rather than cultural knowledge. No human-annotation agreement, second-judge comparison, or example items are reported. This is the load-bearing support for the abstract's 'beating larger 32B models on Indian cultural knowledge' claim.","section":"§3.7, §5.8"},{"comment":"For cultural_knowledge, Sarvam-M (3/5) vs DeepSeek-R1 (0/5) is reported with p=0.222. Fisher's exact test on these counts gives a two-sided p of 0.167 (one-sided p=0.083), so the reported p-value is incorrect. More broadly, only 1 of 5 tested tasks reaches p<0.05, and §5.8 acknowledges the scale limitation; nevertheless, the Abstract and §5.1 state definitive rankings ('came out on top', 'beating larger 32B models'). The conclusions should be explicitly limited to preliminary directional evidence.","section":"§5.4 and Table 2"},{"comment":"The Multilingual MMLU and BharatBBQ datasets were 'developed with reference to native speaker consultation,' but no translation quality checks, back-translation, pilot validation, inter-annotator agreement, or sample items are provided. Without this, language-specific accuracy and bias scores may reflect translation artifacts or cultural misjudgements rather than model capability. This is particularly important because the per-language MMLU percentages in Table 2 are the basis for claims about language-specific weaknesses.","section":"§3.2–3.3"},{"comment":"The abstract states that both Sarvam-M 24B and Gemma 2 27B score 80% on the composite India Fairness Index, but Table 3 reports IFI values of 76.6% and 83.1%, respectively. Section 3.9 excludes Cultural Knowledge from IFI, but Table 1's 'Overall Mean' includes it; the abstract and §5.1 conflate these two different averages. This discrepancy in the headline number should be corrected.","section":"Abstract vs §5.7/Table 3"}],"minor_comments":[{"comment":"The section numbering skips from 3.7 to 3.9; there is no §3.8.","section":"Section numbering"},{"comment":"The text says '50 harmful prompts' translated into English, Hindi, Tamil, Telugu, and Bengali (250 prompts) but later says 'All 200 prompts are classified as High risk.' Clarify the actual count.","section":"§3.5"},{"comment":"The caption notes 'Results are directional; re-evaluation at n ≥ 50 is recommended,' but the p-values are labeled two-sided and the cultural_knowledge p-value is arithmetically incorrect (see major comment).","section":"Table 2 caption"},{"comment":"The abstract says 'sixteen Indian languages,' but the list in §3.2 includes English as a baseline; consider wording such as '16 languages including English.'","section":"Language naming"},{"comment":"Figures lack proper captions and numbering (e.g., 'Figure- Illustrates...' and 'Figure: Performance Heatmap Matrix'). Reference them consistently in the text.","section":"Figures"},{"comment":"Reference [13] cites 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding' for mBERT, but the mBERT model is not described in that paper; cite the appropriate multilingual model release.","section":"References"},{"comment":"The phrase 'full-scale corpus (N=825+)' is undefined; specify what constitutes a sample and how the total is computed.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The framework is potentially valuable, but the current empirical section is not ready for publication as-is. The authors should be encouraged to either run a larger evaluation (n ≥ 50 per task-model), validate the cultural-knowledge judge against human ratings, provide translation quality evidence, and correct the statistical and reporting errors. If they cannot obtain more data, they should reframe the paper as a framework+preliminary pilot and remove the definitive ranking claims from the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know about this paper because it's one of the first open benchmarks aimed squarely at India-specific LLM safety, with code and data actually shipped. The new axes – DPI safety (Aadhaar/UPI/Bhashini misuse vs. over-refusal) and BharatBBQ's caste/region categories – fill a real gap. The framework is sensible: two-layer scoring with deterministic matching before LLM-as-judge, and IFI deliberately excludes cultural knowledge to stay a safety/fairness composite. That's good design.\n\nBut the empirical results as reported are not trustworthy enough to carry the headline. The abstract says Sarvam-M and Gemma both score 80% on IFI, while Table 3 gives 83.1% and 76.6% – a direct inconsistency. More serious: §4.2 says limit=5 samples per task per model, 150 runs total, but the multilingual MMLU table lists per-language percentages across 16 languages. You cannot get those numbers from 5 samples per model. Either the protocol description is wrong or the table comes from a different run. Either way, the MMLU results are uninterpretable as reported.\n\nThe cultural-knowledge claim – Sarvam beats larger 32B models – rests on Llama-3.1-8B-as-judge, which is itself one of the evaluated models, with no human agreement or second judge check. The paper's own discussion admits 'a single judge model can also have its own blind spots baked in.' With n=5, the difference is p=0.222, not significant. So the 'beats 32B' part of the abstract is not supported by the evidence.\n\nOther soft spots: no translation quality checks or example items for the translated MMLU and BharatBBQ, so dataset validity is assumed. The authors do candidly say these are preliminary and recommend n≥50. That honesty counts for something.\n\nBottom line: the framework and datasets are worth engaging with; the numbers are not. This is a paper for a serious referee, but the authors need to fix the counting inconsistency, validate the judge, and re-run at larger n before the empirical rankings are used. If I were editor, I'd send it out with a request for major revision rather than desk reject it.","headline":"Useful open benchmark for Indian LLM safety, but the empirical results are not reliable enough to support the headline claims; the framework and datasets deserve attention, the numbers don't.","tokens_in":14024,"tokens_out":4094,"would_cite":false,"duration_ms":39789,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new open-source framework with six India-specific benchmarks tests multilingual accuracy, caste and regional bias, Digital Public Infrastructure safety, multilingual refusal, jailbreak resistance, and cultural knowledge, reporting that Sa","keywords":["LLM evaluation","Indian languages","multilingual benchmarks","social bias","caste","DPI safety","jailbreak","cultural knowledge"],"falsifier":"Take a random sample of the translated Multilingual MMLU items and have native speakers independently back-translate them to English; if the back-translation changes the correct answer for more than a small fraction (say >5%) of items, then the reported language-wise accuracies do not measure what the paper claims.","tokens_in":13117,"feed_emoji":"🇮🇳","tokens_out":11933,"duration_ms":101630,"temperature":0.7,"pith_summary":"The paper argues that standard English- and Western-centric benchmarks do not surface the safety and fairness failures that matter when large language models are deployed across India's 22 languages and 1.4 billion people. To close this gap, it introduces Inspect India Evals, an open framework of six benchmarks covering multilingual factual accuracy in 16 Indian languages, Indian social bias including caste, safety around Digital Public Infrastructure (Aadhaar, UPI, Bhashini), multilingual safety refusal, multi-turn jailbreak resistance, and rubric-scored Indian cultural knowledge. Testing five open-weight models, the paper claims that Sarvam-M 24B and Gemma 2 27B come out on top with roughly 80% on the composite India Fairness Index, and that Sarvam-M beats larger 32B models on cultural knowledge and DPI safety. The paper also finds a sharp trade-off: the reasoning-focused DeepSeek-R1 14B scores highest on multilingual MMLU but lowest on DPI safety and cultural knowledge, showing that strong reasoning does not imply safe or context-aware behavior in Indian deployments.","feed_headline":"Sarvam-M and Gemma 2 top new India-focused LLM safety evals","feed_subtitle":"Indic-tuned model beats larger rivals on cultural knowledge and Aadhaar/UPI safety; reasoning alone isn't enough.","key_machinery":"The load-bearing mechanism is the set of six benchmark modules plus the composite India Fairness Index (IFI), which averages four normalized sub-scores: multilingual MMLU accuracy, BharatBBQ unbiased accuracy, multilingual safety refusal rate, and DPI safety compliance. The BharatBBQ module adapts the US-centered bias methodology to 13 Indian demographic axes, using ambiguous versus disambiguated question pairs to compute a stereotype-consistent error rate; the DPI module classifies queries by risk level and expected behaviour to separate under-refusal (complying with fraud assistance) from over-refusal (refusing legitimate Aadhaar/UPI guidance); and the cultural knowledge module uses rubric","core_discovery":"The central claim is that India-specific evaluation is both necessary and feasible, and that the proposed six-benchmark framework measures dimensions of LLM behaviour that standard benchmarks miss. Using this framework, the authors report that two open-weight models—Sarvam-M 24B and Gemma 2 27B—are the safest and most accurate for Indian deployment, each reaching about 80% on the composite India Fairness Index. Sarvam-M 24B outperforms even larger 32B models on Indian cultural knowledge (60% vs 30% for Qwen 2.5 32B) and on DPI safety compliance (100%), while Gemma 2 27B shows the highest unbiased accuracy on the Indian bias benchmark (100%). All models refused direct harmful prompts in India","pith_inferences":["A direct extension, which the paper itself recommends, is to run the suite at N≥50 per task; only DPI safety reached statistical significance at N=5, so a larger run would confirm or overturn the reported ranking.","Because cultural knowledge scores rely on a single LLM judge (Llama 3.1 8B), an immediate testable extension is to score the same responses with a second judge model or human raters to measure judge-dependent variance.","The framework's modular design could be transferred to other multilingual, non-Western societies by swapping the demographic axes and the digital-infrastructure module, for example for African or Southeast Asian languages.","The observed vulnerability of the chain-of-thought reasoning model (DeepSeek-R1) to jailbreaks suggests a testable hypothesis: other reasoning-specialized models should show similar drops in jailbreak resistance compared to instruction-tuned models of the same scale."],"forward_implications":["If the framework is valid, choosing models for Indian public-sector and consumer deployments should weigh India-specific safety and bias scores as heavily as general reasoning scores; a model like DeepSeek-R1 that leads multilingual MMLU but fails DPI safety would not be suitable for Aadhaar/UPI-facing chatbots.","Specialized Indic fine-tuning appears to pay off: Sarvam-M 24B, a model tuned for Indian languages, beats general-purpose 32B models on cultural knowledge and DPI safety, supporting further investment in Indic-language alignment rather than just parameter scaling.","The 100% refusal on direct harmful prompts in five Indian languages is not sufficient: multi-turn jailbreak and domain-specific DPI tests reveal a wide range (40–80% and 20–100%), so deployment evaluations should include adversarial and domain-specific probes.","The India Fairness Index gives procurement bodies a single comparable safety/fairness score, which could standardize how Indian government agencies assess AI vendors."],"fun_headline_variants":["India-focused eval finds Sarvam-M and Gemma 2 safest for LLMs","New India benchmarks: Sarvam-M 24B tops cultural and safety tests","Sarvam-M beats 32B rivals on Indian cultural knowledge and safety","India-specific LLM eval: two open models lead on fairness and safety","Sarvam-M and Gemma 2 lead new India-centric LLM benchmarks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework's validity rests on the assumption that the newly created evaluation items—the translated MMLU questions and the BharatBBQ bias probes—are accurate and culturally correct; the paper reports no human validation of these items, so if translations are flawed or stereotypes mislabeled, the reported model differences would be artifacts of the dataset rather than real capability or safety gaps.","fun_headline_variants_meta":{"raw":{"variants":["India-focused eval finds Sarvam-M and Gemma 2 safest for LLMs","New India benchmarks: Sarvam-M 24B tops cultural and safety tests","Sarvam-M beats 32B rivals on Indian cultural knowledge and safety","India-specific LLM eval: two open models lead on fairness and safety","Sarvam-M and Gemma 2 lead new India-centric LLM benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000584,"raw_usage":{"total_tokens":2646,"prompt_tokens":872,"completion_tokens":1774,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":1684}},"tokens_in":616,"tokens_out":1774,"duration_ms":11176,"temperature":1.0,"reasoning_tokens":1684,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:35:35.478767+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the translated Multilingual MMLU items and have native speakers independently back-translate them to English; if the back-translation changes the correct answer for more than a small fraction (say >5%) of items, then the reported language-wise accuracies do not measure what the paper claims.","supporting_citations":[],"review_version":1}