{"id":"2d391dca-66e3-43fa-b8d6-0d8d2df3bfc2","arxiv_id":"2608.07862","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SurakshaEval, a 2,968-prompt safety benchmark in 10 Indic languages plus English, shows that 27 current LLMs pass safety checks far less often in Indic scripts than in English.","lead":"The paper introduces SurakshaEval, a new safety test set with 2,968 human-written prompts in ten Indian languages plus English, and evaluates 27 large language models against it. Generalist readers should care because it gives concrete evidence that today's multilingual models are substantially less safe in Indian languages and native scripts than in English.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Safety is operationalized as alignment with official government positions (§4.2 Q2–Q5), so every EN-vs-IN pass-rate claim inherits this construct-validity threat.","rationale":"The paper builds a genuinely useful resource: human-written prompts across ten Indic languages and English, a seven-category harm taxonomy, a large model sweep, and an honestly reported manual evaluation. The central claim, however, depends on what the questionnaire counts as 'safe.' The reader's weakest assumption identifies the same point I find most load-bearing: Generic Questions 2–5 in §4.2 define harmless in terms of official government approval, societal stability from the official side, suitability for officials to say, and absence of government policy violation. Since §4.3 requires every response to pass this generic gate, all headline numbers are conformity scores. The authors' own §5 statement that they plan to broaden the questions toward general ethical principles is not a peripheral future direction; it flags a defect in the present operationalization. My concrete test would settle whether this construct-validity threat changes the substantive conclusion: re-score a sample with harm-focused generic questions and compare the English-versus-Indic gap. If the gap persists, the benchmark's central finding is robust to the choice of safety definition; if it collapses, the paper's main takeaway is metric-dependent. I do not see a need to change the reader's CONDITIONAL verdict: the concern is real but addressable, and the benchmark's data can be re-analyzed under alternative rubrics without invalidating the resource itself.","tokens_in":32319,"tokens_out":3824,"duration_ms":44458,"concrete_test":"Re-score a stratified random sample of responses (e.g., 50 per Indic language subset plus 50 English responses, drawn from at least three models spanning the performance range) using a modified questionnaire that replaces Generic Questions 2–5 with neutral harm-focused items (e.g., 'Could this response cause physical or psychological harm, spread harmful misinformation, incite violence or discrimination, or violate privacy or consent?') while keeping the specific harm-type questions unchanged. Compare per-model pass rates and the English-versus-Indic gap before and after the substitution. If the cross-lingual gap shrinks materially or model rankings change, the headline 'safety gap' is partly an artifact of the official-position criterion; if the gap persists with the harm-focused rubric, the concern is not load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The safety metric is defined by asking whether a response would be viewed positively by, be suitable for, or risk violating the policies of, an official of the Central or State Government (§4.2, Generic Questions 2–5). This generic gate is applied to every response before specific harm-type checks, and §4.3 states that a response counts as safe only if it passes both generic and specific conditions. Consequently, every reported pass rate, and every English-versus-Indic comparison, encodes a conformity assumption: a response that accurately documents government misconduct, criticizes a state policy, or states a view a government official would disfavor is marked unsafe regardless of whether it causes harm; conversely, a response that echoes official positions can pass the generic gate even when it is socially pernicious. The headline claim that strong multilingual LLMs 'struggle to reliably meet nuanced safety requirements' in Indic languages is therefore a claim about an officially anchored notion of safety, not about harm or ethical principle. The authors disclose this in §5 ('our current questions anchor safety to official government and legal norms') and promise to broaden it, but the disclosure does not repair the metric; it identifies a construct-validity threat that is load-bearing for the central result. Because all models, languages, and harm categories are scored through this same gate, the cross-lingual gap could reflect differential willingness to criticize or conform to official positions in English versus Indic scripts, rather than a deficit in safety behavior per se. This is a correctness risk, not merely a value disagreement, and it needs a robustness check before the benchmark can support the stated 'safety gap' interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SurakshaEval introduces a safety benchmark of 2,968 human-written prompts spanning seven harm types across ten Indic languages and English, and evaluates 27 LLMs using a questionnaire-based Panel-of-LLMs (PoLL) safety judge. The paper reports that safety pass rates are lower for Indic-script prompts than English, that safety degrades most for lower-resource languages and culturally nuanced harm categories, and that English safety does not predict Indic-script safety. The dataset, code, and evaluation pipeline are publicly released.","tokens_in":1803,"tokens_out":1770,"duration_ms":63351,"significance":"If the reported pass rates measured safety, the paper would demonstrate a systematic cross-lingual safety gap and provide a reusable benchmark for a neglected region. The benchmark construction is careful in several respects: human-written prompts, native-script coverage across ten languages, a transparent seven-category taxonomy, public code and data, and a large model roster of 27 open and closed systems. The additional INDIC-BIAS experiments (Section 4.5) provide a useful secondary lens on refusal and stereotype behavior. However, the central metric is not yet validated as a safety measure: the questionnaire's generic gate defines safety largely as alignment with official government positions, and the automated judge agrees with humans at only 69.2% on a single manually checked region, with 44.16% agreement for Societal and Cultural concerns. The headline claim must therefore be re-scoped or the metric re-grounded before the benchmark can support the stated conclusions.","major_comments":[{"comment":"The safety label is defined primarily through the assumed viewpoint of official government positions: Generic Question 2 asks whether an official of the Central or State Government would have a positive attitude toward the statement, Question 4 asks whether the words are suitable for officials, and Question 5 asks whether publication would risk violating government policies or laws. Because §4.3 makes a response safe only when both generic and specific conditions pass, every pass rate in Figures 3–6 encodes the assumption that a response is unsafe if a government official would view it negatively, regardless of harm. A response that documents state misconduct or criticizes a policy would be marked unsafe, while a response echoing official positions can pass the generic gate even when it is socially pernicious. The authors disclose this anchoring in §5, but the disclosure does not remove the threat: all EN-vs-IN comparisons are comparisons on an officially anchored construct, not on harm or ethical principle. The paper should either broaden the generic questions to harm-based criteria and re-run the evaluation, or explicitly re-scope all claims to 'alignment with official government norms.' A concrete test would be to re-annotate a stratified sample with a harm-based rubric and report the disagreement rate with the official-alignment gate.","section":"§4.2, Generic Questions 2–5; §4.3"},{"comment":"The only human agreement check covers 50 prompts from one region (Malayalam) for the top-10 models. Overall agreement is 69.2%, with Societal and Cultural concerns at 44.16% and Regional and Racial issues at 60.91%. These are the harm categories most central to the claim that models miss 'implicit bias' and culturally embedded harms. With agreement near chance on SC, the PoLL-judge pass rates for that category cannot be interpreted as safety rates. Moreover, no manual evaluation is reported for the other nine Indic languages, so the cross-lingual gap in Figures 3–6 rests entirely on an automated judge whose agreement is validated only for Malayalam. I recommend expanding the human sample to at least two or three additional languages, reporting per-language and per-harm-type agreement, and restricting the claim that PoLL 'reliably approximates human judgment' to categories with agreement above a pre-specified threshold.","section":"§4.4, Manual Evaluation; Table 3"},{"comment":"The same GPT-family models (GPT-4.1-mini, GPT-5-mini, GPT-4o-mini) are used both to decide which prompts are unsafe and which harm type they carry (§3) and to judge whether model responses are safe (§4.3). This creates a closed loop in which the benchmark's safety labels are determinations of one model family; the small human check in §4.4 is the only external anchor. The paper should at least report how often the GPT-family judges disagree with non-GPT judges on a sample, ideally an open-weight judge, and consider including such an independent judge in the PoLL ensemble.","section":"§3 and §4.3"},{"comment":"Each of the ten languages had a single native-speaker annotator, with translations via Google Translate and verification by the same annotator plus co-authors. No inter-annotator agreement is reported, and the number of prompts per language varies widely (Table 2: Assamese 128 vs Telugu 211). This does not invalidate the benchmark, but it leaves open the possibility that language-specific differences in prompt difficulty or annotator style drive part of the cross-lingual gap. Reporting at least a second-annotator pass on a subset, with agreement on harm labels and unsafe judgments, would materially strengthen the dataset.","section":"§3, Data Collection"}],"minor_comments":[{"comment":"There are several typographical errors, including 'insufficent' in the abstract, 'official' repeatedly in §4.2, and 'relions' in the Societal and Cultural concerns question in §4.2.","section":"Throughout"},{"comment":"The heatmap captions repeat the N/A explanation inconsistently; I recommend a single consistent statement about unsupported scripts and about the absence of an Avg column for the Indic heatmaps.","section":"Figures 3 and 5"},{"comment":"The row structure for generic/specific/mixed counts is hard to parse; consider splitting the table or using clearer column headers that separate language totals from generic/specific/mixed counts.","section":"Table 2"},{"comment":"The abbreviations B+, B-, and ST are used in the table but defined only in the surrounding text; please define them in the table caption for readability.","section":"Table 6"},{"comment":"Gemini model versions are listed without version numbers or access dates; for reproducibility, please specify the exact model snapshots used and the date of API access.","section":"Appendix B, Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and addresses a genuinely underserved area. The core concern is the construct validity of the safety metric: the authors' own §5 disclosure that the questions 'anchor safety to official government and legal norms' confirms the skeptic's reading that the headline EN-vs-IN comparisons measure conformity to institutional positions rather than harm. This is fixable within the manuscript's scope either by re-scoping the claims or by re-running the evaluation with a harm-grounded rubric. The validation of the automated judge is also thin (one region, one 50-prompt sample, near-chance agreement on Societal and Cultural concerns). I would not reject the paper because the resource itself is valuable and the authors are transparent about limitations; I would require either a re-scoped title and claims or a substantive re-validation before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the resource is real. SurakshaEval gives the field something it didn't have: 2,968 human-written safety prompts in ten Indic scripts plus English, organized into seven harm types with region-specific contexts, and a 27-model evaluation. The two new harm categories (specific individuals, adult content) address real gaps in the earlier English/Hindi-centric benchmarks. Code and data are public. That alone justifies the paper's existence.\n\nThe paper also does some things right. It adapts the Chinese-safety methodology to India rather than just translating prompts; it reports a manual evaluation, albeit small; and it openly discloses in §5 that its safety questions anchor to official government and legal norms. That last disclosure is honest, but it doesn't rescue the metric—it names the problem.\n\nHere's the load-bearing flaw. The generic safety gate in §4.2 asks whether a response would be viewed positively by, be suitable for, or risk violating the policies of a Central or State Government official. A response is safe only if it passes both generic and specific checks. That means every pass rate in Figures 3–6, and every English-vs-Indic comparison, measures alignment with official positions just as much as it measures harm. A model that criticizes a government policy or documents misconduct is marked unsafe; a model that echoes official positions can pass even when the content is socially harmful. The headline finding—that LLMs struggle to meet nuanced safety requirements in Indic scripts—is therefore about a government-anchored notion of safety, not harm or ethical principle. This is a construct-validity threat, not a value disagreement.\n\nThe judge stack adds a second concern. PoLL uses GPT-4.1-mini and GPT-5-mini in both prompt filtering and response scoring. Human agreement is 69.2% overall on a 50-prompt Malayalam subsample, and Societal/Cultural concerns drop to 44.16%. The per-harm-type conclusions for SC, and to a lesser degree RR, are built on a judge whose agreement with humans is barely above chance for that category. The confidence thresholds and consensus criteria are free parameters, and the paper doesn't probe them. These issues are addressable, but as-is they weaken the quantitative claims.\n\nMinor: one annotator per language plus Google Translate, a small human check, and large N/A blocks in the heatmaps complicate cross-lingual comparisons.\n\nWho's this for? Anyone building or evaluating multilingual safety in Indian languages. The dataset deserves a serious referee. I would send it to review, but with expectation of major revision: re-anchor or broaden the safety rubric, re-run the headline analysis under a harm-only scoring rule, expand the human eval, and reframe the abstract to match what the metric actually measures. The resource is valuable; the current conclusion should be read with caution.","headline":"A genuinely useful Indic safety dataset, but the headline 'safety gap' rests on a metric equating safety with official government approval; referee it, and expect major revision.","tokens_in":33249,"tokens_out":3537,"would_cite":true,"duration_ms":39055,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even strong multilingual LLMs fail far more safety checks in Indian-language prompts than in English, and a new benchmark is built to measure this gap.","keywords":["LLM safety","Indic languages","multilingual safety","safety benchmark","regional sensitivity","harm taxonomy","cross-lingual transfer","native scripts"],"falsifier":"A human-rating study in which the same model responses are scored by a diverse panel of Indian community judges using a harm-to-individuals rubric rather than a government-alignment questionnaire: if that rubric does not reproduce the English-versus-Indic gap, or if human judges disagree with the automated labels on a majority of responses, the claim that models are less safe in Indic scripts would be shown to be rubric-dependent.","tokens_in":32149,"feed_emoji":"🛡️","tokens_out":5194,"duration_ms":55720,"temperature":0.7,"pith_summary":"The paper tries to establish that current large language models are systematically less safe in ten major Indian languages, especially in native scripts, than they are in English, and that English-centric safety benchmarks miss this problem. It introduces SurakshaEval, a human-curated set of 2,968 prompts covering seven regionally grounded harm types, with both generic and language-specific items, and a structured questionnaire-based evaluation protocol. Benchmarking 27 multilingual and bilingual models, the paper reports substantial pass-rate drops for Indic-script prompts, with the worst performance in lower-resource languages and in harm categories that require local cultural knowledge. If the finding holds, safety claims for multilingual LLMs cannot be transferred from English, and evaluation must incorporate region-specific data and scripts.","feed_headline":"Indic-language prompts trip up even top LLMs","feed_subtitle":"A 2,968-prompt benchmark shows English safety doesn't transfer to native-script Indian languages.","key_machinery":"The machinery is the SurakshaEval benchmark and its questionnaire-based assessment protocol. The dataset is organized by a seven-type harm taxonomy — regional and racial issues, politically sensitive topics, legal and human rights matters, controversial events, societal and cultural concerns, specific individuals, and adult content — with generic prompts usable across regions and specific prompts tied to one regional context. Safety is scored by an atomic questionnaire whose generic questions define a response as safe if it would be viewed positively by, or would not risk violating the policies of, Indian Central or State government officials, plus harm-specific sub-questions; a Panel of LLMs, six evaluator instances from two models, produces the binary safe/unsafe judgment. This design makes safety assessment decomposable and reproducible, and the paper's headline English-versus-Indic comparisons rest on it.","core_discovery":"The central discovery is that contemporary multilingual LLMs exhibit a consistent cross-lingual safety gap: when prompted in native Indic scripts, they fail to satisfy the paper's safety criteria far more often than when the same prompts are given in English, and English safety performance does not predict Indic-script performance. The best combined pass rates reach about 70 percent for the strongest model, but many models fall below 30 percent on Indic prompts, with the sharpest degradation for lower-resource languages and for harm types such as specific individuals and societal and cultural concerns. The paper also documents recurring failure modes — over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings — and supports its automated judgments with a manual evaluation on a Malayalam subset that agrees about 69 percent of the time.","pith_inferences":["The measured English-versus-Indic gap could partly reflect the government-alignment definition of safety: models that stay neutral on contested political topics may be scored unsafe even when their responses are not harmful to individuals, so an ethics-based rubric might shrink the gap.","A back-translation control study would help isolate how much of the gap comes from script and language difficulty rather than from the content of translated prompts.","The 69 percent automated-human agreement implies that pass-rate differences smaller than roughly a third of judgments could be artifacts of the judge panel; a larger multi-region human study is the natural next check.","Fine-tuning on SurakshaEval preference pairs is a concrete testable extension: if such tuning raises Indic-script pass rates without lowering English pass rates, the benchmark would function as an alignment instrument rather than only a measurement tool."],"forward_implications":["LLM safety evaluation for India must be conducted per language and script, not inferred from English results.","Deployment in Indian contexts should use native-script safety benchmarks for model selection, since lower-resource languages show the largest gaps.","The identified failure modes specify where alignment data is needed: refusal behavior, implicit bias detection, and regional context awareness in Indic languages.","The questionnaire can be converted into preference pairs for supervised fine-tuning or direct preference optimization, turning the benchmark into an alignment tool.","Because the harm taxonomy is concept-level rather than language-specific, the framework can be extended to other regional contexts with relatively few seed prompts.","The benchmark identifies distinct failure modes—over-refusal, implicit bias, and weak contextual awareness—that safety training in Indic languages should target.","If safety does not transfer across languages, then multilingual capability and safety alignment must be tracked as separate axes in model development.","The evaluation pipeline can be reused to track safety in code-mixed and transliterated input settings, an increasingly common real-world usage pattern in India."],"supporting_citations":[{"why":"Supplies the questionnaire-based safety assessment framework and the region-sensitivity taxonomy that SurakshaEval adapts and extends.","marker":"Wang et al. 2024c"},{"why":"Provides the Panel-of-LLMs evaluation strategy used for prompt filtering and for response safety judgments.","marker":"Verga et al. 2024"},{"why":"Supplies the earlier safeguard-evaluation approach and risk taxonomy that the paper moves beyond.","marker":"Wang et al. 2024b"},{"why":"Establishes the prior result that LLMs answer unsafely more often in non-English languages, motivating the cross-lingual hypothesis.","marker":"Wang et al. 2024a"},{"why":"Provides INDIC-BIAS, the fairness benchmark used for the added bias and stereotype experiments, and represents the English-only prior art the paper extends.","marker":"Nawale et al. 2025"},{"why":"Shows existing India-specific bias resources are English-only, highlighting the coverage gap SurakshaEval fills.","marker":"Khandelwal et al. 2024"},{"why":"Shows IndiBias measures social biases in Indian settings but only in English and some Hindi, motivating native-script Indic coverage.","marker":"Sahoo et al. 2024"},{"why":"Exemplifies localized bilingual safety benchmarks outside Indic languages, part of the comparative backdrop for the approach.","marker":"Goloburda et al. 2025"}],"fun_headline_variants":["English safety fails to transfer to Indic scripts","LLMs stumble on native-script Indian safety prompts","70% ceiling: multilingual LLMs hit Indic safety wall","New benchmark exposes Indic-language safety gaps in LLMs","Indic script prompts reveal over-refusal and bias blind spots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that a response counts as safe when it would be viewed positively by, or would not risk violating the policies of, Indian Central or State government officials; if safety is instead understood as avoiding harm to individuals and communities regardless of official stance, the reported pass rates would be measuring conformity to institutional positions. The paper itself flags this anchoring in its conclusion and says it plans to broaden the questions toward general ethical principles.","fun_headline_variants_meta":{"raw":{"variants":["English safety fails to transfer to Indic scripts","LLMs stumble on native-script Indian safety prompts","70% ceiling: multilingual LLMs hit Indic safety wall","New benchmark exposes Indic-language safety gaps in LLMs","Indic script prompts reveal over-refusal and bias blind spots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2604,"prompt_tokens":959,"completion_tokens":1645,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":1582}},"tokens_in":575,"tokens_out":1645,"duration_ms":11507,"temperature":1.0,"reasoning_tokens":1582,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:44:55.476668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-rating study in which the same model responses are scored by a diverse panel of Indian community judges using a harm-to-individuals rubric rather than a government-alignment questionnaire: if that rubric does not reproduce the English-versus-Indic gap, or if human judges disagree with the automated labels on a majority of responses, the claim that models are less safe in Indic scripts would be shown to be rubric-dependent.","supporting_citations":[],"review_version":1}