{"id":"2899ff7b-e4ea-4ab8-8b20-14a4e4ad3bd8","arxiv_id":"2505.24119","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.","lead":"A systematic review of nearly 300 papers at ACL-family venues shows that LLM safety research is overwhelmingly English-only, and the gap widened every year from 2020 to 2024. The authors release their annotations and propose three directions for making safety evaluation and training more multilingual.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 1's own counts undermine the 'widening gap' claim: English-only share fell from 85.7% in 2020 to 77.1% in 2024, and the English-to-other ratio fell from 6.0 to 3.4.","rationale":"The most load-bearing part of the central claim is the trend assertion, not venue coverage. The paper's own Figure 1 counts contradict 'growing gap' when imbalance is measured in relative terms: the English-only share and the English-to-other ratio both decline from 2020 to 2024. Corpus representativeness would affect the magnitude and precision of the gap, but it is unlikely to reverse the qualitative finding of English dominance; by contrast, the widening claim is internally falsified by the reported numbers. The absolute gap is not evidence of increasing imbalance because it scales with total publication volume; a constant relative gap would produce an increasing absolute gap. The paper's other contributions, such as the released annotation dataset and the high inter-annotator agreement, appear sound, and the worst-case metric illustration is a valid caution. I would not reject the survey, but because a headline quantitative claim is not supported by its own data, acceptance should be conditional on correcting the trend language or reanalyzing with relative imbalance metrics.","tokens_in":23015,"tokens_out":7332,"duration_ms":86467,"concrete_test":"Using the released annotation data (CohereLabsCommunity/multilingual_safety_survey2025), recompute per-year imbalance with relative metrics: (i) proportion of English-only papers; (ii) English-only divided by (non-English + multilingual); (iii) the same ratio restricted to conference-only and workshop-only subsets. Then fit a simple logistic or log-linear trend with year as a predictor. If the slope is non-positive or not significantly positive, the abstract and Figure 1 headline should be revised to say the gap is large and persistent but not widening; if a relative measure does widen, the claim survives and only the presentation needs clarification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Abstract, Section 1, Section 2.2) is that the language gap is 'growing', 'widening', and 'more pronounced over time'. The counts in Figure 1 (English-only: 6, 8, 11, 26, 118; non-English monolingual + multilingual: 1, 3, 2, 8, 35) imply the opposite when imbalance is measured relatively. The English-only share across 2020-2024 is 85.7%, 72.7%, 84.6%, 76.5%, 77.1%, and the ratio of English-only to other-language papers is 6.0, 2.7, 5.5, 3.3, 3.4. Neither series shows a widening trend; the endpoint comparison shows the gap narrowing. The only metric that widens is the absolute difference (5 to 83), which is expected under roughly proportional growth of the whole field and is not a direct measure of imbalance. The sentence 'the increase is disproportionately concentrated in English-only research' is also unsupported: non-English/multilingual papers grew about 35x versus about 19.7x for English-only. The paper's own phrase 'the proportional imbalance remains' is consistent with the relative data, but the abstract and Figure 1 caption assert a growing gap. Since the trend component is part of the headline central claim, the paper's quantitative support for that component is internally inconsistent.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic survey of LLM safety research published at *ACL venues from 2020 to 2024, based on abstracts containing 'safe' or 'safety'. The authors manually annotate the languages studied and safety subtopics for nearly 300 publications, report high inter-annotator agreement, and use the resulting data to argue that safety research is overwhelmingly English-centric, that non-English languages are usually studied only in broad multilingual evaluations, and that English-only papers often fail to document their language coverage. They then propose recommendations for ACL organizers and outline three future research directions: multilingual safety evaluation, culturally contextualized synthetic training data, and crosslingual safety generalization. The paper also includes a short limitations section acknowledging the venue restriction and annotation imprecision.","tokens_in":23168,"tokens_out":4182,"duration_ms":46679,"significance":"If its central claims hold, the paper is a useful community resource: it provides a manually curated, publicly released annotation of language coverage across a substantial slice of LLM safety research, with transparent methodology and high pairwise inter-annotator agreement (0.80-0.96 in Table 2). The survey connects a well-documented multilingual safety problem to a concrete measurement of publication practice, and the proposed future directions, especially the worst-case evaluation metric illustrated in Table 4, are actionable. The paper is not circular: the main measurements are external annotations of the literature rather than derivations from fitted parameters. The principal weakness is that the paper's headline 'growing gap' claim is not supported by its own Figure 1 counts when imbalance is measured relatively, which affects the abstract, the introduction, and Section 2.2.","major_comments":[{"comment":"The paper's central claim that the language gap is 'growing', 'widening', and 'more pronounced over time' (Abstract, Section 1, Figure 1 caption, Section 2.2) is not supported by the counts reported in Figure 1. The English-only counts are 6, 8, 11, 26, 118 for 2020-2024, while the monolingual non-English plus multilingual counts are 1, 3, 2, 8, 35, giving English-only shares of 85.7%, 72.7%, 84.6%, 76.5%, and 77.1%, and English-to-other ratios of 6.0, 2.7, 5.5, 3.3, and 3.4. Neither series shows a widening trend, and the endpoint comparison indicates a narrowing relative gap. The only metric that widens is the absolute difference (from 5 to 83), which is expected under roughly proportional growth and is not a direct measure of imbalance. The statement in Section 2.2 that 'the increase is disproportionately concentrated in English-only research' is also contradicted by the growth rates: the non-English/multilingual category grew about 35x (from 1 to 35) versus about 19.7x for English-only (from 6 to 118). The sentence later in the same section that 'the proportional imbalance remains' is consistent with the data, but the abstract, introduction, and Figure caption assert a growing gap. The authors should revise the trend claims to describe a persistent, not growing, relative imbalance, or explicitly and consistently frame the result as a growing absolute gap with the relative trend stated as a caveat.","section":"Section 2.2 and Figure 1"},{"comment":"The quantitative conclusions (gap size, growth rate, documentation percentages) depend on the corpus defined by two restrictions: venue selection limited to *ACL conferences and workshops, and keyword filtering restricted to abstracts containing 'safe' or 'safety'. Section 2.1 justifies these choices but does not quantify their effect, and the Limitations section acknowledges the venue exclusion without discussing its likely direction or magnitude. Safety work published at ICLR, NeurIPS, ICML, or in journals, and safety work framed as 'jailbreak', 'harm', 'refusal', or 'offensive language' without the selected keywords, is excluded. Because the paper's headline numbers are presented as a measurement of the field's state ('a significant and growing language gap'), the authors should either add a sensitivity analysis over alternative keyword sets and venue sets or explicitly state that the quantitative claims are scoped to *ACL venues with safe/safety abstracts, with an assessment of how the exclusion could bias the reported gap and its trend.","section":"Section 2.1"},{"comment":"The documentation comparison in Table 3 is partly tautological given the annotation protocol described in Section 2.1. A paper is classified as monolingual non-English or multilingual only if the annotation identifies non-English languages, but the documentation metric asks whether the paper explicitly mentions the languages studied; if language identification often depends on explicit mention, then the 100% documentation rates for non-English and multilingual papers may be an artifact of how the categories were constructed. The paper reports that annotators followed up on datasets when languages were not explicitly mentioned (footnote 1), but the magnitude of such inference is not reported. The authors should clarify how many papers in each category were classified without explicit language mention, and should re-analyze the documentation comparison on the subset where the language was identified independently of explicit mention, so that the documentation claim is not circular.","section":"Table 3 and Section 2.2"},{"comment":"The abstract says the paper reviews 'nearly 300 publications', but Section 2.1 reports that 28% of the keyword-matched papers were false positives filtered out before analysis, which would leave roughly 216 analyzed papers. Reporting only the raw count overstates the analyzed corpus and could mislead readers about the strength of the evidence. The authors should state both the initial keyword-matched count and the final analyzed count in the abstract or at the start of Section 2.2.","section":"Abstract and Section 2.1"}],"minor_comments":[{"comment":"The axis label 'JailbreakingattacksT oxicity' in Figure 3(a) appears to have rendering or spacing errors; it should read 'Jailbreaking attacks' and 'Toxicity and bias'.","section":"Figure 3(a)"},{"comment":"The header of Table 4 contains spacing artifacts such as 'A verage↑' and 'W orst Case ∗ ↑'; these should be cleaned to 'Average' and 'Worst Case'.","section":"Table 4"},{"comment":"The reference 'OpenAI. Openai gpt-4.5 system card' should be capitalized as 'OpenAI GPT-4.5 System Card'.","section":"References"},{"comment":"The sentence about 'the increase is disproportionately concentrated in English-only research' is not only unsupported but also inconsistent with the growth-rate comparison; this should be corrected even if the Authors prefer a non-relative framing.","section":"Section 2.2"},{"comment":"The notation in Table 4 uses a red exclamation mark and bold text, but the caption does not define what the red text indicates for color-blind readers; consider replacing color cues with explicit labels.","section":"Section 3.1"},{"comment":"The inter-annotator agreement table reports means and standard deviations, but not confidence intervals or the number of repeated annotations per category beyond the 4x20 design; adding this detail would strengthen the reliability claim.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid survey contribution with a valuable released annotation dataset, and the qualitative finding of a persistent English-centricity in LLM safety research is robust. My main concern is that the headline 'growing gap' claim is contradicted by the paper's own Figure 1 when measured relatively, and this contradiction appears in the abstract, the introduction, the figure caption, and Section 2.2. This is fixable with careful rewording and a clearer distinction between absolute and relative imbalance, so I am recommending major revision rather than rejection. The venue and keyword restrictions and the documentation-comparison circularity also need explicit treatment, but they are secondary to the trend-claim problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely useful survey with a reusable artifact, but the headline trend claim is contradicted by the paper's own Figure 1. English-only share went from 85.7% in 2020 to 77.1% in 2024, and the English-to-other ratio went from 6.0 to 3.4. The only metric that widens is the absolute difference (5 to 83), which is expected under field-wide growth. Non-English/multilingual papers grew about 35x versus about 19.7x for English-only, so the sentence 'the increase is disproportionately concentrated in English-only research' is not supported. The abstract, introduction, and Figure 1 caption all say the gap is growing; that part of the central claim should be withdrawn or reanalyzed.\n\nWhat is actually new: the annotated corpus of nearly 300 papers with language, subtopic, and documentation labels, released on Hugging Face. Prior diversity audits covered NLP broadly and prior safety surveys did not quantify the language dimension. Inter-annotator agreement of 0.80 to 0.96 is solid. The subtopic and venue breakdowns are informative, and the worst-case harmlessness reanalysis of Wang et al. is a small but valid illustration.\n\nSoft spots beyond the trend flaw: the corpus is restricted to *ACL venues and abstracts containing 'safe' or 'safety'; the authors disclose this but do not quantify recall, and jailbreak/harm-framed work at ML venues is excluded. The language documentation comparison is partly tautological: non-English papers are recognized as such because their languages are mentioned or inferable, so the 100% documentation figure is close to built into the categorization. These affect exact percentages, not the qualitative direction. The qualitative English-centricity claim holds across subtopics and years.\n\nThe paper deserves a serious referee; the dataset and measurement are worth engaging. A referee should require fixing the trend claim and validating the keyword filter. I would cite the dataset if I worked in multilingual safety. Reading group: yes, as a case study in a survey's conclusions outrunning its own counts.","headline":"Useful survey and dataset, but the headline 'growing language gap' is contradicted by the paper's own Figure 1.","tokens_in":23818,"tokens_out":2651,"would_cite":true,"duration_ms":28508,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM safety research is overwhelmingly English-centric, and the imbalance has grown wider every year from 2020 to 2024, according to a hand-annotated survey of nearly 300 papers.","keywords":["multilingual safety","LLM safety","English-centric bias","language gap","safety evaluation","language documentation","crosslingual generalization","survey"],"falsifier":"Run the same keyword-and-annotation procedure on the machine-learning conferences the paper excludes (for example, ICLR, NeurIPS, and ICML) or on a year of fresh arXiv submissions, and compare the share of multilingual and non-English safety papers. If that share is close to the English-only share, or if the absolute count of non-English safety papers in those venues is large, the claim that the gap is widening over time would be weakened. A second check: re-annotate the papers the survey counts as 'multilingual' for whether they report per-language results; if most report only aggregate multilingual scores, the claim that non-English coverage is shallow is confirmed, and if many report per-language breakdowns, it is not.","tokens_in":22693,"feed_emoji":"🌐","tokens_out":10303,"duration_ms":88369,"temperature":0.7,"pith_summary":"This paper is a systematic survey of nearly 300 LLM safety publications from 2020 to 2024, drawn from the ACL family of NLP venues and annotated by hand for which languages each paper actually studies. Its central claim is that LLM safety research is overwhelmingly English-centric, that even the second-most-studied language (Mandarin Chinese) receives about ten times less attention than English, and that the absolute gap is widening over time, from 5 English-only papers more than multilingual ones in 2020 to 83 more in 2024. It also finds that non-English languages are usually folded into broad multilingual evaluations rather than studied in depth, and that only half of English-only safety papers even name the language they study. If the survey is right, the field's safety assurance is systematically under-validated for non-English speakers, exactly the populations models are increasingly deployed to serve.","feed_headline":"LLM safety research ignores nearly every language but English","feed_subtitle":"A five-year survey of ~300 papers finds the English-only gap widening, with half of papers never naming the language","key_machinery":"The carrying object is the survey corpus itself: nearly 300 publications from 2020 to 2024 at *ACL conferences and workshops, selected by keyword matching of 'safe' and 'safety' in abstracts, manually categorized into a seven-way safety taxonomy (jailbreaking attacks, toxicity and bias, factuality and hallucination, AI privacy, policy, LLM alignment, and 'not related to safety'), and annotated for the languages each work actually studies. The annotation scheme (English-only, monolingual non-English, or multilingual) shows inter-annotator agreement between 0.80 and 0.96, which is what lets the paper report gap statistics with confidence. Two analytical devices carry the argument: a frequency-versus-multilinguality plot showing that high-resource non-English languages are studied mostly in bulk multilingual papers, and a language-documentation measure, following Bender's rule that papers should name the languages they study, which reveals that 50.6% of English-only papers never mention English. The paper also re-scores an existing ten-language harmlessness table with a 'worst-case' column, showing that average scores hide catastrophic per-language failures — Vicuna's Bengali score of 18.4 versus a high average — as a concrete illustration of what English-centric reporting obscures.","core_discovery":"The paper's central claim, stated in its own words, is that 'the vast majority of safety research is centered on English-language models, while comparatively little work addresses safety in non-English or multilingual contexts,' and that this imbalance has become more pronounced over time. The evidence is a hand-annotated corpus of nearly 300 papers from 2020 to 2024 at *ACL venues, filtered by the keywords 'safe' and 'safety' in abstracts and grouped into six safety subtopics. English-only work outnumbers multilingual and non-English work combined in every year, every subtopic, and both conference and workshop settings, and the proportional imbalance has persisted even as overall publication counts rose. Non-English languages appear mostly 'in herds,' as items inside large multilingual test suites — Swahili, Telugu, and Afrikaans, for example, appear almost exclusively that way — and 50.6% of English-only papers never state explicitly that they studied English. The paper reads these findings as a safety failure, not merely a diversity gap, because refusal training and other alignments have repeatedly been shown not to transfer across languages, leaving language-specific harms undetected as models deploy globally.","pith_inferences":["Extending the same keyword-and-annotation count to the machine-learning venues the paper excludes (for example, ICLR, NeurIPS, and ICML) could plausibly show an even wider gap, since those venues have weaker language-diversity norms than the ACL family; the Limitations acknowledge the exclusion but do not quantify its effect.","The 'studied in herds' pattern generates a testable prediction: re-annotating the multilingual papers for whether they report per-language failure rates rather than aggregate scores would show that much non-English coverage is inclusion-by-checklist rather than depth.","The worst-case score device could be generalized into a reporting standard for model cards and leaderboards, since a per-language minimum would have flagged Vicuna as unsafe for Bengali deployment despite its acceptable average."],"forward_implications":["Reporting only average safety scores can certify a model as safe while one language remains badly unprotected; adding worst-case per-language scores to evaluations and leaderboards is a concrete minimal fix.","Current safety benchmarks are built almost entirely from English and Chinese content, so alignment validated on those tests cannot be assumed to hold for other languages; models evaluated on the full language set can fail badly in languages that were exempted from red-teaming.","Making the language-coverage metadata field in the submission system public would let the community track linguistic representation in safety research with essentially no extra effort.","Dedicated conference tracks and shared workshop tasks on multilingual safety would give non-English safety work a more accessible outlet, since the survey finds such work already appears disproportionately in workshops.","Future research should prioritize culturally grounded evaluation benchmarks, diverse multilingual safety training data (including constitutional-AI pipelines and machine translation with cultural checks), and mechanistic or influence-based study of how safety alignment transfers across languages."],"supporting_citations":[{"why":"Supplies the opening evidence that low-resource languages can jailbreak GPT-4, grounding the claim that safety alignment does not transfer across languages.","marker":"[Yong et al., 2023a]"},{"why":"Provides the ten-language harmlessness benchmark that the paper re-scores with a worst-case column, and is the main cited evidence of weaker safety in non-English prompting.","marker":"[Wang et al., 2024a]"},{"why":"The prior safety survey establishing that every public safety evaluation dataset it reviewed includes English, with only two bilingual (English–Chinese) datasets, anchoring the English-centricity claim.","marker":"[Dong et al., 2024]"},{"why":"Defines 'Bender's rule' that researchers must name the languages they study, the standard behind the language-documentation finding.","marker":"[Bender, 2011; 2019]"},{"why":"Supplies the seven-category safety taxonomy the paper adopts for manual categorization of the corpus.","marker":"[Cui et al., 2024]"},{"why":"The Llama-3 system card showing red-teaming results for only eight languages, evidence for the gap between multilingual deployment and safety evaluation.","marker":"[Grattafiori et al., 2024]"},{"why":"The Arabizi case study showing standardized-script evaluation misses jailbreaks in natural writing, motivating evaluation with real linguistic patterns.","marker":"[Al Ghanim et al., 2024]"},{"why":"The code-switching red-teaming study that motivates the call for multilingual, multi-turn safety evaluation.","marker":"[Yoo et al., 2024]"}],"fun_headline_variants":["Safety research skips nearly all non-English languages","300-paper survey: safety work is overwhelmingly English-only","Non-English languages barely studied in LLM safety research","Half of English-only safety papers never name the language","Safety gap grows: multilingual research lags in every subtopic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's quantitative conclusions rest on the premise that LLM safety papers published at *ACL venues whose abstracts contain 'safe' or 'safety' fairly represent the whole field; if a substantial body of multilingual safety work lives in machine-learning venues or under different framing, the measured gap would be overstated, and the Limitations note the venue exclusion without quantifying it.","fun_headline_variants_meta":{"raw":{"variants":["Safety research skips nearly all non-English languages","300-paper survey: safety work is overwhelmingly English-only","Non-English languages barely studied in LLM safety research","Half of English-only safety papers never name the language","Safety gap grows: multilingual research lags in every subtopic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1402,"prompt_tokens":927,"completion_tokens":475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":397}},"tokens_in":543,"tokens_out":475,"duration_ms":5072,"temperature":1.0,"reasoning_tokens":397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:34:31.220850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same keyword-and-annotation procedure on the machine-learning conferences the paper excludes (for example, ICLR, NeurIPS, and ICML) or on a year of fresh arXiv submissions, and compare the share of multilingual and non-English safety papers. If that share is close to the English-only share, or if the absolute count of non-English safety papers in those venues is large, the claim that the gap is widening over time would be weakened. A second check: re-annotate the papers the survey counts as 'multilingual' for whether they report per-language results; if most report only aggregate multilingual scores, the claim that non-English coverage is shallow is confirmed, and if many report per-language breakdowns, it is not.","supporting_citations":[],"review_version":1}