{"id":"d86e2576-cd5f-46e3-83cb-b8a79eb063f4","arxiv_id":"2504.12767","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Language models show higher offensive stereotyping bias against overlooked marginalized groups in English and German, while Arabic models show high bias against both marginalized and dominant groups, especially for religion and ethnicity.","lead":"Researchers measured stereotyping bias in 23 AI language models against 270 marginalized groups in 25 countries, including many groups never tested before. The results suggest models are more biased against minority groups than dominant ones, and that measured bias rises when prompts use a low-resource dialect like Egyptian Arabic.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOS_MLM metric (Eq. 1) likely measures template-level toxicity preference rather than identity-specific stereotyping; without a neutral-noun control, the marginalized-vs-dominant and dialect claims are not yet supported.","rationale":"The reader's weakest assumption already identifies template wording and the lack of validation for Eq. 1 as the central risk; my stress-test sharpens the mechanism by noting that the identity term is in the scored set U rather than the manipulated set M, so the marginalized-vs-dominant comparison can only be driven by the identity token's conditional probability given the toxic/non-toxic filler. That identity-specific term is never isolated or validated. I therefore agree with the reader's conditional verdict: the empirical claims should not be treated as definitive until a control condition demonstrates that Eq. 1 tracks identity-specific offensiveness. The paper does contribute a substantial multilingual identity resource and candid qualitative examples, but those do not repair the unvalidated central metric. If the neutral-noun control fails, the verdict should move to REJECT; absent that test, CONDITIONAL remains appropriate, so I leave the reader's verdict unchanged.","tokens_in":21692,"tokens_out":9278,"duration_ms":105352,"concrete_test":"Run Eq. 1 on the same 37 toxic/non-toxic template pairs in each language with the [Identity] slot filled by 20 neutral, non-demographic nouns (e.g., 'person', 'neighbor', 'colleague') matched for corpus frequency with the identity terms. If the mean neutral-condition SOS_MLM score is not close to 0.5 and does not differ materially from the marginalized/dominant group means, the metric is dominated by template-level lexical effects rather than identity-specific bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 1 is the load-bearing measurement for RQ2 and for the Egyptian-vs-MSA claim, yet its construct validity is unestablished. In Eq. 1, S = U ∪ M: M is the toxic or non-toxic filler (adjective/verb), while U contains the identity term plus the shared frame tokens ('Being', 'a', 'person', 'is'). The score is the pseudo-log-likelihood of U conditioned on M, and the comparison score(S) > score(S') is dominated by the shared frame's fit with the toxic filler, which is identical for every identity. The only identity-specific term is log P(identity | M) inside the sum, and nothing in the paper shows that this term is large enough, stable enough, or aligned with human judgments of stereotyping to drive the reported marginalized-vs-dominant and dialect gaps. There is no control condition with neutral fillers, no matched toxic/non-toxic word sets controlled for frequency and register, no human validation, and no significance testing. The paper does acknowledge HONEST/HurtLex coverage limitations (Sec. 4.2, 5.1), but it does not acknowledge this structural confound for SOS_MLM, which is the metric singled out for RQ2 in Sec. 5.2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale audit of offensive stereotyping bias (SOS) in 23 language models across 270 marginalized groups and 60 dominant groups in Egypt, the remaining 21 Arab countries, Germany, the UK, and the US, using English, German, Modern Standard Arabic (MSA), and Egyptian Arabic. The authors build a new SOS dataset of 72,000 sentences, extend the HONEST dataset, and propose a new masked-language-model metric, SOS_MLM, based on pseudo-log-likelihood comparisons of toxic and non-toxic template pairs. They report that instruction-following models hallucinate substantially more on Arabic instructions, that HONEST scores are lower in non-English languages partly because of HurtLex coverage gaps, that Egyptian Arabic yields higher bias scores than MSA, that most MLMs show higher bias against marginalized than dominant groups except Arabic MLMs on religion and ethnicity, and that intersectional bias is especially high for non-binary and Black women. The paper also tests the dialect hypothesis with CamelBERT-Da and finds higher Egyptian than MSA scores. The authors conclude that existing bias evaluation, which is largely US/English-centered, is incomplete for regional and low-resource contexts.","tokens_in":21965,"tokens_out":4177,"duration_ms":44426,"significance":"If the measurement claims hold, this is a valuable contribution because it broadens bias evaluation beyond US/English settings, introduces regional identity groups that are rarely studied, and provides a multilingual dataset and code. The paper has concrete strengths: it ships data and code, it includes an additional dialect-trained model (CamelBERT-Da) as a falsifiable probe of the pretraining-data hypothesis, it reports qualitative examples of toxic completions, and it explicitly acknowledges the HurtLex coverage limitations for non-English languages. However, the significance is conditional because the paper's central comparisons for RQ2 rest on the SOS_MLM metric, whose construct validity is not established, and because the repeated use of 'significantly higher' is not backed by inferential tests. The dataset and descriptive findings are useful, but the main scientific claims require additional validation.","major_comments":[{"comment":"The construct validity of the proposed SOS_MLM metric is not established, and this is load-bearing for RQ2 and for the Egyptian-vs-MSA claim. In Eq. (1), score(S) is the sum of log pseudo-likelihoods of all unmodified tokens U, which include the identity token plus shared frame tokens such as 'Being', 'a', 'person', 'is'. The frame-token contribution is identical across identities, so the comparison score(S) > score(S') can be dominated by whether the shared frame fits toxic versus non-toxic fillers, rather than by an identity-specific stereotype. The identity-specific signal is confined to log P(identity | M), and the paper provides no neutral-noun control, no matched toxic/non-toxic filler sets controlled for frequency and register, and no human-annotation validation to show that this term drives the reported marginalized-vs-dominant and dialect differences. Without such a control or validation, the metric may be measuring template-level toxicity preference rather than bias against specific identity groups.","section":"4.3/Eq. 1"},{"comment":"The paper repeatedly uses 'significantly higher' (e.g., English HONEST scores, Egyptian versus MSA SOS_MLM scores) without any inferential statistics. Table 3 and Fig. 1a report only means; there are no confidence intervals, paired tests, or effect sizes. Because SOS_MLM and HONEST scores are aggregates over many sentence pairs, bootstrap or mixed-effects analyses are feasible and would substantiate the claimed differences. As written, the dialect and cross-language comparisons are descriptive trends, not statistically supported findings.","section":"4.2/4.3/5.1"},{"comment":"The cross-language comparison of HONEST scores is confounded by lexicon coverage. The paper itself notes that HurtLex has 3360 English entries versus 1147 Arabic and 2043 German entries, and that the Arabic lexicon includes English hurtful words, yet Sec. 4.2 states that 'HONEST scores are significantly higher for the English dataset' without adjusting for coverage. The authors do acknowledge this limitation later, but the main quantitative claim in Sec. 4.2 remains misleading and should be either re-analyzed with coverage-matched lexicons or explicitly re-framed as a metric-artifact hypothesis. This does not necessarily invalidate the qualitative examples, but it weakens the quantitative cross-language and cross-dialect HONEST comparisons.","section":"4.2/5.1"}],"minor_comments":[{"comment":"The title as submitted to arXiv contains 'Out of Sight Out of Mind' twice; the in-text title uses it once. This duplication should be corrected.","section":"Title"},{"comment":"The dataset-size statement is hard to verify: 72,000 sentences is not derived from the stated 37 toxic and 37 non-toxic templates, the number of identity terms, and the three gender variants. Please provide the exact arithmetic or a table clarifying how the counts are obtained.","section":"3.2"},{"comment":"In the third item of Sec. 5.1, 'we hypothesis that' should be 'we hypothesize that'; also the later sentence 'This internet access gap relates to income' is a fragment that should be joined to the previous sentence.","section":"5.1"},{"comment":"Reference [91] 'Askari, S.' has no title, venue, or year, and reference [49] 'Queerinai, O. O. et al.' is malformed. Please complete these entries.","section":"References"},{"comment":"The appendix figures are dense and hard to read; consider labeling the panels with the model names and sensitive attributes directly in the figure rather than only in the caption.","section":"A.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful empirical contribution with a substantial new dataset, and the authors are transparent about several limitations. The main risk is that the paper's central conclusions about marginalized-vs-dominant bias and dialect differences depend on a metric whose construct validity is not yet demonstrated. If the authors add a neutral-noun control or an identity-conditioned analysis and provide inferential statistics, the claims could become much stronger. I would not reject the paper; the dataset and the qualitative findings are valuable, but the current analysis does not yet support the headline quantitative claims as stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first bias audit to cover marginalized groups in the Arab world at scale, and the dataset is a real contribution: 270 marginalized and 60 dominant groups across 25 countries, with templates in English, German, Modern Standard Arabic, and Egyptian Arabic. Second, the metric used for the MLM results—the part that supports the paper’s main claims about group differences and dialect effects—is not yet validated, and the paper’s quantitative conclusions should be read as provisional.\n\nThe paper does several things well. It extends HONEST and CrowS-Pairs-style evaluation to a much wider set of identities and languages, and it runs a useful extra model (CamelBERT-Da) to test the hypothesis that Egyptian dialect under-representation drives higher bias scores. The qualitative examples of toxic completions are disturbing and illustrative. The authors also acknowledge the HurtLex coverage problem for cross-language HONEST comparisons, which is responsible for some of the English-vs-Arabic gap.\n\nThe soft spots are real. The SOS_MLM metric in Eq. 1 compares pseudo-log-likelihoods of toxic vs. non-toxic template pairs that differ in more than toxicity. The shared frame, including the identity term, is the same in both sentences; what changes is the filler adjective or verb. The score is dominated by how well the toxic filler fits the frame, which is identical across identities, and the identity-specific term—log P(identity | M)—is not shown to be large or stable enough to drive the results. There is no neutral-noun control, no matching of toxic and non-toxic words for frequency or register, and no human validation. So the marginalized-vs-dominant and Egyptian-vs-MSA claims rest on an instrument whose construct validity is unestablished. The paper does not acknowledge this confound. Also, 'significantly higher' appears repeatedly without confidence intervals or inferential tests. HONEST scores across languages are partly artifacts of lexicon coverage, as the authors note. And the complete identity list, translations, and decoding settings are not in the preprint, which makes replication difficult.\n\nIf you work on bias evaluation, the dataset alone is worth having on your radar. But treat the headline numbers as hypotheses, not measurements. A serious referee should engage with this paper, because the resource is important and the flaws are fixable. The authors should validate SOS_MLM against human judgments, add significance testing, and release the full artifacts. If those happen, this becomes a solid contribution.","headline":"A genuinely new bias-audit resource for overlooked Arab-world groups, but the MLM metric that carries the main claims needs validation before those claims should be trusted.","tokens_in":22461,"tokens_out":2731,"would_cite":true,"duration_ms":26025,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 23 language models and 270 marginalized groups in 25 countries, the paper reports that English and German models show higher offensive stereotyping bias against marginalized groups than dominant groups, while Arabic MLMs score high…","keywords":["offensive stereotyping bias","language model bias","marginalized groups","regional contexts","Egyptian Arabic","Modern Standard Arabic","low-resource languages","intersectionality"],"falsifier":"To settle the central claim, run the SOS_MLM and HONEST metrics on matched templates where toxic and non-toxic versions are equated for length, frequency, and register, and have human raters score offensiveness for the same prompts. If the metric's group-level ordering diverges from human judgments, or if the Egyptian-MSA gap disappears under matched templates, the paper's bias comparisons are measurement artifacts rather than model bias.","tokens_in":21519,"feed_emoji":"⚖️","tokens_out":7108,"duration_ms":70409,"temperature":0.7,"pith_summary":"Across 23 language models and 270 marginalized groups in 25 countries, the paper claims that English and German models score higher offensive stereotyping bias against marginalized groups than against dominant groups, while Arabic masked language models show high bias against both marginalized and dominant groups on religion and ethnicity. It also reports that bias scores are higher when prompts are written in Egyptian Arabic than in Modern Standard Arabic, and attributes this gap to dialect under-representation in pretraining data. The study builds a new multilingual dataset covering six sensitive attributes and introduces an MLM-specific bias metric to make regional measurement possible. If the results hold, current bias evaluation and debiasing, which center largely on US and English-speaking contexts, miss the populations most exposed to harm.","feed_headline":"Egyptian Arabic prompts trigger more model bias than standard Arabic","feed_subtitle":"Studying 270 marginalized groups across 25 countries, the paper shows regional and dialect bias that current metrics undercount.","key_machinery":"Two measurement instruments carry the argument. The SOS_MLM metric, introduced here, compares the pseudo-log-likelihood that a masked language model assigns to a toxic sentence versus its non-toxic template twin containing the same identity term; the bias score is the fraction of identity-prompt pairs for which the toxic version is more probable. For generative models, the paper adapts the HONEST metric, which counts how often the model's top-K completions contain a hurtful word from the HurtLex lexicon. The dataset built for this study contains 270 marginalized and 60 dominant identity groups across six sensitive attributes, in English, German, Modern Standard Arabic, and Egyptian Arabic, and is what allows group-level and dialect-level comparisons for the first time.","core_discovery":"The central claim is that offensive stereotyping bias in language models is a regional and low-resource-language problem that current English-centered benchmarks fail to capture. The paper's three main results are: English and German models consistently show higher SOS bias against marginalized groups than against dominant groups; Arabic MLMs instead score high bias against both marginalized and dominant groups for religion and ethnicity, which the authors trace to pretraining on Western news translated into Arabic; and measuring with Egyptian Arabic yields significantly higher bias scores than Modern Standard Arabic, a gap attributed to under-representation of Egyptian content in pretraining corpora. The paper also reports pronounced intersectional bias against non-binary, LGBTQIA+, and Black women, and documents that multilingual instruction-following models hallucinate far more in Arabic than in English. Finally, it shows that the HONEST metric's reliance on a smaller Arabic HurtLex makes cross-language bias comparisons unreliable.","pith_inferences":["Because template pairs differ by more than toxicity, part of the measured SOS_MLM score could reflect lexical or syntactic preferences; the authors do not validate the metric against human offensiveness judgments, so a human-annotation study would separate bias from artifact.","Low SOS_MLM scores for identities absent from Arabic training data, such as 'Muhamash' or 'Ahmadi', may indicate that the model has no association for them at all; interpreting those low scores as 'low bias' would be a misreading.","If the dialect gap is driven by pretraining-corpus composition, the same protocol applied to another underrepresented Arabic dialect should show a similar gap; this is a testable extension the paper does not run.","The Arabic-MLM result of high bias against dominant groups predicts that retraining on local rather than translated data would lower bias scores for both dominant and marginalized groups; that prediction could be checked by fine-tuning the same architecture on curated local corpora."],"forward_implications":["Existing bias benchmarks that cover only US and English-speaking groups will miss the majority of marginalized groups worldwide; regional group inventories are needed.","Low-resource languages and dialects cannot be audited with English-derived metrics: HurtLex's 3,360 English entries versus 1,147 Arabic and 2,043 German entries mean cross-language HONEST comparisons understate non-English bias.","If Arabic models inherit Western-media stereotypes through translated pretraining data, training on local, representative sources is a concrete debiasing lever.","Higher measured bias in Egyptian Arabic than MSA implies dialect-specific evaluation is required and that dialect under-representation is itself a bias mechanism.","Intersectional identities, especially non-binary and Black women, need separate measurement; aggregate group scores hide them."],"supporting_citations":[{"why":"HONEST dataset and hurtful-completion metric; supplies the generative-model bias measure the paper extends.","marker":"[42]"},{"why":"CrowS-Pairs; provides the pseudo-log-likelihood sentence-scoring method that the proposed SOS_MLM metric adapts.","marker":"[5]"},{"why":"Borkan et al. toxicity dataset; supplies the 37 toxic and 37 non-toxic template pairs used to build the SOS dataset.","marker":"[53]"},{"why":"HurtLex; the lexicon whose entry counts drive HONEST scores and the cross-language coverage limitation.","marker":"[59]"},{"why":"CamelBERT-Da; dialect-pretrained Arabic MLM used to test the hypothesis that dialect data raise bias scores.","marker":"[84]"},{"why":"Sap et al.; the African American English finding that motivates the Egyptian-versus-MSA dialect comparison.","marker":"[30]"},{"why":"Masud et al.; provides rectified F1 and hallucination measurement used to evaluate instruction-following models.","marker":"[54]"},{"why":"Minority Rights platform; source of the country-specific ethnic and religious marginalized group lists.","marker":"[39]"},{"why":"Introduces the offensive stereotyping bias (SOS) concept that the paper operationalizes in LMs.","marker":"[15]"}],"fun_headline_variants":["Dialect gap: Egyptian Arabic exposes more LM bias than standard","English-centric tests miss LM bias against overlooked groups","Arabic LMs show high bias for all groups, not just minorities","Intersectional bias spikes for non-binary, LGBTQIA+, Black women","First regional study maps stereotyping bias in 23 language models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two bias scores measure bias against the identity groups themselves; if template wording, translation choices, or the different sizes of the hurtful-word lexicon across languages dominate the scores, the paper's group and dialect comparisons do not measure what they claim.","fun_headline_variants_meta":{"raw":{"variants":["Dialect gap: Egyptian Arabic exposes more LM bias than standard","English-centric tests miss LM bias against overlooked groups","Arabic LMs show high bias for all groups, not just minorities","Intersectional bias spikes for non-binary, LGBTQIA+, Black women","First regional study maps stereotyping bias in 23 language models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1745,"prompt_tokens":1061,"completion_tokens":684,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":597}},"tokens_in":677,"tokens_out":684,"duration_ms":7119,"temperature":1.0,"reasoning_tokens":597,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:22:38.694261+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"To settle the central claim, run the SOS_MLM and HONEST metrics on matched templates where toxic and non-toxic versions are equated for length, frequency, and register, and have human raters score offensiveness for the same prompts. If the metric's group-level ordering diverges from human judgments, or if the Egyptian-MSA gap disappears under matched templates, the paper's bias comparisons are measurement artifacts rather than model bias.","supporting_citations":[{"cited_title":"HONEST: Measuring Hurtful Sentence Completion in Language Models","cited_arxiv_id":null,"evidence_quote":"HONEST dataset and hurtful-completion metric; supplies the generative-model bias measure the paper extends."},{"cited_title":"Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification","cited_arxiv_id":null,"evidence_quote":"Borkan et al. toxicity dataset; supplies the 37 toxic and 37 non-toxic template pairs used to build the SOS dataset."},{"cited_title":"CEUR Workshop proceedings","cited_arxiv_id":null,"evidence_quote":"HurtLex; the lexicon whose entry counts drive HONEST scores and the cross-language coverage limitation."},{"cited_title":"The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models","cited_arxiv_id":null,"evidence_quote":"CamelBERT-Da; dialect-pretrained Arabic MLM used to test the hypothesis that dialect data raise bias scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sap et al.; the African American English finding that motivates the Egyptian-versus-MSA dialect comparison."},{"cited_title":"Hate Personified: Investigating the role of LLMs in content moderation","cited_arxiv_id":null,"evidence_quote":"Masud et al.; provides rectified F1 and hallucination measurement used to evaluate instruction-following models."},{"cited_title":"https://minorityrights.org/world-map/, 2023","cited_arxiv_id":null,"evidence_quote":"Minority Rights platform; source of the country-specific ethnic and religious marginalized group lists."}],"review_version":1}