{"id":"ff7b2d01-14a3-4316-bf91-8d08e644c2b8","arxiv_id":"2608.03532","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Bias in GPT-5.2 and Gemini 2.5 Flash changes rather than transfers between English and Swahili, with GPT-5.2 refusal behavior appearing only in English.","lead":"This paper tested whether AI chatbots show the same social biases in English and Swahili by sending thousands of matched prompts to two models. It found that bias changes across languages, and one model refused English prompts but never refused the same questions in Swahili, suggesting safety rules are anchored to English.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Translation errors and an English-prompted judge can create, not just attenuate, cross-lingual differences; the neutral-doubling, >55% dissimilarity, and even the refusal asymmetry may be artifacts.","rationale":"The reader's weakest assumption correctly identified the directional claim that translation errors and judge insensitivity only attenuate cross-lingual differences. My stress-test sharpens this into the single most load-bearing concern: the paper's own Sections 3.3, 5.5, and 8.1 provide concrete mechanisms by which these confounds can create the observed effects. The neutral-rate doubling and >55% dissimilarity are exactly the patterns a less-sensitive Swahili judge and non-equivalent Swahili prompts would produce. Because the refusal asymmetry is also not independent of translation quality, the most 'unambiguous' finding is not fully shielded. These are addressable with native prompt correction and human annotation of Pipeline 1, so the conditional verdict remains appropriate; no change to the reader's verdict is needed.","tokens_in":9420,"tokens_out":6045,"duration_ms":75110,"concrete_test":"Run a corrected-prompt replication: have two native Swahili speakers independently correct or validate all 4,900 Swahili prompts for semantic equivalence to the English originals, then re-submit the corrected set to GPT-5.2 and Gemini 2.5 Flash and re-run the Claude judge. Additionally, have the same annotators label a stratified sample of at least 300 Swahili completions per model for stereotype presence and sentiment. If the 169-vs-0 refusal gap, Gemini's ~2x neutral rate, and >55% semantic dissimilarity persist on corrected prompts and human labels, the central claim is supported; if any of these collapse, the confound created the effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'bias transforms rather than transfers' rests on the assumption that the two main confounds — machine-translated prompts and an English-prompted judge — only attenuate cross-lingual differences, never create them. That assumption is not established and is undermined by the paper's own evidence. Section 8.1 acknowledges systematic translation errors, including hallucinated descriptors ('Sikh' translated as 'Msikh') and inappropriate lexical choices. Such errors can make a Swahili prompt semantically non-equivalent to its English original, so the >55% semantic dissimilarity rate may reflect prompt non-equivalence rather than language-dependent bias. Section 5.5 shows the judge is 8 percentage points less accurate on Swahili stressor detection, a task analogous to the sentiment classification in Pipeline 1. A less sensitive judge would over-label Swahili completions as neutral, which is exactly the direction of the headline result that Gemini's neutral sentiment rate doubled (19.8% to 40.1%). The paper validates only 200 semantic-similarity judgments with native annotators (§3.3); Pipeline 1's stereotype and sentiment labels have no direct human validation in Swahili. Finally, the claim in §8.1 that the refusal asymmetry is 'entirely independent of translation quality' is too strong: if a Swahili descriptor is mistranslated or hallucinated, the model may not be processing the same demographic concept, so zero refusals could reflect prompt non-equivalence rather than English-anchored refusal. All three headline results are thus vulnerable to confounds that can create, not merely attenuate, the observed differences.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether social biases generalize across English and Swahili in two commercial LLMs. Using 4,900 symmetric English–Swahili prompt pairs across nine demographic axes, the authors collect 19,600 completions from GPT-5.2 and Gemini 2.5 Flash and evaluate them with Claude Sonnet 4.5 as a cross-provider judge. The headline findings are: aggregate stereotype rates are roughly equal across languages, but per-axis stereotype rates shift by up to 12 percentage points; Gemini's neutral sentiment rate doubles in Swahili; GPT-5.2 refuses 169 English prompts but zero Swahili prompts; and over 55% of English–Swahili completion pairs are judged semantically dissimilar. A logit-lens experiment on Qwen-2.5-3B is presented as preliminary mechanistic evidence. The paper concludes that English-only bias audits are insufficient for multilingual deployment.","tokens_in":9774,"tokens_out":3358,"duration_ms":39802,"significance":"If the results hold, the paper makes an important contribution: it provides one of the first systematic English–Swahili generative-bias comparisons, and it explicitly challenges the assumption that bias transfers language-neutrally. The design has notable strengths: a large, symmetric prompt set; temperature-0 decoding; a judge from a different provider than the evaluated models; an external Swahili-accuracy check on SAD; native-annotator validation of a subset of outputs; and an open dataset/code release. The logit-lens analysis is a reasonable preliminary attempt to distinguish representational divergence from decoding-stage divergence. However, the central empirical claims about 'transformation' rather than 'transfer' depend on two confounds that the manuscript itself documents: machine-translated prompts with systematic errors, and an English-prompted judge with reduced Swahili sensitivity. These confounds are not convincingly shown to only attenuate cross-lingual differences, so the strongest claims currently outrun the evidence.","major_comments":[{"comment":"The 'conservative lower bound' argument is not established. The manuscript documents translation errors that change referents, e.g., 'Sikh' hallucinated as 'Msikh' and 'mixed-race' translated literally rather than as 'chotara' (§8.1). Such errors make the Swahili prompt semantically non-equivalent to its English original. The >55% semantic dissimilarity rate in §5.4 could therefore reflect prompt non-equivalence rather than language-dependent bias. The claim that anglicized translations are 'closer to English' and thus reduce divergence is plausible for some lexical errors, but not for errors that substitute or drop a demographic descriptor. The paper should either filter out/augment validation for translation-faithful prompts, or provide an analysis showing that the headline results persist when only natively validated prompt pairs are used.","section":"§4.1, §5.4, §8.1"},{"comment":"The judge's lower Swahili sensitivity can create, not merely attenuate, the observed sentiment shift. §5.5 reports an 8.0-point accuracy drop on Swahili stressor detection, a task the authors themselves note is close to the subjective sentiment/stereotype classification. A less sensitive judge would be expected to over-label Swahili outputs as neutral, and the headline result for Gemini is exactly that: the neutral rate doubles from 19.8% to 40.1% (§5.2). Pipeline 1 has no direct human validation in Swahili; the 200 judged outputs in §3.3 cover only semantic similarity. To support the claim that sentiment framing changes across languages, the manuscript should report a human-validated Swahili subset for stereotype and sentiment labels, or re-run with a native-Swahili-prompted judge.","section":"§5.2, §5.5, §8.2"},{"comment":"The statement that the refusal asymmetry is 'entirely independent of translation quality' is too strong. If a descriptor is mistranslated or hallucinated, the model may not be processing the same demographic concept, so zero refusals in Swahili could reflect prompt non-equivalence rather than a language-dependent safety mechanism. The asymmetry is certainly consistent with English-surface-form-triggered refusals, but the paper currently presents it as the 'most unambiguous finding' (§6.2) without ruling out the translation confound. A more defensible claim would require testing natively written Swahili prompts or demonstrating that the refusal pattern persists across faithful and erroneous translations.","section":"§5.3, §8.1"}],"minor_comments":[{"comment":"There is a stray '1.' at the end of the paragraph describing the number of completions; it appears to be a footnote artifact and should be removed.","section":"§3.1 / §4.2"},{"comment":"The logit-lens experiment is acknowledged as preliminary because Qwen-2.5-3B is much smaller than GPT-5.2/Gemini 2.5 Flash. This caveat should also be stated in the Conclusion, where the mechanistic interpretation is restated with less hedging.","section":"§4.5 / §6.3"},{"comment":"The SAD Swahili translations were not native-validated, as the text notes; this makes the 'likely conservative' claim in §5.5 speculative. Adding a few translated examples would help reviewers judge input quality.","section":"Table 3"},{"comment":"The 'utoto' example is useful, but the paper does not report inter-annotator agreement for the human validation. A single annotator for 200 items is thin support; agreement statistics with at least two annotators would strengthen the 95% figure.","section":"§8.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a plausible and important central message, and the overall design is better than many cross-lingual bias studies. The main concern is not circularity or misrepresentation, but that the two documented confounds—translation quality and judge sensitivity—are load-bearing for the headline claims. The requested additions (a translation-faithful subsample analysis, human validation of Pipeline 1 in Swahili, or a native-Swahili judge) are feasible within the manuscript's scope and would convert the 'conservative lower bound' assertion into a demonstrated robustness result. I do not see a need for rejection, but the current version oversells the 'transforms rather than transfers' conclusion relative to the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: worth a serious referee, but the core 'bias transforms rather than transfers' claim is undercut by a confound that can create the observed differences, not just mask them. The refusal asymmetry is the most interesting observation, but even that is not as airtight as the paper claims.\n\nWhat is actually new: first open-ended English–Swahili bias comparison; a 169-vs-zero refusal asymmetry on neutral templates; a >55% semantic dissimilarity rate; and an attempt to probe internal representations with logit lens. The authors did some things right: they used a cross-provider judge, ran an external judge-accuracy check on SAD, did a small human validation of outputs, were transparent about translation defects, and used appropriate stats (chi-square, Bonferroni, Cramer's V). The logit-lens section is explicitly preliminary and does not overclaim.\n\nSoft spots in proportion: The 'conservative lower bound' assumption is not established. The paper argues that anglicized translations and a judge with lower Swahili accuracy can only attenuate cross-lingual differences. But translation errors can make the Swahili prompt non-equivalent to the English original, which would overstate semantic dissimilarity and could explain why GPT-5.2 refuses zero in Swahili (a mistranslated or hallucinated descriptor may no longer name the same demographic). The judge's lower Swahili sensitivity could over-label neutral completions, which is exactly the direction of the headline neutral doubling. The claim in Section 8.1 that the refusal asymmetry is 'entirely independent of translation quality' is too strong, contradicted by the paper's own examples. Also, the external SAD validation used Google-translated inputs without native checks, so part of the accuracy gap may be input quality, not judge competence. Minor: the dataset/code link is not actually a link, and human validation is one annotator on 200 samples with no IAA.\n\nNet: this is a useful empirical contribution for people studying multilingual safety, and it demonstrates why English-only audits are insufficient. It deserves peer review; a good referee would ask for native-constructed prompts, multi-annotator IAA, and a judge validated in Swahili, or at least evidence that the confounds work in the direction the authors assume.","headline":"Worth refereeing, but the main claim is built on a confound that may create the cross-lingual differences it reports.","tokens_in":10226,"tokens_out":3338,"would_cite":false,"duration_ms":39909,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Social bias in large language models is not language-neutral: the same model stereotypes, frames sentiment, and refuses prompts differently in English and Swahili, so English-only bias audits do not cover multilingual deployment.","keywords":["cross-lingual bias","Swahili","large language models","stereotype","refusal behavior","semantic similarity","safety alignment","multilingual evaluation"],"falsifier":"Replace the machine-translated Swahili prompts with natively written Swahili and have native annotators label semantic similarity; if the refusal asymmetry and the more-than-55% dissimilarity disappear, the cross-lingual transformation was an artifact of translation and judge bias rather than a property of the models.","tokens_in":9344,"feed_emoji":"🌍","tokens_out":9258,"duration_ms":94740,"temperature":0.7,"pith_summary":"The paper tries to establish that bias in large language models transforms rather than transfers across languages. Using 4,900 symmetric English–Swahili prompt pairs across nine demographic axes, it runs two commercial models and scores 19,600 completions for stereotypes, sentiment, refusals, and cross-lingual semantic similarity. Aggregate stereotype rates look similar in both languages, but per-axis rates shift by up to 12 percentage points, Gemini's neutral-sentiment rate doubles in Swahili, GPT-5.2 refuses 169 English prompts and zero Swahili prompts, and over 55% of prompt pairs produce semantically dissimilar completions. If these results hold, English-only safety audits cannot be assumed to protect users of lower-resource languages.","feed_headline":"Bias transforms, not transfers, across English and Swahili","feed_subtitle":"GPT-5.2 refused 169 English prompts but zero Swahili; over half of paired outputs diverged.","key_machinery":"The load-bearing object is the symmetric English–Swahili prompt pair: the same sentence template with the same demographic descriptor in each language, decoded deterministically so language is the only variable. Fifty templates times descriptors on nine bias axes make 4,900 pairs; two evaluation pipelines then classify every completion for stereotype/sentiment and compare each English–Swahili pair for semantic similarity. A logit-lens probe on an open-weight model provides the mechanistic check of whether divergence lives in internal representations or only in decoding.","core_discovery":"The central discovery is that cross-lingual bias transforms rather than transfers. Although overall stereotype rates are statistically equivalent across languages for both models, the details change: GPT-5.2's Race/Ethnicity stereotype rate rises 12 percentage points in Swahili; Gemini shifts sentiment toward blander, more neutral language in Swahili; GPT-5.2 refuses 169 English prompts and zero Swahili prompts on identical neutral templates; and both models produce semantically similar completions less than half the time. The refusal asymmetry is consistent with safety behavior triggered by English surface forms, and a logit-lens probe (projecting each layer's hidden state onto the model's","pith_inferences":["This is our extension, not the paper's: if corpus asymmetry drives the divergence, languages with even less web presence than Swahili should show larger safety gaps, a claim the paper does not test.","The paper's judge agreed with a native annotator only 95% of the time on semantic similarity, so an automated cross-lingual audit would likely need native adjudication to avoid flattening culturally embedded nuance.","The proprietary black-box results cannot separate representation-level from decoding-level divergence; a direct test would compare per-language fine-tuning against output-layer interventions on the same model.","A practical test of the transformation thesis would be to vary prompt formality or dialect in Swahili and see whether refusal and sentiment outcomes track linguistic distance from English."],"forward_implications":["English-only bias audits will not reliably predict behavior in Swahili, and the gap is likely larger for less-resourced languages.","Refusal guardrails can be language-specific: neutral English prompts may be refused while their Swahili equivalents are answered, leaving bias unmitigated outside English.","Equal stereotype rates across languages can still hide different harms because sentiment framing and semantic content shift.","Divergence is largest on culturally specific axes such as race and migration, pointing to training-data asymmetries as the driver.","Token-level alignment will not generalize; the paper concludes that safety alignment must operate at a conceptual level across languages."],"supporting_citations":[{"why":"Supplies the descriptor-template prompt construction used to build the 4,900 symmetric prompt pairs from 50 templates and demographic descriptors.","marker":"(Smith et al., 2022)"},{"why":"Defines the stereotype-in-ambiguous-context benchmark and the nine bias axes that structure the prompt set.","marker":"(Parrish et al., 2022)"},{"why":"Establishes that translating prompts into low-resource languages can bypass GPT-4 guardrails, the prior result the refusal asymmetry extends to neutral templates.","marker":"(Yong et al., 2023)"},{"why":"Demonstrates cross-lingual transferability of structural representations, the claim the paper argues does not extend to social concepts.","marker":"(Artetxe et al., 2020)"},{"why":"Describes the logit lens technique used to probe hidden states for language-specific representations.","marker":"(nostalgebraist, 2020)"},{"why":"Validates logit-lens analysis across model scales and supports the layer-wise divergence interpretation.","marker":"(Wendler et al., 2024)"},{"why":"Provides the SAD dataset used in the external validation of the judge's cross-lingual classification accuracy.","marker":"(Laine et al., 2024)"},{"why":"Prior finding that bias transforms rather than transfers in decision-making tasks, which this paper extends to open-ended generation.","marker":"(Huijzer and Chen, 2025)"},{"why":"Frames the open questions about how multilingual bias manifests that this study addresses.","marker":"(Gamboa et al., 2025)"}],"fun_headline_variants":["Bias shifts, not transfers, across English and Swahili","GPT-5.2 refuses 169 English prompts, zero Swahili","English-only bias tests miss Swahili divergence","Over half of AI outputs diverge across languages","Stereotype rate jumps 12 points in Swahili"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that machine-translation errors and the English-language judge can only blur cross-lingual differences, never create them; if translated prompts or judge misclassification can fabricate dissimilarity or refusal gaps, the headline numbers could be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Bias shifts, not transfers, across English and Swahili","GPT-5.2 refuses 169 English prompts, zero Swahili","English-only bias tests miss Swahili divergence","Over half of AI outputs diverge across languages","Stereotype rate jumps 12 points in Swahili"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1144,"prompt_tokens":698,"completion_tokens":446,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":362}},"tokens_in":442,"tokens_out":446,"duration_ms":4892,"temperature":1.0,"reasoning_tokens":362,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:07:02.678548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the machine-translated Swahili prompts with natively written Swahili and have native annotators label semantic similarity; if the refusal asymmetry and the more-than-55% dissimilarity disappear, the cross-lingual transformation was an artifact of translation and judge bias rather than a property of the models.","supporting_citations":[],"review_version":1}