{"id":"0e2f7b06-e43d-4caa-b018-1c2cb37c686d","arxiv_id":"2608.01436","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The same coercive-control disclosure meets different levels of protective response from chatbots depending on the language it is written in, and the unevenness is attributed to the systems, not to the languages or their speakers.","lead":"Researchers put a scripted domestic-violence scenario to seven AI chatbots in nine languages: a woman asks for help writing an apology letter accepting her partner's phone tracking. They found the chatbots' willingness to comply, and their recognition of coercion, varied sharply by language, with the two strongest systems refusing everywhere.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'design outcome' conclusion rests on two proprietary systems' perfect refusal in one scripted vignette; a robustness check across phrasings and time is needed before generalizing.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the protective-floor inference depends on GPT-5.5 and Claude Haiku 4.5's perfect refusal reflecting a stable design property rather than an artifact of a single vignette or proprietary advantages. I agree. The paper's strongest empirical contribution—the cross-linguistic unevenness measured over 3,528 responses with human-validated coding—stands on its own. However, the abstract and conclusion elevate this into 'failures elsewhere are a design outcome,' which requires more than an existence proof. Two systems succeeding does not show that every other developer could have achieved the same result with reasonable design choices; it could reflect differences in scale, data, or proprietary safety work. The paper is admirably transparent about this in §5.7, but the central claim as stated goes beyond the data. The concrete test I propose directly checks the stability of the floor across phrasings and time; if it fails, the unevenness remains but the moralized 'design outcome' claim would need to be tempered. The reader's CONDITIONAL verdict is therefore appropriate; my read does not change it.","tokens_in":13525,"tokens_out":4171,"duration_ms":41645,"concrete_test":"Re-run the audit with GPT-5.5 and Claude Haiku 4.5 on the same nine languages, expanding the single vignette to at least 10 paraphrase variants per cell (e.g., more colloquial, emotionally flooded, indirect help-seeking, different coercive-control scenarios beyond the apology letter), and repeat sampling on two separate dates. Score the responses with the same codebook. If either system produces any compliant or non-protective response in any language variant, the 'attainable floor' is scenario-dependent and the design-outcome conclusion weakens; if both remain perfect across all variants and dates, the inference gains real support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central attribution—that uneven recognition is 'a property of the systems and not of the languages or their speakers'—leans on the §4.1 claim that GPT-5.5 and Claude Haiku 4.5 held the line in all nine languages, making a protective floor 'attainable' and failures 'a design outcome.' This is an existence proof from two proprietary systems on one fixed vignette, yet the conclusion generalizes it to all builders. If their perfect refusal reflects scale, proprietary safety tuning, or an accidental match to this particular script rather than a stable, replicable design property, then the floor is not established and the moralized 'design outcome' attribution overreaches. The paper itself concedes in §5.7 that this is a single-scenario snapshot and that commercial systems change over time, but the conclusion states the strong claim anyway. The empirical unevenness (e.g., Chinese 43% vs. Hebrew 83% held-line in Table 1) is robust and important; what is load-bearing is the inference from two systems' success to a general capability/design conclusion—especially because the two ceiling systems are proprietary, and other developers may face resource or data constraints that the design does not control for.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper audits seven commercial conversational AI systems (Mistral Medium 3.5, DeepSeek V4 Flash, Gemini 3.5 Flash, Qwen 3.6 Flash, Llama 4 Maverick, GPT-5.5, Claude Haiku 4.5) on one fixed coercive-control scenario in nine languages, scoring 3,528 responses on whether the model writes an apology letter requested by a woman whose partner tracks her phone, and on three protective-content axes (naming control, countering self-blame, affirming agency). It reports substantial cross-lingual differences in held-line rates (Chinese 43% vs Hebrew 83% for all systems), a motive-by-language interaction, a 'home-language' capitulation pattern among three non-anglophone developers, and uniform distress effects. Because GPT-5.5 and Claude Haiku 4.5 refuse in all nine languages, the paper concludes that a protective floor is attainable and failures elsewhere are a design outcome.","tokens_in":13786,"tokens_out":5180,"duration_ms":49044,"significance":"The empirical core is valuable: it is a large, systematically coded audit with inter-coder reliability (κ = .90 on the boundary outcome), back-translation checks, and open data/code, and it addresses a socially important question. The finding that protection varies sharply by language, holding scenario and system constant, is robust and policy-relevant. The paper's two load-bearing interpretive claims—the 'design outcome' attribution and the 'home-language effect'—go beyond what the design can support, and need tempering or additional evidence.","major_comments":[{"comment":"The inference that failures elsewhere are a 'design outcome' rests entirely on two proprietary systems having refused in all nine languages on one fixed scripted vignette. Section 5.7 concedes this is a single-scenario snapshot and that commercial systems change over time. Perfect performance by two large proprietary systems is an existence proof of attainability under this specific script, but not evidence that other builders could achieve the same floor as a design choice; it may reflect scale, safety-tuning investment, or an accidental match between the script and the systems' training. To make the 'design outcome' claim load-bearing, the authors would need robustness across multiple phrasings/scenarios and longitudinal sampling, or at minimum should qualify the conclusion to 'within this scenario family and at this point in time.' As written, the abstract and conclusion state the str","section":"§4.1 and §5.6"},{"comment":"The home-language effect is based on three non-anglophone systems (Mistral, DeepSeek, Qwen), and the paper itself notes the alternative explanation of thinner native-language safety work. With n=3 and heterogeneous developer contexts, the data cannot distinguish a home-language effect from an investment effect. The heading 'Capitulation in the home language' and the claim that 'each gave way most readily in that language' overstate the evidence. This should be reframed as an exploratory pattern, and the 'systems-not-languages' attribution should not lean heavily on it.","section":"§4.5/§5.3"},{"comment":"The endorsement-in-model's-own-voice rates (49% in Hindi/Chinese vs 10% in English) are presented without separate human validation, and the paper acknowledges this axis 'carries wider measurement uncertainty.' Since this register feeds the theoretical discussion of social reproduction, it should be clearly labeled exploratory and separated from the validated results in the abstract's claims.","section":"§4.2"}],"minor_comments":[{"comment":"Citation inconsistency: 'Vergés and Gil-Juárez, 2021' in the text versus 'Vergés Bosch and Gil-Juárez, 2021' in the reference list.","section":"References"},{"comment":"Figure 1 is referenced but not reproduced in the manuscript text; ensure the published version includes it.","section":"Figure 1"},{"comment":"The terms 'protective ceiling' and 'protective floor' are used interchangeably; pick one consistent term to avoid confusion.","section":"Throughout"},{"comment":"The 'resource class' claim for Catalan versus Spanish relies on Joshi et al. (2020), but no resource measure is given; a brief operationalization would strengthen the claim.","section":"§4.5"},{"comment":"The statement that the two top-performing systems 'are reachable at no cost' would benefit from a date and source, since free-tier policies change over time.","section":"§5.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to attract public attention. The empirical disparities are solid and should be highlighted, but the 'design outcome' wording and home-language framing should be tempered before publication. Editors may wish to ask for a robustness check or a clearly qualified conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for its empirical result: the same coercive-control disclosure gets sharply different protection depending on the language it's written in. The design is straightforward—one scripted vignette, nine languages, seven systems, 3,528 responses—and the measurement is careful: independent human coding on a stratified sample, reported kappas, back-translation checks, open code and data. The finding that Chinese, despite being a well-resourced language, has the lowest held-line rate, while Hebrew is near the top, is a genuine counter-example to the usual resource-gradient story. The two-axis separation (refusal vs. protective content) and the language-specific motive effects are also new and worth taking seriously.\n\nWhat's genuinely new: prior audits of intimate-partner-violence disclosures are English-only (Foriest et al.) or adversarial (Yong et al.); this is the first I know of to test a sincere victim-side request across languages. The home-language capitulation pattern—three non-anglophone systems each failing most in their builder's language—is suggestive, and the paper is admirably honest that the design can't distinguish deference from thinner native-language safety work.\n\nSoft spots, in proportion. The 'design outcome' inference rests on two proprietary systems holding the line in all nine languages. That's an existence proof, and the paper scopes it to 'this scenario family' (§5.7) and concedes the snapshot problem. The abstract's phrasing 'failures elsewhere are a design outcome' is a step too strong—accidental or brittle perfection would not be 'design'—but the underlying claim that the unevenness is a property of the systems rather than the languages is logically fine, because some systems avoid it. I'd ask them to soften 'design outcome' to something like 'contingent on system-specific choices.' The endorsement-gradient result (49% vs. 10% endorsing self-blame) lacks separate human validation; the paper flags this itself, so minor. The home-language effect has only three non-anglophone systems, and they are heterogeneous; again, the paper says this.\n\nOverall, a solid, honest empirical audit. The central finding holds up; the interpretive edges are mostly contained by the paper's own limitations section. I'd bring it to reading group and would cite it if I worked on multilingual safety. It deserves a serious referee, not a desk reject; with modest softening of the causal language, it could be a strong venue paper. Recommend peer review.","headline":"Solid cross-lingual audit showing language-dependent protection; the 'design outcome' framing slightly overreaches but the paper's own caveats mostly contain it.","tokens_in":14275,"tokens_out":2823,"would_cite":true,"duration_ms":27289,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A woman's disclosure of coercive control receives materially different protection from widely used chatbots depending on which of nine languages she writes in, and two systems show the gap is a design outcome rather than a language property","keywords":["coercive control","digital gender-based violence","large language models","multilingual safety","algorithmic justice","cross-lingual bias","intimate partner violence","self-blame"],"falsifier":"Re-run the same seventy-two prompt cells with slightly looser wording (a more colloquial or fragmented disclosure) or after a public update to one of the seven models. If either of the two systems that refused everywhere starts writing the apology letter in any language, or if the language ordering reverses substantially, the paper's claim that the unevenness is a stable design outcome would be undercut.","tokens_in":1732,"feed_emoji":"🌐","tokens_out":5284,"duration_ms":113246,"temperature":0.7,"pith_summary":"This paper asks whether a woman disclosing coercive control to a conversational AI gets the same protection regardless of the language she writes in. It puts one scripted scenario - a partner tracks her phone and she asks for a self-blaming apology letter accepting the surveillance - to seven widely used chatbots in nine languages, scoring whether the bot refuses to write the letter and whether it names the control, counters her self-blame, and affirms her agency. The answer is no: refusal and recognition broke unevenly along language lines, with Chinese at the bottom for holding the line and Hebrew at the top, and the gaps did not track how well resourced the language is. Two frontier systems refused every prompt in all nine languages, which the authors read as proof that a protective floor is reachable and that failures elsewhere are a design outcome. The paper argues that recognition of coercive control should be held to a floor, one language at a time.","feed_headline":"Same abuse plea gets unequal AI protection by language","feed_subtitle":"Seven chatbots met one coercive-control scenario in nine languages; refusal and recognition split along language lines.","key_machinery":"The load-bearing object is a single scripted disclosure vignette, varied only by the partner's stated motive (affection, jealousy, paternalistic protection, past betrayal) and his level of distress, rendered identically in nine languages. Responses are scored on two independent axes: a four-level behavioural outcome from outright refusal to full compliance, and a structural index (0-3) counting whether the reply names the control, counters the woman's self-blame, and affirms her agency. The argument's decisive move is the ceiling comparison: two systems that refuse and recognize in all nine languages convert the observed unevenness from an unavoidable language effect into a design outcome.","core_discovery":"The central discovery is that recognition and refusal are separable and unevenly distributed across languages. In a fixed scenario where a woman asks for an apology letter accepting her partner's phone tracking, a model that holds the line (refuses to write) is not the same as one that does protective work (naming the control, countering self-blame, affirming agency), and different languages pull these apart in different ways. Chinese, despite being richly represented in training data, drew the lowest refusal rate among the five varying systems, while low-resource Catalan drew one of the highest; Arabic and Hebrew, sibling Semitic languages, sat at opposite ends. Two of the seven systems - G","pith_inferences":["The authors' explanations for the Catalan and Arabic results - corpus-borne feminist discourse for Catalan, borrowed English safety for Hebrew - are explicitly not measured in this study; a direct test would inspect training-data composition or run controlled fine-tuning experiments varying only the safety-evaluation language.","The authors name the cross of language with an explicitly stated country as the decisive untested experiment; if an explicitly permissive or protective locale overrides the language effect, the injustice becomes easier to remedy, while if language dominates, developers would have to do per-language safety work.","The paper's observation that the best-protecting systems ration free use more tightly implies that the women with least money and support will systematically meet the least protective systems even after a floor is technically attainable; a distributional audit of free-tier limits would quantify that."],"forward_implications":["The same written disclosure garners materially different protection depending on its language; within the varying systems, Chinese held the line in 20% of cases versus 76% for Hebrew.","Good refusal and good recognition are not the same; an evaluation that scores only refusals, or merges both into one index, misses half of what protection is.","A universal protective floor is attainable in this scenario family, because two systems refused and did protective work in every language; shortfalls are design outcomes, not language limitations.","Sympathetic excuses for the partner work differently by language: paternalism raised naming most in Hindi, while Russian's recognition stayed flat, so safety tuning cannot be done once in English and then translated.","The systems most likely to be free at the point of need are the ones that protect least; the better-protecting systems ration sustained use on free tiers."],"supporting_citations":[{"why":"Defines coercive control as a sustained pattern of domination and supplies the three protective practices (naming the control, countering self-blame, affirming agency) that the structural index scores.","marker":"Stark, 2007"},{"why":"Documents technology-facilitated coercive control and the competing roles of digital platforms, grounding the digital dimension of the scenario and the protective-response standard.","marker":"Dragiewicz et al., 2018"},{"why":"Shows low-resource languages jailbreak GPT-4, the adversarial multilingual-safety gap the paper tests on a sincere disclosure and finds does not follow a resource logic.","marker":"Yong et al., 2023"},{"why":"Supplies the language-resource classification used to design the Catalan/Spanish resource pair and to test the resource-based explanation.","marker":"Joshi et al., 2020"},{"why":"The closest prior audit, measuring across English only the interpretive resources generative AI gives to intimate-partner-violence disclosures; the paper's vignette is the reverse design.","marker":"Foriest et al., 2026"},{"why":"Provides the ambivalent-sexism account of benevolent framing (restriction as care) that the motive manipulations are built on.","marker":"Glick and Fiske, 1996"},{"why":"Supplies the testimonial/hermeneutical injustice concepts used to interpret Arabic's naming-the-control-without-uptake as a measurable recognition failure.","marker":"Fricker, 2007"},{"why":"Frames recognition as participatory parity, turning language-wise uneven protection into a status injustice rather than a mere engineering gap.","marker":"Fraser, 2000"},{"why":"Documents deep gender bias in machine translation for Hebrew and Arabic, motivating the matched-pair design that contrasts two sibling Semitic languages.","marker":"Stanovsky et al., 2019"}],"fun_headline_variants":["AI's protective response to abuse varies by language","Same coercive-control plea, different AI answers across languages","Language splits how AI responds to coercive control","Seven chatbots, nine languages: one abuse plea, unequal help"],"cache_read_input_tokens":16128,"weakest_assumption_plain":"The charge that failures elsewhere are a design outcome rests on the assumption that the two ceiling systems' perfect refusal in this one scripted scenario is a stable, replicable property of their design rather than a quirk of this particular story or of proprietary advantages other developers cannot match.","fun_headline_variants_meta":{"raw":{"variants":["AI's protective response to abuse varies by language","Same coercive-control plea, different AI answers across languages","Language splits how AI responds to coercive control","Seven chatbots, nine languages: one abuse plea, unequal help"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2761,"prompt_tokens":743,"completion_tokens":2018,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":1954}},"tokens_in":487,"tokens_out":2018,"duration_ms":16114,"temperature":1.0,"reasoning_tokens":1954,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:09:12.562097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same seventy-two prompt cells with slightly looser wording (a more colloquial or fragmented disclosure) or after a public update to one of the seven models. If either of the two systems that refused everywhere starts writing the apology letter in any language, or if the language ordering reverses substantially, the paper's claim that the unevenness is a stable design outcome would be undercut.","supporting_citations":[],"review_version":1}