{"id":"25b3323f-1ae0-4ae3-ad69-bc9c512b15f1","arxiv_id":"2504.15941","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new English-French benchmark shows that current LLM translators rarely produce French inclusive forms and often misread the singular 'they' as plural.","lead":"This paper introduces FairTranslate, a human-annotated English-French dataset for measuring how well translation systems handle non-binary and inclusive language. Evaluations of four large language models show they translate inclusive forms, such as the singular 'they', far worse than masculine or feminine forms.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'almost never' and 'misunderstands singular they' conclusions depend on a narrow inclusive-orthography scorer that is not shown to use the §4.3 alternatives dictionary; a semantic audit is needed.","rationale":"FairTranslate is a useful and largely well-constructed benchmark: the three-variant counterfactual design controls for source content, the human annotation is a real asset, the public dataset and code are important contributions, and the BLEU/COMET gaps with ANOVA significance are consistent evidence of lower translation quality on the authors' inclusive references. I therefore do not reject the paper. However, the reader's weakest assumption correctly identifies that the alternative-forms dictionary may not have been applied in scoring. I extend this concern: even with the dictionary applied, the conclusion that models systematically misunderstand singular 'they' requires semantic evidence, not just orthographic matching. Translating singular 'they' as 'ils' or using standard French generic forms is not necessarily a failure to understand the English source; it may reflect target-language convention. Since the paper's headline and Section 9 make a stronger causal claim, the authors should either report a human semantic evaluation of model outputs or soften the conclusion. These are addressable conditions, so the appropriate verdict remains CONDITIONAL rather than ACCEPT or REJECT. My agreement with the reader is partial because I share the scoring-transparency concern but see the semantic-validation issue as equally load-bearing.","tokens_in":18558,"tokens_out":6508,"duration_ms":69681,"concrete_test":"Re-run Table 4 and Figure 5 on a stratified sample of at least 200 inclusive-source sentences per model (or a random 200 total) using two independent scoring procedures: (1) apply the published alternative-forms dictionary so that all recognized inclusive variants are counted, and (2) have native French speakers, including users of inclusive writing, label each model output as inclusive-acceptable, standard-French-acceptable, wrong gendered form, or wrong-number/plural misunderstanding. If the dictionary-aware count and the human 'inclusive-acceptable' count remain below roughly 15%, the 'almost never' claim holds and the deeper conclusion is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim rests on two quantitative analyses: Table 4's count of inclusive indicators and Figure 5's pronoun-choice breakdown. Both operationalize 'inclusive correctness' through the authors' chosen French inclusive orthography, yet the paper never states that the alternative-forms dictionary from Section 4.3 was applied when counting. Table 4's indicator list is exact ('Iel', 'iel', 'Lea', 'lea', 'Un.e', 'un.e', 'Ce.tte', 'ce.tte' plus suffix endings), and the dictionary is not referenced in the counting protocol. Any valid variant such as 'un·e', 'la·le', 'iel·le', or 'un(e)' would therefore be counted as non-inclusive. With the raw list the best result is 86/806 (10.7%), so the 'almost never' conclusion could be overstated if variants were accepted. More importantly, Section 8.2 interprets the frequent translation of singular 'they' as 'ils' as evidence that the model 'has not recognized the singular usage of they.' This is an interpretive leap: French has no single standardized inclusive singular pronoun, and 'ils' or 'on' are common generic or plural renderings. The paper's own footnote admits the plural may sometimes be justified, but it does not quantify how many of the 246 examples are legitimate double interpretations. Consequently, the deeper claim that models 'fail to adequately understand inclusive constructs in English' (Section 9) is inferred from non-use of a prescriptive orthography rather than demonstrated through semantic evaluation. This is the load-bearing issue because if a human acceptability study showed that many outputs preserve the intended non-binary or singular reference in standard French, the paper's central condemnation of the models would be substantially weakened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FairTranslate, an English-French dataset of 2,418 sentence pairs centered on occupations, annotated with gender labels (male, female, inclusive), ambiguity type, and occupational stereotype. The English side is produced by back-translating French reference sentences, with all three gender variants per base example. The authors evaluate four LLMs (Gemma2-2B, Mistral-7B, Llama3.1-8B, Llama3.3-70B) under four prompting strategies, measuring BLEU, COMET, frequency of inclusive indicators (Table 4), occupation-form gender (Section 7), and pronoun choice for singular 'they' (Section 8.2). They report that inclusive forms score significantly lower than binary forms, that French inclusive indicators are rarely generated (0-86/806), and that models frequently translate singular 'they' as plural 'ils', leading to the conclusion that LLMs fail to understand established inclusive constructs in English.","tokens_in":18808,"tokens_out":6614,"duration_ms":58040,"significance":"The dataset is a valuable resource: it is publicly released, built with a counterfactual design that controls for sentence content across gender variants, and annotated with stereotype and ambiguity metadata. The finding that inclusive forms are translated worse than binary forms across all models and metrics is statistically robust (ANOVA p<1e-10). If the counting and interpretation issues identified below are resolved, the benchmark could support progress in gender-inclusive MT. The paper's strongest contribution is the dataset itself rather than the specific model rankings.","major_comments":[{"comment":"The counting of inclusive indicators uses a fixed list (\"Iel\", \"iel\", \"Lea\", \"lea\", \"Un.e\", \"un.e\", \"Ce.tte\", \"ce.tte\" plus suffix endings), but Section 4.3 introduces a Python dictionary to map the dataset's chosen forms to recognized alternatives such as 'ul' instead of 'iel' and 'un·e' instead of 'un.e'. The counting protocol in Section 8.1 does not state that this dictionary was applied. If it was not applied, valid inclusive outputs such as 'un·e', 'la·le', or 'iel·le' would be counted as non-inclusive, inflating the \"almost never\" result. Please clarify whether the dictionary was used in the counts reported in Table 4 and Figure 5; if it was not, re-run the analysis with the dictionary applied and report both raw and normalized counts.","section":"Section 8.1, Table 4"},{"comment":"The claim that translating singular 'they' as 'ils' shows the model \"has not recognized the singular usage of 'they'\" is an interpretive leap. French has no single standardized inclusive singular pronoun; 'ils' can function as a generic or grammatical plural rendering, and footnote 8 admits that a plural may be justified in some examples. The paper does not quantify how many of the 246 examples are legitimate double interpretations. To support the stronger conclusion in Section 9 that models \"fail to adequately understand inclusive constructs in English,\" the authors should provide a semantic audit, such as human judgments on a sample of the 246 source sentences to determine the intended referent number and human acceptability judgments on the French outputs. Without this, the \"misinterpretation as plural\" conclusion is not established.","section":"Section 8.2, footnote 8"},{"comment":"The English source sentences are back-translations from French, generated by GPT-4o from the French reference sentences and then human-verified. This means the English is constructed and may not reflect natural patterns of singular 'they' usage, including contexts where the referent is unambiguously singular. Since the paper's broader claim in Section 9 concerns models' handling of \"established inclusive constructs in English,\" the benchmark's ecological validity depends on the naturalness of the English side. The paper should either provide evidence that the back-translated English is natural (e.g., human naturalness ratings) or explicitly temper the claim to say the evaluation uses constructed stimuli. A brief discussion of this limitation would strengthen the paper.","section":"Section 4.2, Step 2"}],"minor_comments":[{"comment":"The abstract calls FairTranslate \"fully human-annotated,\" but Section 4.2 describes LLM-assisted sentence generation with human supervision; consider using \"human-verified\" or clarifying the annotation workflow.","section":"Abstract"},{"comment":"The phrases \"singular 'they' (person 3)\" and \"plural 'they' (person 6)\" are unclear; they should read \"third-person singular\" and \"third-person plural\".","section":"Section 2"},{"comment":"The interpretation of BLEU/COMET differences is thoughtful, but reporting effect sizes (e.g., partial eta-squared) alongside the ANOVA p-values would help readers gauge the magnitude of gender disparities.","section":"Section 6.4"},{"comment":"The linguistic prompting includes \"to be applied only if explicitly requested\" followed by the appended \"Otherwise, use the classic feminine or masculine form.\" This instruction may bias against inclusive forms; a sentence-level analysis of whether the appended sentence suppresses inclusive output would clarify the prompting effect.","section":"Section 5"},{"comment":"Consider reporting percentages rather than raw counts in the text; \"ranges between 0 and 86 out of 806\" is clear, but percentages would aid comparison across models and prompting conditions.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is solid and the core finding of lower inclusive-form performance is robust. The main risks are the unstated application of the alternatives dictionary in the indicator counts and the strong interpretive claim about singular 'they' without a semantic audit. Both are fixable with additional analysis. If the authors cannot show that the dictionary was applied, they should rerun the counts and temper the \"almost never\" wording accordingly. The paper would also benefit from acknowledging the back-translation limitation in the main text, not only implicitly through the construction description."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FairTranslate is worth knowing about: it's the first English-French MT benchmark that explicitly evaluates non-binary gender, built as counterfactual triples (male/female/inclusive) over 62 occupations, with human annotations and stereotype metadata. The dataset is released on Hugging Face, and the code is on GitHub. That alone is a solid contribution. The empirical pattern—inclusive translations score lower on BLEU and COMET, and models almost never emit inclusive French orthography even when prompted—is consistent across four models and statistically strong.\n\nThe soft spots are real but addressable. The biggest one is the leap from \"models don't produce 'iel' or midpoint forms\" to \"models fail to understand the singular 'they'.\" The paper's own footnote admits that the plural may sometimes be justified, but it never quantifies how many of the 246 singular-they examples have a legitimate double reading. And French inclusive orthography is not standardized; the alternatives dictionary in Section 4.3 is available, but the counting protocol for Table 4 doesn't state that it was applied. Variants like 'un·e' or 'iel·le' may have been counted as failures. That said, even the best raw count is 86/806 (~11%), so the \"almost never\" conclusion probably survives a generous rescoring—but the paper needs to show that.\n\nThe second soft spot is the interpretive claim. Translating singular 'they' as 'ils' could mean the model treated it as plural, but it could also be a generic-masculine strategy, which is a legitimate French convention for gender-neutral reference. Without a semantic audit (does the translation preserve the singular referent? is it acceptable to native speakers?), saying the model \"has not recognized the singular usage\" is stronger than the evidence supports.\n\nMinor issues: the English sources are back-translations from French rather than naturally occurring text, there are no commercial MT baselines, and the proportion analyses lack confidence intervals. All fixable.\n\nWho this is for: anyone building or evaluating gender-inclusive MT, and people who care about how to operationalize \"inclusive correctness\" in a morphologically gendered target language. It deserves a serious referee, and with a revised analysis it would be a useful published resource. I'd send it to review but ask the authors to (1) clarify the scoring protocol with and without the alternatives dictionary, (2) soften the \"misunderstands English\" conclusion or support it with a human acceptability study, and (3) add the missing baselines and error bars.","headline":"Useful dataset and a clear empirical pattern, but the strong 'models misunderstand singular they' conclusion needs a semantic audit and a clarified scoring protocol.","tokens_in":19411,"tokens_out":2638,"would_cite":true,"duration_ms":24839,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs almost never produce inclusive French, even when prompted","keywords":["gender bias","machine translation","inclusive language","non-binary gender","singular they","English-French","LLM evaluation","FairTranslate"],"falsifier":"Count inclusive French markers in model outputs while applying the alternative-forms dictionary to accept any recognized inclusive orthography; if the proportion of inclusive-labeled sentences with such markers exceeds a low threshold (e.g., 15%), the 'almost never' claim fails.","tokens_in":18297,"feed_emoji":"⚖️","tokens_out":7477,"duration_ms":60128,"temperature":0.7,"pith_summary":"This paper introduces FairTranslate, a human-annotated English–French dataset of 2,418 occupation-based sentence pairs, and uses it to evaluate four large language models under four prompting strategies. The central claim is that LLMs translate inclusive gender forms far worse than masculine or feminine forms: French inclusive indicators such as 'iel' and midpoint forms are almost never generated, even when the prompt explicitly requests them. More fundamentally, the paper argues that this failure is not just a matter of models being unfamiliar with recent French inclusive writing practices; models also misread long-established English inclusive constructs, particularly the singular 'they', translating it as plural instead of as a singular neutral form. If correct, FairTranslate provides a valid benchmark for non-binary gender bias in machine translation and demonstrates a prompting-resistant limitation of current LLMs.","feed_headline":"LLMs almost never produce inclusive French, even when prompted","feed_subtitle":"A 2,418-pair benchmark shows models default to masculine and misread singular 'they' as plural.","key_machinery":"The load-bearing mechanism is the FairTranslate dataset itself, built on a counterfactual design: every occupation-based sentence exists in three gender variants (male, female, inclusive) that differ only in gender markers, allowing model behavior to be compared across genders while controlling for content. Each sentence carries metadata for stereotype alignment, ambiguity type (ambiguous, unambiguous, long unambiguous), and a list of three French occupational forms (e.g., infirmier, infirmière, infirmier.ière). The central analytic device is the inclusive-indicator detection: the paper counts occurrences of inclusive markers (iel, lea, un.e, ce.tte, and midpoint endings such as ier.ère) in model outputs to measure how often inclusive forms are produced. A companion dictionary maps the chosen inclusive orthography to recognized alternatives so that valid inclusive forms other than the dataset's canonical ones are not penalized.","core_discovery":"The paper's core discovery is that the gender gap in LLM translation is systematic and prompting-resistant: across all four models, male translations receive the highest BLEU and COMET scores, female translations next, and inclusive translations lowest, with ANOVA p-values below $10^{-10}$; and inclusive French forms (iel, un.e, lea, midpoint occupational endings) appear in at most 86 of 806 inclusive-labeled sentences, and in most configurations fewer than 10. Prompting with moral or linguistic instructions improves inclusive output only slightly and at the cost of degrading binary translations. The deeper finding is that models predominantly translate the singular 'they' as the plural 'ils', indicating a failure to interpret the inclusive function of 'they' in English, which the authors argue underlies the poor inclusive output in French. The paper therefore claims that the challenge is not merely adapting to new French conventions but a fundamental representational gap in how LLMs handle established English inclusive constructs.","pith_inferences":["The dictionary's role in the Table 4 and Figure 5 counts is not explicit; if it was not applied, the 'almost never generated' result may overstate model failure for systems producing alternative recognized inclusive forms.","The counterfactual design could be ported to other gendered target languages, such as German or Spanish, to test whether the singular-'they' misinterpretation generalizes across typologically distinct languages.","The observed misreading of singular 'they' as plural suggests a representational issue that may also affect zero-shot translation into other languages with gendered pronouns, beyond French."],"forward_implications":["Current open LLMs are not reliable for inclusive English-to-French translation: without further intervention, inclusive forms are rarely produced, and prompting alone does not close the gap.","Translation quality for binary genders is also affected by prompting: moral and linguistic prompts improve inclusive output slightly but degrade masculine and feminine translations, so fairness interventions carry a cost.","Models' failure to interpret the singular 'they' as singular is a distinct error source, not merely a lack of French inclusive vocabulary: the dataset isolates this by labeling ambiguity and coreference distance.","The FairTranslate dataset can serve as a reusable benchmark for measuring non-binary gender bias in any English-to-French MT system, including future models.","Because each sentence has three gender variants, the dataset enables counterfactual evaluations that separate model bias from dataset bias."],"supporting_citations":[{"why":"The prior WinoMT benchmark that established binary gender bias evaluation in MT; this work extends that methodology beyond the binary.","marker":"[28]"},{"why":"Supplies the counterfactual dataset construction method (systematically modifying sentences) that FairTranslate adapts.","marker":"[9]"},{"why":"Provides the moral prompting instruction and the AmbGIMT non-binary evaluation approach that FairTranslate builds on.","marker":"[5]"},{"why":"The COMET neural metric used to measure semantic translation quality, which is robust to lexical variation and anchors the paper's gender-gap claims.","marker":"[26]"},{"why":"The BLEU n-gram metric used for surface-level translation quality in all experiments.","marker":"[23]"},{"why":"Documents the variety of French inclusive 'graphic tumult', grounding the paper's choice of one inclusive orthography among many.","marker":"[1]"},{"why":"A systematized inclusive French grammar that informs the dataset's inclusive forms (iel, un.e, lea, midpoint endings).","marker":"[3]"}],"fun_headline_variants":["Singular 'they' trips up LLMs in French translation","New benchmark shows LLMs ignore inclusive language in French","Prompting fails to fix LLMs' gender bias in French translation","LLMs default to masculine and plural in French, despite cues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation's ground truth rests on the authors' chosen French inclusive orthography (iel, un.e, lea, midpoint forms), and if the scoring did not accept other recognized inclusive forms via the provided dictionary, the 'almost never generated' result would overstate model failure.","fun_headline_variants_meta":{"raw":{"variants":["Singular 'they' trips up LLMs in French translation","New benchmark shows LLMs ignore inclusive language in French","Prompting fails to fix LLMs' gender bias in French translation","LLMs default to masculine and plural in French, despite cues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000509,"raw_usage":{"total_tokens":2499,"prompt_tokens":983,"completion_tokens":1516,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1446}},"tokens_in":599,"tokens_out":1516,"duration_ms":9613,"temperature":1.0,"reasoning_tokens":1446,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:13:42.557586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count inclusive French markers in model outputs while applying the alternative-forms dictionary to accept any recognized inclusive orthography; if the proportion of inclusive-labeled sentences with such markers exceeds a low threshold (e.g., 15%), the 'almost never' claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior WinoMT benchmark that established binary gender bias evaluation in MT; this work extends that methodology beyond the binary."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the counterfactual dataset construction method (systematically modifying sentences) that FairTranslate adapts."},{"cited_title":"Beyond Binary Gender: Evaluating Gender-Inclusive Machine Translation with Ambiguous Attitude Words","cited_arxiv_id":"2407.16266","evidence_quote":"Provides the moral prompting instruction and the AmbGIMT non-binary evaluation approach that FairTranslate builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the variety of French inclusive 'graphic tumult', grounding the paper's choice of one inclusive orthography among many."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A systematized inclusive French grammar that informs the dataset's inclusive forms (iel, un.e, lea, midpoint endings)."}],"review_version":1}