REVIEW 3 major objections 5 minor 78 references
Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read GAND, a new benchmark of 5,047 naturally occurring English sentences that are strictly gender-ambiguous for a singular referent, shows which contextual words steer an MT system's gendered translation and quantifies a masculine default.
desk verdict GAND is a genuinely useful new natural-data benchmark for gender-ambiguous MT, but its core strict-ambiguity guarantee rests on a single annotator's manual pass with no agreement measure — a real, fixable weakness, not a fatal one. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the GAND dataset itself—5,047 strictly gender-ambiguous English sentences—combined with a contrastive-translation interpretability pipeline. For each sentence, the model's original translation (target) is paired with a manually constructed translation that flips the referent's gender (foil); the contrastive gradient norm, computed from the difference in next-token probabilities for the two genders, assigns a saliency score to each source token. A threshold of 15% cumulative saliency selects the most influential words, and linguistic analysis uses part-of-speech tags and dependency distances to characterize them.
What would settle it
Independently re-annotate a random sample of the 5,047 GAND sentences with at least two annotators to check whether the referent is strictly gender-ambiguous; if a substantial fraction are judged to contain disambiguating cues—a name, a gender pronoun, or a stereotyped predicate—the core guarantee fails. Alternatively, run the same contrastive-attribution analysis on another open encoder-decoder model: if the masculine default and the local-context saliency pattern do not replicate, the findings are specific to OPUS-MT rather than general to neural MT.
Extended reading notes
Core claim
The paper's central claim is that GAND is the first large-scale, fully gender-ambiguous natural-language resource for machine translation, with every included sentence referring to a singular entity whose gender is not implied by the source text. Using a randomly selected 1,000-sentence subset translated from English into Spanish and German with the OPUS-MT models, the authors manually create contrastive translations that differ only in the gender of the referent, then compute saliency scores based on the contrastive probability difference. This reveals that the model's choice of masculine versus feminine translation is driven primarily by content words—nouns, verbs, adjectives, and proper n
Load-bearing premise
The load-bearing premise is that every GAND sentence really is strictly gender-ambiguous; this rests on a single annotator's manual vetting of 14,511 candidates, with no second annotator, so a hidden disambiguating cue in any included sentence would corrupt both the benchmark and the attribution findings built on it.
Editorial extensions
If this is right
- GAND provides a natural-data benchmark for evaluating how different machine translation systems and large language models handle gender ambiguity, moving beyond synthetic templates.
- The quantified masculine default—roughly 50–56% higher probability for masculine than feminine translations in this setup—gives a concrete baseline for measuring bias and the effect of mitigation strategies.
- Salient cues are predominantly content words within a dependency distance of one to three from the referent, suggesting that local syntactic context, not distant discourse, drives gender assignment in these models.
- The contrastive-attribution methodology can flag when a gendered translation is driven by salient stereotyped cues rather than genuine ambiguity, supporting diagnostic rather than normative uses.
- Preliminary intervention experiments—removing, masking, or flipping salient words—appear to change the translated gender, pointing toward a causal link between the identified cues and model output.
Reading between the lines
- If GAND's strict-ambiguity guarantee holds, it could become a shared testbed for gender-neutral translation strategies across many target languages; however, the single-annotator vetting means independent re-annotation is needed before relying on it as a gold standard.
- The saliency findings suggest a practical debiasing recipe: editing, masking, or counterfactually flipping high-saliency gendered content words may shift translations away from masculine defaults—this is a testable extension the paper only begins to explore.
- The roughly half-overlap in salient words between German and Spanish hints that the attribution pattern may be more model-driven than target-language-specific; testing other open encoder-decoder models would clarify whether this is a general property or specific to OPUS-MT.
- Because the resource is explicitly diagnostic rather than normative, it could support user-facing ambiguity warnings in translation interfaces, alerting readers when a gendered output is not grounded in the source.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces GAND, a dataset of 5047 English sentences claimed to be strictly gender-ambiguous with respect to a singular referent, compiled from C4 and OpenSubtitles using list-based filtering, rule-based cleaning, and manual verification. A 1000-sentence subset (GAND-CT) is translated into German and Spanish with OPUS-MT, manually extended with contrastive translations, and analyzed via gradient-based saliency attribution to identify source words that influence the model's choice of masculine versus feminine target gender. The authors report a masculine-default pattern (higher contrastive probability differences for masculine translations) and find that salient words are predominantly nouns, verbs, and adjectives at dependency distance 1–2 from the referent.
Significance. If the strict gender-ambiguity guarantee holds, GAND fills a genuine gap: existing benchmarks are largely synthetic or include unambiguous sentences, while GAND offers natural, linguistically varied English sentences for studying gender behavior in MT. The contrastive-attribution pipeline is a useful methodological template, and the public release of the dataset and scripts supports reproducibility and future benchmarking. The main risk is that the central ambiguity property is asserted from a single annotator's manual vetting, and the interpretability findings rely on threshold and statistical choices that are not yet fully validated. These concerns are addressable, and the resource itself is a solid contribution if they are resolved.
major comments (3)
- [§3.3, Appendix A.2.2] The central property of GAND—strict gender ambiguity for a singular referent—rests entirely on one annotator's manual verification of 14,511 candidates, with no inter-annotator agreement, second pass, or independent audit. The automatic rules in Table 10 target structural coreference with gender pronouns/proper nouns; they cannot remove non-pronominal disambiguating cues such as semantically gendered activities, titles, or world knowledge. Since 65.33% of manually checked candidates were rejected, the manual pass is doing substantial load-bearing work. If a non-negligible fraction of included sentences contain residual disambiguating cues, both the benchmark property and all downstream attribution findings inherit that error. Please report reliability (e.g., Cohen's κ on a sample annotated by a second annotator, or an independent audit with disagreement analysis), and ideally make the pe
- [§4.2, §4.4.2] The 15% cumulative-saliency threshold is adopted from Hackenbuchner et al. (2025c) with no sensitivity analysis in this paper. The counts and POS/dependency-distance distributions reported in §4.4.2 are properties of the word set selected by that threshold; a 10% or 20% threshold could change which words are considered salient and therefore affect the main interpretability conclusions. Please report robustness across a range of thresholds, or otherwise demonstrate that the findings are stable with respect to this choice.
- [§4.4.1, Table 3] The claim that the model shows higher contrastive probability differences for masculine than feminine translations is supported only by descriptive means. The feminine subsets are small (64 for EN→DE, 65 for EN→ES), and the reported standard deviations (~0.3) overlap substantially with the masculine means. Without confidence intervals, effect sizes, or a significance test, the quantitative comparison is not established. Similarly, the POS and dependency-distance comparisons against the overall distributions in §4.4.2 are not accompanied by any uncertainty quantification. Please add appropriate inferential statistics, or explicitly frame these as descriptive observations for the current sample.
minor comments (5)
- [§4.2 heading] Typo: 'Saliency Attibution' should be 'Saliency Attribution'.
- [Table 3] The counts for EN→DE (893 + 64 + 41 = 998) and EN→ES (814 + 65 + 120 = 999) do not sum to 1000. Please clarify whether some sentences were excluded for another reason or whether these are rounding artifacts.
- [§4.4.1] The sentence beginning 'More frequently than in masculine scenarios, the CPD of an original feminine target was negative...' is difficult to parse. It would be clearer to state explicitly that a negative CPD means the model assigned higher probability to the foil than to the observed target in that contrastive pair.
- [Figure 3 caption] The caption references 'red vertical lines' and 'horizontal red lines' without explaining in the caption what the red lines represent; please add this information to the figure caption or legend.
- [§3.1, Table 4] The male/female/neutral labels for the 183 referents are derived from word embeddings and an LLM, which encode societal stereotypes. Using these labels to define 'matches' and 'mismatches' in §4.4.1 is descriptively reasonable, but the wording risks reifying the labels as ground truth. Please clarify that these are association-based labels, not objective gender categories.
Circularity Check
No significant circularity: GAND's construction is independent of the tested model, and the attribution analysis is a direct measurement; only a minor same-team methodological self-citation appears.
full rationale
The paper's central resource claim—that GAND contains 5047 naturally sourced English sentences that are strictly gender-ambiguous for a singular referent—is not derived from the model being analyzed. Sentences are filtered from C4 and OpenSubtitles using automated rules and manual vetting (Sections 3.1–3.3); the ambiguity property is an annotation decision, not a model output. The interpretability analysis is likewise a direct measurement: OPUS-MT translations are contrasted with manually created gender-contrastive translations, and saliency is computed via the standard contrastive gradient norm of Yin and Neubig (2022). The finding that nouns, verbs, and adjectives near the referent are salient is an empirical result of that computation, not an input to it. The only same-team self-citation that could be flagged is the 15% cumulative-saliency threshold and the preprocessing choices taken from Hackenbuchner et al. (2025c). This is a methodological hyperparameter inherited from prior work, explicitly described as replaceable ('Future work could include analyses based on alternative thresholds'), and the main conclusions do not reduce to that threshold. A substantive correctness risk, but not a circularity, is the single-annotator manual verification of gender ambiguity without inter-annotator agreement; if label noise exists, downstream attribution findings inherit it, but this is a validity concern rather than a derivation that assumes its own conclusion. Overall, the paper contains no step where a 'prediction' is equivalent by construction to a fitted input or where the central result is forced by a self-citation chain.
Assumptions & free parameters
free parameters (5)
- 183-referent gender-association list =
72 male-associated, 57 female-associated, 54 neutral
- Sentence length cutoffs =
min 5 tokens, max 50 tokens
- Coreference exclusion rules =
10 handcrafted rules plus manually compiled gender (pro)noun list
- Saliency threshold =
15% cumulative attribution (top-15% subset)
- Token removal set for attribution =
target referent token, EOS, punctuation, and {a, an, the, this, that, these, those}
assumptions (6)
- domain assumption English sentences containing one of 183 hand-selected referent nouns, after automatic filtering and single-annotator manual vetting, are strictly gender-ambiguous with respect to that referent.
- domain assumption The gender association labels for referent nouns (masculine/feminine/neutral) derived from prior word-embedding studies and ChatGPT-4 reflect meaningful genderedness of the referent.
- domain assumption Saliency attribution (gradient-norm contrastive explanations from inseq) reflects the contextual cues that actually inform the model's gender choice.
- domain assumption Manually created contrastive translations differ from the original only in gender-marked target tokens, so the attribution contrast isolates gender.
- ad hoc to paper The 15% cumulative-saliency threshold yields plausible salient cues.
- domain assumption Stanza and spaCy dependency parses are accurate enough for coreference exclusion checks and dependency-distance measurements.
Cite this review
Pith. "Pith review of Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution." pith.science (2026). https://pith.science/paper/YCG4FOP5
@misc{pith2026260722546,
author = {Pith},
title = {Pith review of: Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution},
year = {2026},
howpublished = {\url{https://pith.science/paper/YCG4FOP5}},
note = {Machine review of arXiv:2607.22546}
}
read the original abstract
Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenarios in a natural way. To this end, we present GAND, a gender-ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations. A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Attanasio, Giuseppe and Plaza del Arco, Flor Miriam and Nozza, Debora and Lauscher, Anne. A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.243
-
[2]
Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the G e NTE Corpus
Piergentili, Andrea and Savoldi, Beatrice and Fucci, Dennis and Negri, Matteo and Bentivogli, Luisa. Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the G e NTE Corpus. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023.emnlp-main.873
-
[3]
Interpreting Language Models with Contrastive Explanations
Yin, Kayo and Neubig, Graham. Interpreting Language Models with Contrastive Explanations. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. doi:10.18653/v1/2022.emnlp-main.14
-
[4]
The elephant in the interpretability room:
Bastings, Jasmijn and Filippova, Katja , year =. The elephant in the interpretability room:. Proceedings of the. doi:10.18653/v1/2020.blackboxnlp-1.14 , language =
-
[5]
Yin, Kayo and Fernandes, Patrick and Pruthi, Danish and Chaudhary, Aditi and Martins, André F. T. and Neubig, Graham , year =. Do. Proceedings of the 59th. doi:10.18653/v1/2021.acl-long.65 , language =
-
[6]
Bastings, Jasmijn and Ebert, Sebastian and Zablotskaia, Polina and Sandholm, Anders and Filippova, Katja , year =. “. Proceedings of the 2022. doi:10.18653/v1/2022.emnlp-main.64 , language =
-
[7]
Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation
Manna, Chiara and Alishahi, Afra and Blain, Fr \'e d \'e ric and Vanmassenhove, Eva. Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation. Proceedings of the 3rd Workshop on Gender-Inclusive Translation Technologies (GITT 2025). 2025
2025
-
[8]
Inseq: An Interpretability Toolkit for Sequence Generation Models
Sarti, Gabriele and Feldhus, Nils and Sickert, Ludwig and van der Wal, Oskar. Inseq: An Interpretability Toolkit for Sequence Generation Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 2023. doi:10.18653/v1/2023.acl-demo.40
Show all 78 references
-
[9]
A decade of gender bias in machine translation
Savoldi, Beatrice and Bastings, Jasmijn and Bentivogli, Luisa and Vanmassenhove, Eva. A decade of gender bias in machine translation. Patterns. 2025. doi:10.1016/j.patter.2025.101257
2025
-
[10]
2024 , eprint=
A Primer on the Inner Workings of Transformer-based Language Models , author=. 2024 , eprint=
2024
-
[11]
Getting Gender Right in Neural Machine Translation
Vanmassenhove, Eva and Hardmeier, Christian and Way, Andy. Getting Gender Right in Neural Machine Translation. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. doi:10.18653/v1/D18-1334
2018 doi
-
[12]
Analysing Neural Language Models: Contextual Decomposition Reveals Default Reasoning in Number and Gender Assignment
Jumelet, Jaap and Zuidema, Willem and Hupkes, Dieuwke. Analysing Neural Language Models: Contextual Decomposition Reveals Default Reasoning in Number and Gender Assignment. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019. doi:10.1865...
2019 doi
-
[13]
First the Worst: Finding Better Gender Translations During Beam Search
Saunders, Danielle and Sallis, Rosie and Byrne, Bill. First the Worst: Finding Better Gender Translations During Beam Search. Findings of the Association for Computational Linguistics: ACL 2022. 2022. doi:10.18653/v1/2022.findings-acl.301
2022 doi
-
[14]
2023 , eprint=
Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks , author=. 2023 , eprint=
2023
-
[15]
The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
Conti, Lina and Fucci, Dennis and Gaido, Marco and Negri, Matteo and Wisniewski, Guillaume and Bentivogli, Luisa. The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models. Proceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Network...
2025 doi
-
[16]
https://aclanthology.org/2020.eamt-1.61/
J. Proceedings of the 22nd Annual Conferenec of the European Association for Machine Translation (EAMT) , url="https://aclanthology.org/2020.eamt-1.61/", year =
2020
-
[17]
Gender Bias and the Role of Context in Human Perception and Machine Translation
Hackenbuchner, Jani c a and Tezcan, Arda and Daems, Joke. Gender Bias and the Role of Context in Human Perception and Machine Translation. Computational Linguistics in the Netherlands Journal. 2025
2025
-
[18]
Automatic detection of (potential) factors in the source text leading to gender bias in machine translation
Hackenbuchner, Jani c a and Tezcan, Arda and Daems, Joke. Automatic detection of (potential) factors in the source text leading to gender bias in machine translation. Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 2). 2024
2024
-
[19]
James Murdoch and Chandan Singh and Karl Kumbier and Reza Abbasi-Asl and Bin Yu
W. James Murdoch and Chandan Singh and Karl Kumbier and Reza Abbasi-Asl and Bin Yu. Definitions, methods, and applications in interpretable machine learning. Proceedings of the National Academy of Sciences. 2019. doi:10.1073/pnas.1900654116. https://www.pnas.org/doi/pdf/10.107...
2019 doi
-
[20]
Contrastive Explanation
Lipton, Peter. Contrastive Explanation. Royal Institute of Philosophy Supplement. 1990. doi:10.1017/S1358246100005130
1990 doi
-
[21]
and Garcke, Jochen , journal=
Roscher, Ribana and Bohn, Bastian and Duarte, Marco F. and Garcke, Jochen , journal=. Explainable Machine Learning for Scientific Insights and Discoveries , year=
-
[22]
Biran, Or, and Courtenay V. Cotton. Explanation and Justification in Machine Learning: A Survey. Proceedings of the IJCAI-17 Workshop on Explainable Artificial Intelligence (XAI). 2017
2017
-
[23]
Interpretable Machine Learning , subtitle=
Christoph Molnar , year=. Interpretable Machine Learning , subtitle=
-
[24]
Proceedings of The Fourth Widening Natural Language Processing Workshop , publisher=
Towards Mitigating Gender Bias in a decoder-based Neural Machine Translation model by Adding Contextual Information , author=. Proceedings of The Fourth Widening Natural Language Processing Workshop , publisher=
-
[25]
, month = jul, year =
Caliskan, Aylin and Ajay, Pimparkar Parth and Charlesworth, Tessa and Wolfe, Robert and Banaji, Mahzarin R. , month = jul, year =. Gender. Proceedings of the 2022. doi:10.1145/3514094.3534162 , language =
2022
-
[26]
https://dl.acm.org/doi/10.5555/3157382.3157584
Man is to. CoRR , url="https://dl.acm.org/doi/10.5555/3157382.3157584", author =
- [27]
-
[28]
Proceedings of the 5th Conference on Machine Translation (WMT) , publisher=
Gender Coreference and Bias Evaluation at WMT 2020 , author=. Proceedings of the 5th Conference on Machine Translation (WMT) , publisher=
2020
-
[29]
Collecting a
Levy, Shahar and Lazar, Koren and Stanovsky, Gabriel , year =. Collecting a. Findings of the. doi:10.18653/v1/2021.findings-emnlp.211 , language =
2021 doi
-
[30]
Gendered Technology in Translation and Interpreting , pages=
Gender bias in machine translation and the era of large language models , author=. Gendered Technology in Translation and Interpreting , pages=. 2024 , publisher=
2024
-
[31]
g EN der- IT : An Annotated E nglish- I talian Parallel Challenge Set for Cross-Linguistic Natural Gender Phenomena
Vanmassenhove, Eva and Monti, Johanna. g EN der- IT : An Annotated E nglish- I talian Parallel Challenge Set for Cross-Linguistic Natural Gender Phenomena. Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing. 2021. doi:10.18653/v1/2021.gebnlp-1.1
2021 doi
-
[32]
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with m G e NTE
Savoldi, Beatrice and Attanasio, Giuseppe and Cupin, Eleonora and Gkovedarou, Eleni and Hackenbuchner, Jani c a and Lauscher, Anne and Negri, Matteo and Piergentili, Andrea and Thind, Manjinder and Bentivogli, Luisa. Mind the Inclusivity Gap: Multilingual Gender-Neutral Transl...
2025 doi
-
[33]
2025 , eprint=
What Triggers my Model? Contrastive Explanations Inform Gender Choices by Translation Models , author=. 2025 , eprint=
2025
-
[34]
Gender bias amplification during Speed-Quality optimization in Neural Machine Translation
Renduchintala, Adithya and Diaz, Denise and Heafield, Kenneth and Li, Xian and Diab, Mona. Gender bias amplification during Speed-Quality optimization in Neural Machine Translation. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the...
2021 doi
-
[35]
M i TT en S : A Dataset for Evaluating Gender Mistranslation
Robinson, Kevin and Kudugunta, Sneha and Stella, Romina and Dev, Sunipa and Bastings, Jasmijn. M i TT en S : A Dataset for Evaluating Gender Mistranslation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp...
2024 doi
-
[36]
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Zhao, Jieyu and Wang, Tianlu and Yatskar, Mark and Ordonez, Vicente and Chang, Kai-Wei. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: ...
2018 doi
-
[37]
Automatic Gender Identification and Reinflection in A rabic
Habash, Nizar and Bouamor, Houda and Chung, Christine. Automatic Gender Identification and Reinflection in A rabic. Proceedings of the First Workshop on Gender Bias in Natural Language Processing. 2019. doi:10.18653/v1/W19-3822
2019 doi
-
[38]
2023 , isbn =
Rarrick, Spencer and Naik, Ranjita and Mathur, Varun and Poudel, Sundar and Chowdhary, Vishal , title =. 2023 , isbn =. doi:10.1145/3600211.3604675 , booktitle =
2023
-
[39]
Analyzing Gender Translation Errors to Identify Information Flows between the Encoder and Decoder of a NMT System
Wisniewski, Guillaume and Zhu, Lichao and Ballier, Nicolas and Yvon, Fran c ois. Analyzing Gender Translation Errors to Identify Information Flows between the Encoder and Decoder of a NMT System. Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neura...
2022 doi
-
[40]
Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms
Mastromichalakis, Orfeas Menis and Filandrianos, Giorgos and Symeonaki, Maria and Stamou, Giorgos. Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms. Proceedings of the 2025 Conference on Empirical Methods in Natural Lang...
2025 doi
-
[41]
2025 , eprint=
Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation , author=. 2025 , eprint=
2025
-
[42]
Gender Bias in Machine Translation
Savoldi, Beatrice and Gaido, Marco and Bentivogli, Luisa and Negri, Matteo and Turchi, Marco. Gender Bias in Machine Translation. Transactions of the Association for Computational Linguistics. 2021. doi:10.1162/tacl_a_00401
2021 doi
-
[43]
Language (
Blodgett, Su Lin and Barocas, Solon and Daumé III, Hal and Wallach, Hanna , editor =. Language (. Proceedings of the 58th. 2020 , pages =. doi:10.18653/v1/2020.acl-main.485 , urldate =
2020 doi
-
[44]
Context-Aware Neural Machine Translation Learns Anaphora Resolution
Voita, Elena and Serdyukov, Pavel and Sennrich, Rico and Titov, Ivan. Context-Aware Neural Machine Translation Learns Anaphora Resolution. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. doi:10.18653/v1/P18-1117
2018 doi
-
[45]
Quantifying the
Sarti, Gabriele and Chrupał a, Grzegorz and Nissim, Malvina and Bisazza, Arianna , editor =. Quantifying the. International. 2024 , pages =
2024
-
[46]
Sennrich, Rico , year =. How. Proceedings of the 15th. doi:10.18653/v1/E17-2060 , language =
-
[47]
Improving
Rios Gonzales, Annette and Mascarell, Laura and Sennrich, Rico , year =. Improving. Proceedings of the. doi:10.18653/v1/W17-4702 , language =
-
[48]
, author=
Measuring nominal scale agreement among many raters. , author=. Psychological bulletin , volume=. 1971 , publisher=
1971
-
[49]
Educational and psychological measurement , volume=
A coefficient of agreement for nominal scales , author=. Educational and psychological measurement , volume=. 1960 , publisher=
1960
-
[50]
You Shall Know a Word ' s Gender by the Company it Keeps: Comparing the Role of Context in Human Gender Assumptions with MT
Hackenbuchner, Jani c a and Daems, Joke and Tezcan, Arda and Maladry, Aaron. You Shall Know a Word ' s Gender by the Company it Keeps: Comparing the Role of Context in Human Gender Assumptions with MT. Proceedings of the 2nd International Workshop on Gender-Inclusive Translati...
2024
-
[51]
Proceedings of the AAAI Conference on Artificial Intelligence , author =
Interpreting. Proceedings of the AAAI Conference on Artificial Intelligence , author =. 2022 , pages =. doi:10.1609/aaai.v36i11.21442 , language =
2022 doi
- [52]
-
[53]
Contrastive
Vamvas, Jannis and Sennrich, Rico , year =. Contrastive. Proceedings of the 2021. doi:10.18653/v1/2021.emnlp-main.803 , language =
2021 doi
-
[54]
Jacovi, Alon and Goldberg, Yoav , year =. Towards. Proceedings of the 58th. doi:10.18653/v1/2020.acl-main.386 , language =
2020 doi
-
[55]
ACM Computing Surveys , author =
Post-hoc. ACM Computing Surveys , author =. 2022 , pages =. doi:10.1145/3546577 , language =
2022 doi
-
[56]
Implementing Gender-Inclusivity in MT Output using Automatic Post-Editing with LLM s
Nunziatini, Mara and Diego, Sara. Implementing Gender-Inclusivity in MT Output using Automatic Post-Editing with LLM s. Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1). 2024
2024
-
[57]
N eu T ral R ewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives
Vanmassenhove, Eva and Emmery, Chris and Shterionov, Dimitar. N eu T ral R ewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18...
2021 doi
-
[58]
A Rewriting Approach for Gender Inclusivity in P ortuguese
Veloso, Leonor and Coheur, Luisa and Ribeiro, Rui. A Rewriting Approach for Gender Inclusivity in P ortuguese. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.585
2023 doi
-
[59]
Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model
Amrhein, Chantal and Schottmann, Florian and Sennrich, Rico and L. Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/20...
2023 doi
-
[60]
Gender, names and other mysteries: Towards the ambiguous for gender-inclusive translation
Saunders, Danielle and Olsen, Katrina. Gender, names and other mysteries: Towards the ambiguous for gender-inclusive translation. Proceedings of the First Workshop on Gender-Inclusive Translation Technologies. 2023
2023
-
[61]
Cho, Won Ik and Kim, Ji Won and Kim, Seok Min and Kim, Nam Soo , year =. On. Proceedings of the. doi:10.18653/v1/W19-3824 , language =
-
[62]
Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale
Costa-juss \`a , Marta and Andrews, Pierre and Smith, Eric and Hansanti, Prangthip and Ropers, Christophe and Kalbassi, Elahe and Gao, Cynthia and Licht, Daniel and Wood, Carleigh. Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in L...
2023 doi
-
[63]
and Cattoni, Roldano and Turchi, Marco
Bentivogli, Luisa and Savoldi, Beatrice and Negri, Matteo and Di Gangi, Mattia A. and Cattoni, Roldano and Turchi, Marco. Gender in Danger? Evaluating Speech Translation Technology on the M u ST - SHE Corpus. Proceedings of the 58th Annual Meeting of the Association for Comput...
2020 doi
-
[64]
Reducing
Saunders, Danielle and Byrne, Bill , year =. Reducing. Proceedings of the 58th. doi:10.18653/v1/2020.acl-main.690 , language =
2020 doi
-
[65]
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics , volume=
Savoldi, Beatrice and Piergentili, Aandrea and Fucci, Dennis and Negri, Matteo and Bentivogli, Luisa , title=. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics , volume=
-
[66]
GAMBIT +: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics
Filandrianos, George and Menis Mastromichalakis, Orfeas and Mohammed, Wafaa and Attanasio, Giuseppe and Zerva, Chrysoula. GAMBIT +: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics. Proceedings of the Tenth Conference on Machine Tran...
2025 doi
-
[67]
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing) , publisher=
Prequel: Quality estimation of machine translation outputs in advance , author=. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing) , publisher=
2022
-
[68]
Proceedings of the Fourth Workshop on Discourse in Machine Translation (DiscoMT) , publisher=
When and Why is Document-level Context Useful in Neural Machine Translation? , author=. Proceedings of the Fourth Workshop on Discourse in Machine Translation (DiscoMT) , publisher=
-
[69]
and Zettlemoyer, Luke , year =
Stanovsky, Gabriel and Smith, Noah A. and Zettlemoyer, Luke , year =. Evaluating. Proceedings of the 57th. doi:10.18653/v1/P19-1164 , language =
-
[70]
Gender bias and stereotypes in
Kotek, Hadas and Dockum, Rikker and Sun, David , month = nov, year =. Gender bias and stereotypes in. Proceedings of. doi:10.1145/3582269.3615599 , language =
-
[71]
and Casas, Noe , year =
Basta, Christine and Costa-jussà, Marta R. and Casas, Noe , year =. Evaluating the. Proceedings of the. doi:10.18653/v1/W19-3805 , language =
-
[72]
MT - G en E val: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation
Currey, Anna and Nadejde, Maria and Pappagari, Raghavendra Reddy and Mayer, Mia and Lauly, Stanislas and Niu, Xing and Hsu, Benjamin and Dinu, Georgiana. MT - G en E val: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation. Proceedings...
2022 doi
-
[73]
Sewunetie, Walelign Tewabe and Tonja, Atnafu Lambebo and Belay, Tadesse Destaw and Nigatu, Hellina Hailu and Kidanu, Gashaw and Mossie, Zewdie and Seid, Hussien and Yimam, Seid Muhie , year=. Gender. Proceedings of the 2nd International Workshop on Gender-Inclusive Translation...
-
[74]
Gkovedarou, Eleni and Daems, Joke and Bruyne, Luna De , year =. Gender. Proceedings of the 3rd Workshop on Gender-Inclusive Translation Technologies (GITT 2025) , publisher =
2025
-
[75]
Gender Neutralization for an Inclusive Machine Translation: from Theoretical Foundations to Open Challenges
Piergentili, Andrea and Fucci, Dennis and Savoldi, Beatrice and Bentivogli, Luisa and Negri, Matteo. Gender Neutralization for an Inclusive Machine Translation: from Theoretical Foundations to Open Challenges. Proceedings of the First Workshop on Gender-Inclusive Translation T...
2023
-
[76]
Perspectives , author =
Gender-fair translation: a case study beyond the binary , volume =. Perspectives , author =. 2024 , pages =. doi:10.1080/0907676X.2023.2268654 , language =
2024
-
[77]
Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair G erman Machine Translation
Pranav, A and Hackenbuchner, Jani c a and Attanasio, Giuseppe and Lardelli, Manuel and Lauscher, Anne. Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair G erman Machine Translation. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025....
2025 doi
- [78]
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.