REVIEW 3 major objections 5 minor 59 references
Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes GIFI, a single 0-100 score that averages seven fairness metrics, and uses it to rank 22 large language models on gender inclusivity, finding that all models handle non-binary pronouns worse than binary ones.
desk verdict Useful benchmark with a real neopronoun finding, but the SA/OF metric makes the headline GIFI ranking partly an artifact; worth refereeing with major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the GIFI composite, defined as the average of seven sub-metrics, each constructed to fall in [0,1] with higher meaning fairer. GDR and PE use $1/(1+CV)$, where $CV$ is the coefficient of variation of per-pronoun accuracies; SN and NTS use $1 - \text{Average MAD}$ of sentiment or toxicity scores across pronoun-swapped prompt variants; CF counts the fraction of counterfactual output pairs whose embedding cosine similarity stays above a threshold; and SA and OF measure the squared deviation of generated pronoun distributions from a uniform target on stereotype and occupation templates. The uniform target and the exclusion of singular 'they' from SA and OF are the normative choices that make the scores comparable across models.
What would settle it
Run a community-grounded audit in which transgender and non-binary users rate model completions for respect and misgendering, then check whether the GIFI rankings correlate with those ratings; if they do not, or if the SA and OF rankings flip when singular 'they' is included or when the uniform target is replaced by a base-rate target estimated from real pronoun usage, the index is measuring something other than gender inclusivity.
Extended reading notes
Core claim
The central claim is that gender inclusivity in an LLM can be compressed into the GIFI score, the unweighted average of seven normalized metrics: Gender Diversity Representation, Sentiment Neutrality, Non-Toxicity Score, Counterfactual Fairness, Stereotypical Association, Occupational Fairness, and Performance Equality. The authors argue the resulting rankings are meaningful: GPT-4o (73), Claude 3 (71), and DeepSeek V3 (70) lead, while Vicuna (49), GPT-2 (55), and LLaMA 2 (56) trail. They also claim the most consistent finding across all 22 models is that neopronouns are handled far worse than binary pronouns, with no model spontaneously generating neopronouns in stereotype or occupation prompts, and with 'they' rarely exceeding 30 percent of completions.
Load-bearing premise
The SA and OF metrics assume that the fair output is a uniform distribution of pronouns across the ten non-'they' groups, and they drop singular 'they' from that computation because it would 'largely skew the results'; if that normative target is wrong, the stereotype and occupation scores, and therefore the composite GIFI, would change.
Editorial extensions
If this is right
- Any LLM can be assigned a 0-100 GIFI score, making cross-model gender-inclusivity comparisons direct.
- Non-binary and neopronoun handling is the weakest axis for every model tested, so progress on binary gender fairness should not be mistaken for inclusivity.
- High Performance Equality on math reasoning mostly tracks overall reasoning ability: strong models are fair because they solve every variant, while weak models fail every variant.
- The framework's prompt sets and pronoun lists are extensible, so new pronouns or new task types such as coding and planning can be added without redesigning the index.
- Current models over-generate 'she' in female-stereotyped contexts and 'he' in male-stereotyped contexts, while neopronouns are never spontaneously generated.
Reading between the lines
- The uniform-distribution target baked into SA and OF may penalize models whose outputs track real-world pronoun base rates; a base-rate-adjusted variant would test whether the stereotype rankings survive.
- Because singular 'they' is excluded from SA and OF, models that heavily overuse 'they' are shielded from a penalty that could change their GIFI rank; re-inserting 'they' is a direct sensitivity test.
- Equal weighting of the seven sub-metrics is an arbitrary choice; weight perturbation could reveal whether the reported model ordering is stable or driven by one or two components.
- The paper notes intersectionality is unexplored; a natural extension is to condition the same seven probes on race, disability, or other demographic markers and look for compounding effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Gender Inclusivity Fairness Index (GIFI), a composite score defined as the arithmetic mean of seven metrics (GDR, SN, NTS, CF, SA, OF, PE) that probe LLM behavior under eleven pronoun groups (binary he/she, singular they, and eight neopronoun families). The authors evaluate 22 open and proprietary LLMs, report a GIFI ranking (GPT-4o 73, Vicuna 49), and conclude that binary pronouns are handled far better than neopronouns. The code and data are released, and appendices include qualitative error analyses and a classifier ablation.
Significance. The paper's most robust empirical contribution is the consistent finding that models recognize and generate binary pronouns far more reliably than neopronouns, a trend that appears across GDR, PE, and the substitution analysis in Appendix D. The benchmark's breadth (22 models, multiple metric families, released code/data) is useful for future fairness evaluations, and the VADER ablation in Appendix C is a good robustness check. That said, GIFI's value as a headline index is weakened by the fact that two of its seven components (SA and OF) are computed over categories that are never generated, and by the absence of uncertainty intervals for the rankings. The equal weighting of the seven components is an axiomatic choice, not an empirically justified one.
major comments (3)
- [§3.3, Footnote 6; Table E.6; §6] The SA/OF formula in §3.3 (1 − (1/M)Σ_m Σ_g (O_mg − 1/G)^2) is internally inconsistent with the stated exclusion of 'they' and with the reported scores. Under the declared G=10 after Footnote 6 excludes 'they', and with the paper's own observation in §6 that neopronouns are entirely absent across all models in these tasks, eight of the ten O_mg terms are identically zero. The metric is then a distance to a uniform distribution over categories that receive zero probability mass, its [0,1] range is unattainable, and the score does not primarily measure stereotypical or occupational association. Reported values such as Phi-3 SA=0.72 (Table E.6) are only explicable if 'they' is actually retained as a target category with G=11, so the computation behind Table E.6 does not match the text. Because SA and OF are two of seven equally weighted GIFI components, the ranking in Figure 2a is affected. The authors must clarify the actual computation and either re-define SA/OF over the support that can be generated (e.g., using only pronoun families that appear in the completions, with an explicit treatment of absent families) or provide a sensitivity analysis that recomputes GIFI without these two components.
- [Figure 2a; Table E.6; §5] The headline GIFI rankings are point estimates without uncertainty intervals or significance tests, although generation uses stochastic decoding (temperature 0.95, top-p 0.95, §5) and several component metrics are averaged over four runs. Adjacent scores differ by 1–2 points (e.g., Claude 3 and GPT-4o-mini at 71 and 68; multiple models at 67), so the abstract's claim of 'significant variations' in gender inclusivity is not statistically supported. The Limitations section explicitly defers the computation of 'average fairness indexes and significance intervals' to future work; since the ranking is the central result, this uncertainty quantification should be included in the present paper.
- [§3.4, Eq. (1); Table E.6; Appendix D.4] PE is defined as the coefficient-of-variation-based consistency of accuracy across pronoun groups, so a model that fails all math variants for every pronoun receives a high PE score. The paper itself acknowledges in Appendix D.4 that 'a model can be uniformly correct—or uniformly incorrect' and that weak models' failures are due to task difficulty rather than pronoun bias. Since PE is one of seven equally weighted GIFI components, the composite index partially rewards uniformly poor performance and therefore conflates overall reasoning capability with gender fairness. The authors should condition PE on mean accuracy (e.g., PE × Acc) or exclude it from GIFI and report both versions.
minor comments (5)
- [§6, Figure 5 discussion] The text describing Claude 4 ('skews heavily towardshe... overhe... andthey...') and LLaMA 4 ('disproportionately favorshe') has missing quotes and garbled comparisons; these phrases should be corrected to read 'towards he', 'over she', and 'favors she'.
- [Figure D.10 caption] The caption reads 'Pronoun generation bias by cccupation'; 'cccupation' should be 'occupation'.
- [§3.3] The symbol G is used both for the number of pronoun groups and for the number of repeated generations ('we collect model generation G times'); rename one of them (e.g., T for repetitions) to avoid confusion.
- [Appendix B / E] The number of generation runs is specified only for GDR and PE (4 runs); please specify the number of runs for SN, NTS, CF, SA, and OF to support reproducibility.
- [§3.5] The statement that 'all the included metrics lie in the range of [0,1]' is formally true but misleading for SA/OF, whose attainable range under the current protocol is much narrower; this should be corrected once the metric definition is fixed.
Circularity Check
No significant circularity: GIFI is a defined aggregate index whose components are computed from model outputs, not fitted inputs or self-citation chains.
full rationale
GIFI is introduced as a defined aggregate, not as a quantity derived from hidden assumptions: the paper states 'GIFI... is a single number that averages across all axes of evaluation (GDR, SN, NTS, CF, SA, OF, PE) for easy interpretation.' Each component has an explicit formula (Eq. 1 for GDR and PE; MAD-based definitions for SN and NTS; a cosine-similarity proportion for CF; a squared deviation from the uniform distribution for SA and OF), and none of these formulas is fitted to the evaluated models and then relabeled as a prediction. The one tunable constant, the CF threshold gamma=0.3, is declared configurable rather than inferred from outcomes. Citations to prior work (Lauscher et al., Hossain et al., Ovalle et al., Dong et al.) supply external datasets and pronoun inventories; they are not self-citations of the present authors, and no uniqueness theorem is invoked to force the framework's choice of pronouns or metric. The SA/OF support mismatch (neopronouns are never generated, so with G=10 after excluding 'they' the attainable ceiling is approximately 1 - [3(1/3 - 0.1)^2 + 7(0.1)^2] = 0.767) is a genuine construct-validity limitation that compresses scores and may distort rankings, but it is not a circular step: the score is still computed from observed model outputs, and the ranking does not reduce by construction to the metric's own inputs. The paper's Limitations section candidly flags external-classifier bias, data contamination, reproducibility, and incomplete metrics, all of which are validity concerns rather than circularity. There is no derivation chain in which an output quantity is identical to an input by definition or in which a fitted parameter is renamed as a prediction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- cosine similarity threshold gamma for counterfactual fairness =
0.3
assumptions (5)
- domain assumption RoBERTa sentiment and Perspective API toxicity scores are valid proxies for sentiment and toxicity of model outputs.
- ad hoc to paper Uniform pronoun distribution is the fairness target for stereotype and occupational association.
- ad hoc to paper Excluding the 'they' pronoun group from SA/OF does not bias the fairness estimate.
- domain assumption Performance on GSM8K math problems with substituted pronouns is a valid measure of intrinsic gender bias.
- ad hoc to paper The GIFI average over seven metrics is a meaningful composite of gender inclusivity.
Cite this review
Pith. "Pith review of Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models." pith.science (2026). https://pith.science/paper/UZ4CD7WY
@misc{pith2026250615568,
author = {Pith},
title = {Pith review of: Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/UZ4CD7WY}},
note = {Machine review of arXiv:2506.15568}
}
read the original abstract
We present a comprehensive evaluation of gender fairness in large language models (LLMs), focusing on their ability to handle both binary and non-binary genders. While previous studies primarily focus on binary gender distinctions, we introduce the Gender Inclusivity Fairness Index (GIFI), a novel and comprehensive metric that quantifies the diverse gender inclusivity of LLMs. GIFI consists of a wide range of evaluations at different levels, from simply probing the model with respect to provided gender pronouns to testing various aspects of model generation and cognitive behaviors under different gender assumptions, revealing biases associated with varying gender identifiers. We conduct extensive evaluations with GIFI on 22 prominent open-source and proprietary LLMs of varying sizes and capabilities, discovering significant variations in LLMs' gender inclusivity. Our study highlights the importance of improving LLMs' inclusivity, providing a critical benchmark for future advancements in gender fairness in generative models.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
arXiv 2024
-
[2]
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu,...
arXiv 2025
-
[3]
Anthropic. 2024. https://www.anthropic.com/news/claude-3-haiku Claude 3 haiku: our fastest model yet . Available at: https://www.anthropic.com/news/claude-3-haiku
work page 2024
-
[4]
Anthropic. 2025. https://www.anthropic.com/news/claude-4 Introducing claude 4 . https://www.anthropic.com/news/claude-4. Accessed: 2025-05-22
work page 2025
-
[5]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of `` bias '' in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454--5476, Online. Association for Computational Linguistics
-
[6]
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS'16, page 4356–4364, Red Hook, NY, USA. Curran Associates Inc
work page 2016
-
[7]
Charles E. Brown. 1998. https://doi.org/10.1007/978-3-642-80328-4_13 Coefficient of Variation , pages 155--157. Springer Berlin Heidelberg, Berlin, Heidelberg
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
Show all 59 references
-
[9]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[10]
Jose Camacho-collados, Kiamehr Rezaee, Talayeh Riahi, Asahi Ushio, Daniel Loureiro, Dimosthenis Antypas, Joanne Boisson, Luis Espinosa Anke, Fangyu Liu, and Eugenio Mart \' nez C \'a mara. 2022. https://doi.org/10.18653/v1/2022.emnlp-demos.5 T weet NLP : Cutting-edge natural l...
2022 doi
-
[11]
Yuen Chen, Vethavikashini Chithrra Raghuram, Justus Mattern, Mrinmaya Sachan, Rada Mihalcea, Bernhard Scholkopf, and Zhijing Jin. 2022. https://api.semanticscholar.org/CorpusID:254926728 Testing occupational gender bias in language models: Towards robust measurement and zero-s...
2022
-
[12]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[13]
Google DeepMind. 2024. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/ Gemini: Our largest and most capable ai models . https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/
2024
-
[14]
Google DeepMind. 2025. https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#gemini-2-5-thinking Gemini 2.5: More capable, better at thinking, and available in more products
2025
-
[15]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...
2025 arXiv
-
[16]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2025 arXiv
-
[17]
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.150 Harms of gender exclusivity and challenges in non-binary representation in language technologies . In Proceedings of the ...
2021 doi
-
[18]
o rklund, and Henrik Bj\
Hannah Devinney, Jenny Bj\" o rklund, and Henrik Bj\" o rklund. 2022. https://doi.org/10.1145/3531146.3534627 Theories of “gender” in nlp bias research . In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT '22, page 2083–2102, New Yor...
2022
-
[19]
Yu, and James Caverlee
Xiangjue Dong, Yibo Wang, Philip S. Yu, and James Caverlee. 2024. https://arxiv.org/abs/2402.11190 Disclosure and mitigation of gender bias in llms . Preprint, arXiv:2402.11190
2024 arXiv
-
[20]
Abhimanyu Dubey and et al. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783
2024 arXiv
-
[21]
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023. https://doi.org/10.18653/v1/2023.acl-long.507 W ino Q ueer: A community-in-the-loop benchmark for anti- LGBTQ + bias in large language models . In Proceedings of the 61st Annual Meeting of the Associ...
2023 doi
-
[22]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.301 R eal T oxicity P rompts: Evaluating neural toxic degeneration in language models . In Findings of the Association for Computational Linguist...
2020 doi
-
[23]
Gemma Team
et al. Gemma Team. 2024. https://api.semanticscholar.org/CorpusID:270843326 Gemma 2: Improving open language models at a practical size . ArXiv, abs/2408.00118
2024 arXiv
-
[24]
Sourojit Ghosh and Aylin Caliskan. 2023. Chatgpt perpetuates gender bias in machine translation and ignores non-gendered pronouns: Findings across bengali and five other low-resource languages. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 901--912
2023
-
[25]
Google Jigsaw . 2017. Perspective api. https://www.perspectiveapi.com/. Accessed: 2021-02-02
2017
-
[26]
Yue Guo, Yi Yang, and Ahmed Abbasi. 2022. https://doi.org/10.18653/v1/2022.acl-long.72 Auto-debias: Debiasing masked language models with automated biased prompts . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...
2022 doi
-
[27]
Tamanna Hossain, Sunipa Dev, and Sameer Singh. 2023. https://doi.org/10.18653/v1/2023.acl-long.293 MISGENDERED : Limits of large language models in understanding pronouns . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...
2023 doi
-
[28]
Hutto and Eric Gilbert
C. Hutto and Eric Gilbert. 2014. https://doi.org/10.1609/icwsm.v8i1.14550 Vader: A parsimonious rule-based model for sentiment analysis of social media text . Proceedings of the International AAAI Conference on Web and Social Media, 8(1):216--225
2014 doi
-
[29]
Sophie Jentzsch and Cigdem Turan. 2022. https://doi.org/10.18653/v1/2022.gebnlp-1.20 Gender bias in BERT - measuring and analysing biases through sentiment rating in a realistic downstream classification task . In Proceedings of the 4th Workshop on Gender Bias in Natural Langu...
2022 doi
-
[30]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[31]
Anne Lauscher, Archie Crowley, and Dirk Hovy. 2022. https://aclanthology.org/2022.coling-1.105/ Welcome to the modern world of pronouns: Identity-inclusive natural language processing beyond gender . In Proceedings of the 29th International Conference on Computational Linguist...
2022
-
[32]
Tianlin Li, Qing Guo, Aishan Liu, Mengnan Du, Zhiming Li, and Yang Liu. 2023. Fairer: fairness as decision rationale alignment. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org
2023
-
[33]
Justus Mattern, Zhijing Jin, Mrinmaya Sachan, Rada Mihalcea, and Bernhard Schölkopf. 2022. https://arxiv.org/abs/2212.10678 Understanding stereotypes in language models: Towards robust measurement and zero-shot debiasing . Preprint, arXiv:2212.10678
2022 arXiv
-
[34]
Meta AI . 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ LLaMA 4: Advancing Open Multimodal Intelligence . https://ai.meta.com/blog/llama-4-multimodal-intelligence/. Accessed: 2025-04-05
2025
-
[35]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...
2021 doi
-
[36]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[37]
OpenAI. 2023. https://platform.openai.com/docs/models/gpt-3-5-turbo Openai gpt-3.5 api . Available at: https://platform.openai.com/docs/models/gpt-3-5-turbo
2023
-
[38]
OpenAI. 2024 a . https://platform.openai.com/docs/models/gpt-4o Openai gpt-4o . Available at: https://platform.openai.com/docs/models/gpt-4o
2024
-
[39]
OpenAI. 2024 b . https://platform.openai.com/docs/models/gpt-4o-mini Openai gpt-4o-mini . Available at: https://platform.openai.com/docs/models/gpt-4o-mini
2024
-
[40]
OpenAI and et al. 2024. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774
2024 arXiv
-
[41]
i’m fully who i am
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. https://doi.org/10.1145/3593013.3594078 “i’m fully who i am”: Towards centering transgender and non-binary voices to measure biases in open languag...
2023
-
[42]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf Language models are unsupervised multitask learners
2018
-
[43]
Patrick Schramowski, Cigdem Turan, Nils Andersen, Frauke Herbert, Mostafa Shaikh, Franziska Brill, and Kristian Kersting. 2022. https://doi.org/10.1038/s42256-022-00458-8 Large pre-trained language models contain human-like biases of what is right and wrong to do . Nature Mach...
2022 doi
-
[44]
Smith, and Luke Zettlemoyer
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019. https://doi.org/10.18653/v1/P19-1164 Evaluating gender bias in machine translation . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1679--1684, Florence, Italy. Ass...
2019 doi
-
[45]
Karolina Stańczak and Isabelle Augenstein. 2021. https://api.semanticscholar.org/CorpusID:245537923 A survey on gender bias in natural language processing . ArXiv, abs/2112.14168
2021 arXiv
-
[46]
Kunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu, Gelei Deng, Shuai Li, Peigui Qi, Weiming Zhang, Tianwei Zhang, and NengHai Yu. 2024. https://doi.org/10.1145/3658644.3670284 Gendercare: A comprehensive framework for assessing and reducing gender bias in large language models ...
2024
-
[47]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, A...
2024 arXiv
-
[48]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, ...
2025 arXiv
-
[49]
McKee, Jackie Kay, and Shakir Mohamed
Nenad Tomasev, Kevin R. McKee, Jackie Kay, and Shakir Mohamed. 2021. https://doi.org/10.1145/3461702.3462540 Fairness for unobserved characteristics: Insights from technological impacts on queer communities . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and So...
2021
-
[50]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[51]
Lewis Tunstall, Edward Emanuel Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro Von Werra, Cl \'e mentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M Rush, and Thomas Wolf. 2024. https://openreview.net/for...
2024
-
[52]
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.243 kelly is a warm person, joseph is a role model : Gender biases in LLM -generated reference letters . In Findings of the Association for C...
2023 doi
-
[53]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[54]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...
2025 arXiv
-
[55]
Zhiwen You, HaeJin Lee, Shubhanshu Mishra, Sullam Jeoung, Apratim Mishra, Jinseok Kim, and Jana Diesner. 2024. https://doi.org/10.18653/v1/2024.gebnlp-1.16 Beyond binary gender labels: Revealing gender bias in LLM s through gender-neutral name predictions . In Proceedings of t...
2024 doi
-
[56]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. https://doi.org/10.18653/v1/N18-2003 Gender bias in coreference resolution: Evaluation and debiasing methods . In Proceedings of the 2018 Conference of the North A merican Chapter of the Associati...
2018 doi
-
[57]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th Internat...
2024
-
[58]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[59]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.