REVIEW 3 major objections 1 cited by
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
T0 review · 3 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Five LLMs show confidence gaps of up to 40 percentage points on coreference resolution for intersectional identities, with the greatest uncertainty for doubly disadvantaged groups in anti-stereotypical settings.
desk verdict The abstract promises an intersectional bias benchmark, but the full text is an unrelated robotics paper—there is no actual paper here to review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two components carry the argument. First, WinoIdentity, a benchmark built by augmenting WinoBias with 25 demographic markers across 10 attributes (including age, nationality, race, body type, sexual orientation, and socio-economic status) intersected with binary gender, yielding 245,700 prompts and 50 distinct bias patterns. Second, Coreference Confidence Disparity, a group (un)fairness metric that measures whether models are more or less confident for some intersectional identities than others. The paper's investigative lens is uncertainty: rather than looking only at task accuracy, it treats a model's output confidence as a signal of harm from omission and underrepresentation.
What would settle it
If the same confidence disparities appear when demographic markers are replaced by matched neutral tokens, or if the models are shown to be well calibrated within each identity group, the claim that confidence gaps reflect social bias would collapse; the memorization claim would fail if confidence remains high on novel identity-name combinations.
Extended reading notes
Core claim
The paper's central claim is that coreference resolution systems exhibit measurable, group-dependent confidence differences that signal social bias, and that these differences can be detected even when accuracy looks high. Evidence comes from a new benchmark, WinoIdentity, which augments WinoBias with 25 demographic markers across 10 attributes intersected with binary gender, producing 245,700 prompts covering 50 bias patterns. The proposed group unfairness metric, Coreference Confidence Disparity, quantifies whether models are more or less confident for some intersectional identities than others. The authors report confidence disparities of up to 40% across demographic attributes, with the
Load-bearing premise
The core claims stand on the premise that a model's reported confidence in coreference resolution is a faithful measure of its true uncertainty about an identity, and that lower confidence on privileged markers indicates memorization rather than reasoning.
Editorial extensions
If this is right
- Fairness audits of LLMs in decision-support roles should measure confidence distributions, not just accuracy, because confidence gaps can persist even when outputs are correct.
- WinoIdentity provides a reusable testbed for intersectional coreference bias across 50 identity combinations, allowing systematic comparisons across models and training regimes.
- If lower confidence on privileged markers indeed reflects memorization rather than reasoning, reported accuracy gains in recent LLMs may overstate their reasoning ability.
- Confidence disparities of up to 40% imply that in contexts like hiring or admissions, models could be systematically more certain about applicants from some intersectional groups than others, potentially affecting downstream decisions.
- The benchmark's construction makes it possible to separate stereotype-consistent from anti-stereotypical settings, allowing targeted study of when uncertainty spikes.
Reading between the lines
- Editorial: Confidence disparities may partly reflect calibration differences or task difficulty rather than social bias; a comparison that matches prompts for lexical frequency and syntactic complexity would help separate these explanations.
- Editorial: The memorization inference is indirect; it would be strengthened by testing on identity-name combinations unseen in training or by probing whether confidence drops track training-data frequency.
- Editorial: The intersectional extension could transfer to other NLP tasks with confidence signals, such as question answering or summarization, where omission and underrepresentation harms also matter.
- Editorial: The benchmark's binary-gender intersection excludes non-binary and genderqueer identities; extending the attribute grid would likely reveal additional patterns of uncertainty.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, as identified by its abstract and arXiv metadata, claims a study of intersectional bias in large language models through confidence disparities in coreference resolution. The abstract states that the authors constructed a WinoIdentity benchmark with 25 demographic markers and 10 attributes, propose a Coreference Confidence Disparity metric, evaluate five LLMs, observe confidence disparities as high as 40%, and conclude that LLM performance may reflect memorization rather than logical reasoning. However, the full text provided is an entirely different paper: 'DexFruit: Dexterous Manipulation and Gaussian Splatting Inspection of Fruit,' a robotics manuscript by different authors concerning tactile sensing, grasping, and 3D reconstruction of fruit. The DexFruit text contains none of the components advertised in the abstract: no WinoIdentity benchmark, no coreference-confidence metric, no list of evaluated LLMs, no bias results, and no related analysis. The arXiv identifier in the file header, 2508.07118v3 [cs.RO], also differs from the claimed identifier 2508.07111 [cs.CL]. Because the manuscript body is unrelated to the central claim, no scientific assessment of the bias study is possible.
Significance. If the abstract's claims were supported, the work would address an important gap by extending single-axis bias evaluations to intersectional identity combinations and by proposing confidence disparity as a measurable bias signal. The reported 40% disparity and the distinction between value-alignment and validity failures could be significant contributions to the LLM-bias literature. However, as submitted, the manuscript contains no methods, experiments, benchmark description, metric definition, or results that substantiate these claims. The full text is a robotics paper with a different scope and authorship. There are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions to credit. The submission therefore offers no validatable scientific content for the claimed study, and any potential significance is entirely contingent on external evidence not present in the manuscript.
major comments (3)
- [Full Text (entire document vs. Abstract)] The central claim of the abstract—an empirical study of intersectional confidence disparities in five LLMs using the WinoIdentity benchmark—has no supporting methodology or results in the submitted full text. The full text is a robotics paper titled 'DexFruit: Dexterous Manipulation and Gaussian Splatting Inspection of Fruit' by a different author list. It describes grasping policies, tactile sensing, and 3D reconstruction, with no mention of coreference resolution, demographic attributes, confidence disparities, or bias. This is not a local presentation issue; it makes the central claim completely unverifiable from the manuscript.
- [Header/arXiv metadata] The manuscript header identifies the paper as arXiv:2508.07118v3 [cs.RO], while the submission is claimed to be arXiv:2508.07111 [cs.CL]. The abstract's empirical claims (245,700 prompts, 50 bias patterns, 40% disparity) cannot be traced to any section, equation, table, or dataset in the body. Since the document's apparent identity and content are inconsistent with the abstract, the submission does not constitute a coherent research paper that can be reviewed.
- [Abstract (unsupported central assertions)] Even granting the abstract's framing, two load-bearing premises are asserted without evidence: (i) that output confidence in coreference resolution is a valid proxy for model uncertainty and social bias, and (ii) that lower confidence for privileged markers indicates memorization rather than logical reasoning. Because the full text is unrelated, no derivation, experiment, or argument addresses these premises. As submitted, the paper's claims are unfalsifiable observations with no methodological grounding.
Circularity Check
No circularity can be established; the submitted full text does not contain the claimed methodology.
full rationale
The abstract describes a benchmark (WinoIdentity), a metric (Coreference Confidence Disparity), and empirical findings, but the full text supplied is an unrelated robotics paper (DexFruit) with no equations, no benchmark construction, no metric definition, and no results relevant to the abstract's claims. Circularity requires exhibiting a specific reduction of a claimed derivation to its own inputs or a self-citation chain; because none of the derivation steps are present, no such reduction can be quoted. The mismatch and resulting lack of support are serious validity/correctness concerns, but they are not evidence of circular reasoning. Accordingly, no circular step is identified and the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Model confidence in coreference resolution is a valid proxy for model uncertainty and for socially meaningful bias.
- domain assumption Template prompts derived from WinoBias, augmented with demographic markers, transfer to real-world hiring and admissions contexts.
- domain assumption Decreased confidence on privileged markers implies memorization rather than logical reasoning.
invented entities (2)
-
WinoIdentity benchmark
-
Coreference Confidence Disparity metric
Cite this review
Pith. "Pith review of Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution." pith.science (2026). https://pith.science/paper/ZBVOPFCS
@misc{pith2026250807111,
author = {Pith},
title = {Pith review of: Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZBVOPFCS}},
note = {Machine review of arXiv:2508.07111}
}
read the original abstract
Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and admissions. There is, however, scientific consensus that AI systems can reflect and exacerbate societal biases, raising concerns about identity-based harm when used in critical social contexts. Prior work has laid a solid foundation for assessing bias in LLMs by evaluating demographic disparities in different language reasoning tasks. In this work, we extend single-axis fairness evaluations to examine intersectional bias, recognizing that when multiple axes of discrimination intersect, they create distinct patterns of disadvantage. We create a new benchmark called WinoIdentity by augmenting the WinoBias dataset with 25 demographic markers across 10 attributes, including age, nationality, and race, intersected with binary gender, yielding 245,700 prompts to evaluate 50 distinct bias patterns. Focusing on harms of omission due to underrepresentation, we investigate bias through the lens of uncertainty and propose a group (un)fairness metric called Coreference Confidence Disparity which measures whether models are more or less confident for some intersectional identities than others. We evaluate five recently published LLMs and find confidence disparities as high as 40% along various demographic attributes including body type, sexual orientation and socio-economic status, with models being most uncertain about doubly-disadvantaged identities in anti-stereotypical settings. Surprisingly, coreference confidence decreases even for hegemonic or privileged markers, indicating that the recent impressive performance of LLMs is more likely due to memorization than logical reasoning. Notably, these are two independent failures in value alignment and validity that can compound to cause social harm.
Forward citations
Cited by 1 Pith paper
-
A Novel Computational Thermodynamics Framework with Intrinsic Chemical Short-Range Order
A hybrid CVM-CALPHAD model using the Fowler-Yang-Li transform models chemical short-range order in multicomponent alloys at CVM accuracy and lower cost, benchmarked on fcc binaries and demonstrated on Cu-Au and Cu-Au-Ag.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Junaid Ali, Preethi Lahoti, and Krishna P. Gummadi. Accounting for model uncertainty in algorithmic discrimination. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES '21, pp.\ 336–345, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450384735. doi:10.1145/3461702.3462630. URL https://doi.org/10.1145/346...
arXiv 2021
-
[3]
Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. Do large language models discriminate in hiring decisions on the basis of race, ethnicity, and gender? arXiv preprint arXiv:2406.10486, 2024
arXiv 2024
-
[4]
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
Lena Armstrong, Abbey Liu, Stephen MacNeil, and Dana \"e Metaxa. The silicone ceiling: Auditing gpt's race and gender biases in hiring. arXiv preprint arXiv:2405.04412, 2024
work page Pith review arXiv 2024
-
[5]
Soumya Barikeri, Anne Lauscher, Ivan Vuli \'c , and Goran Glava s . Redditbias: A real-world resource for bias evaluation and debiasing of conversational language models. arXiv preprint arXiv:2106.03521, 2021
work page Pith review arXiv 2021
-
[6]
Unmasking contextual stereotypes: Measuring and mitigating bert's gender bias
Marion Bartl, Malvina Nissim, and Albert Gatt. Unmasking contextual stereotypes: Measuring and mitigating bert's gender bias. arXiv preprint arXiv:2010.14534, 2020
arXiv 2010
-
[7]
On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.\ 610--623, 2021
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.\ 610--623, 2021
2021
-
[8]
Marianne Bertrand and Sendhil Mullainathan. Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination. American Economic Review, 94 0 (4): 0 991–1013, September 2004. doi:10.1257/0002828042002561. URL https://www.aeaweb.org/articles?id=10.1257/0002828042002561
Show all 72 references
-
[9]
Language (technology) is power: A critical survey of `` bias '' in NLP
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. Language (technology) is power: A critical survey of `` bias '' in NLP . In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), ACL, 2020
2020
-
[10]
Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th In...
2021
-
[11]
Language and identity
Mary Bucholtz and Kira Hall. Language and identity. A companion to linguistic anthropology, 1: 0 369--394, 2004
2004
-
[12]
Toward gender-inclusive coreference resolution: An analysis of gender and bias throughout the machine learning lifecycle
Yang Trista Cao and Hal Daum \'e III. Toward gender-inclusive coreference resolution: An analysis of gender and bias throughout the machine learning lifecycle. Computational Linguistics, 47 0 (3): 0 615--661, 2021
2021
-
[13]
Extracting intersectional stereotypes from embeddings: Developing and validating the flexible intersectional stereotype extraction procedure
Tessa ES Charlesworth, Kshitish Ghate, Aylin Caliskan, and Mahzarin R Banaji. Extracting intersectional stereotypes from embeddings: Developing and validating the flexible intersectional stereotype extraction procedure. PNAS nexus, 3 0 (3): 0 pgae089, 2024
2024
-
[14]
Intersectionality as critical social theory: Intersectionality as critical social theory, patricia hill collins, duke university press, 2019
Patricia Hill Collins, Elaini Cristina Gonzaga da Silva, Emek Ergun, Inger Furseth, Kanisha D Bond, and Jone Mart \' nez-Palacios. Intersectionality as critical social theory: Intersectionality as critical social theory, patricia hill collins, duke university press, 2019. Cont...
2019
-
[15]
A validity perspective on evaluating the justified use of data-driven decision-making algorithms
Amanda Coston, Anna Kawakami, Haiyi Zhu, Ken Holstein, and Hoda Heidari. A validity perspective on evaluating the justified use of data-driven decision-making algorithms. In 2023 IEEE conference on secure and trustworthy machine learning (SaTML), pp.\ 690--704. IEEE, 2023
2023
-
[16]
The algorithmic leviathan: Arbitrariness, fairness, and opportunity in algorithmic decision-making systems
Kathleen Creel and Deborah Hellman. The algorithmic leviathan: Arbitrariness, fairness, and opportunity in algorithmic decision-making systems. Canadian Journal of Philosophy, 52 0 (1): 0 26--43, 2022
2022
-
[17]
Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics
Kimberl\'e Crenshaw. Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. The University of Chicago Legal Forum, 140: 0 139--167, 1989
1989
-
[18]
Are ai systems biased against the poor? a machine learning analysis using word2vec and glove embeddings
Georgina Curto, Mario Fernando Jojoa Acosta, Flavio Comim, and Bego \ n a Garcia-Zapirain. Are ai systems biased against the poor? a machine learning analysis using word2vec and glove embeddings. AI & society, 39 0 (2): 0 617--632, 2024
2024
-
[19]
Second order winobias (sowinobias) test set for latent gender bias detection in coreference resolution
Hillary Dawkins. Second order winobias (sowinobias) test set for latent gender bias detection in coreference resolution. arXiv preprint arXiv:2109.14047, 2021
2021 arXiv
-
[20]
Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning
Stefan Depeweg, Jose-Miguel Hernandez-Lobato, Finale Doshi-Velez, and Steffen Udluft. Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning. In International conference on machine learning, pp.\ 1184--1193. PMLR, 2018
2018
-
[21]
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp.\ 7659--7666, 2020
2020
-
[22]
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. Fairness through awareness. In Proceedings of the 3rd innovations in theoretical computer science conference, pp.\ 214--226, 2012
2012
-
[23]
Winoqueer: A community-in-the-loop benchmark for anti-lgbtq+ bias in large language models
Virginia K Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. Winoqueer: A community-in-the-loop benchmark for anti-lgbtq+ bias in large language models. arXiv preprint arXiv:2306.15087, 2023
2023 arXiv
-
[24]
A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition
Susan T Fiske, Amy JC Cuddy, Peter Glick, and Jun Xu. A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology, 82 0 (6): 0 878--902, 2002
2002
-
[25]
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Maria Florina Balcan and Kilian Q. Weinberger (eds.), Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Mac...
2016
-
[26]
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. Counterfactual fairness in text classification through robustness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pp.\ 219--226, 2019
2019
-
[27]
A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges
Usman Gohar and Lu Cheng. A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges. arXiv preprint arXiv:2305.06969, 2023
2023 arXiv
-
[28]
Algorithmic arbitrariness in content moderation
Juan Felipe Gomez, Caio Machado, Lucas Monteiro Paes, and Flavio Calmon. Algorithmic arbitrariness in content moderation. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 2234--2253, 2024
2024
-
[29]
Akal badi ya bias: An exploratory study of gender bias in hindi language technology
Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, Ujwal Gadiraju, Aditya Vashistha, Vivek Seshadri, et al. Akal badi ya bias: An exploratory study of gender bias in hindi language technology. In The 2024 ACM Conference on ...
2024
-
[30]
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. Equality of opportunity in supervised learning. Advances in neural information processing systems, 29, 2016
2016
-
[31]
Misgendered: Limits of large language models in understanding pronouns
Tamanna Hossain, Sunipa Dev, and Sameer Singh. Misgendered: Limits of large language models in understanding pronouns. arXiv preprint arXiv:2306.03950, 2023
2023 arXiv
-
[32]
Socialcounterfactuals: Probing and mitigating intersectional social biases in vision-language models with counterfactual examples
Phillip Howard, Avinash Madasu, Tiep Le, Gustavo Lujan Moreno, Anahita Bhiwandiwalla, and Vasudev Lal. Socialcounterfactuals: Probing and mitigating intersectional social biases in vision-language models with counterfactual examples. In Proceedings of the IEEE/CVF Conference o...
2024
-
[33]
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
Eyke H \"u llermeier and Willem Waegeman. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine learning, 110 0 (3): 0 457--506, 2021
2021
-
[34]
Are female carpenters like blue bananas? a corpus investigation of occupation gender typicality
Da Ju, Karen Ullrich, and Adina Williams. Are female carpenters like blue bananas? a corpus investigation of occupation gender typicality. In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 4254--4274, 2024
2024
-
[35]
Taxonomizing and measuring representational harms: A look at image tagging
Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. Taxonomizing and measuring representational harms: A look at image tagging. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp.\ 1427...
2023
-
[36]
What uncertainties do we need in bayesian deep learning for computer vision? In I
Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 30. Curran ...
2017
-
[37]
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif M Mohammad. Examining gender and race bias in two hundred sentiment analysis systems. arXiv preprint arXiv:1805.04508, 2018
2018 arXiv
-
[38]
Dreyer, Aleksandar Shtedritski, and Yuki M
Hannah Rose Kirk, Yennie Jun, Haider Iqbal, Elias Benussi, Filippo Volpin, Frederic A. Dreyer, Aleksandar Shtedritski, and Yuki M. Asano. Bias out-of-the-box: an empirical analysis of intersectional occupational biases in popular generative language models. In Proceedings of t...
2024
-
[39]
Stereotype content at the intersection of gender and sexual orientation
Amanda Klysing, Anna Lindqvist, and Fredrik Bj \"o rklund. Stereotype content at the intersection of gender and sexual orientation. Frontiers in Psychology, 12: 0 713839, 2021
2021
-
[40]
intersectionally fair
Youjin Kong. Are “intersectionally fair” ai algorithms really fair to women of color? a philosophical analysis. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 485--494, 2022
2022
-
[41]
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, pp.\ 12--24, 2023
2023
-
[42]
Uncertainty as a fairness measure
Selim Kuzucu, Jiaee Cheong, Hatice Gunes, and Sinan Kalkan. Uncertainty as a fairness measure. Journal of Artificial Intelligence Research, 81: 0 307--335, 2024
2024
-
[43]
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30, 2017
2017
-
[44]
Collecting a large-scale gender bias dataset for coreference resolution and machine translation
Shahar Levy, Koren Lazar, and Gabriel Stanovsky. Collecting a large-scale gender bias dataset for coreference resolution and machine translation. arXiv preprint arXiv:2109.03858, 2021
2021 arXiv
-
[45]
A survey on fairness in large language models
Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. A survey on fairness in large language models. arXiv preprint arXiv:2308.10149, 2023
2023 arXiv
-
[46]
Comparing diversity, negativity, and stereotypes in chinese-language ai technologies: a case study on baidu, ernie and qwen
Geng Liu, Carlo Alberto Bono, and Francesco Pierri. Comparing diversity, negativity, and stereotypes in chinese-language ai technologies: a case study on baidu, ernie and qwen. arXiv preprint arXiv:2408.15696, 2024
2024 arXiv
-
[47]
Intersectional stereotypes in large language models: Dataset and analysis
Weicheng Ma, Brian Chiang, Tong Wu, Lili Wang, and Soroush Vosoughi. Intersectional stereotypes in large language models: Dataset and analysis. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp.\ 8589--8597, 2023
2023
-
[48]
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229, 2024
2024 arXiv
-
[49]
Torr, and Yarin Gal
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip H.S. Torr, and Yarin Gal. Deep deterministic uncertainty: A new simple baseline. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.\ 24384--24394, June 2023
2023
-
[50]
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456, 2020
2004 arXiv
-
[51]
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. Crows-pairs: A challenge dataset for measuring social biases in masked language models. arXiv preprint arXiv:2010.00133, 2020
2010 arXiv
-
[52]
Factoring the matrix of domination: A critical review and reimagination of intersectionality in ai fairness
Anaelia Ovalle, Arjun Subramonian, Vagrant Gautam, Gilbert Gee, and Kai-Wei Chang. Factoring the matrix of domination: A critical review and reimagination of intersectionality in ai fairness. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pp.\ 496--...
2023
-
[53]
BBQ : A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. BBQ : A hand-built bias benchmark for question answering. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Findings of the Associat...
2022 doi
-
[54]
Perturbation augmentation for fairer nlp
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. Perturbation augmentation for fairer nlp. arXiv preprint arXiv:2205.12586, 2022
2022 arXiv
-
[55]
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. Gender bias in coreference resolution. arXiv preprint arXiv:1804.09301, 2018
2018 arXiv
-
[56]
The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama
Abel Salinas, Parth Shah, Yuzhong Huang, Robert McCormack, and Fred Morstatter. The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama. In Proceedings of the 3rd ACM Conference on Equity and Access in Algori...
2023
-
[57]
Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36: 0 55565--55581, 2023
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36: 0 55565--55581, 2023
2023
-
[58]
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326, 2019
1909 arXiv
-
[59]
A framework for understanding sources of harm throughout the machine learning life cycle
Harini Suresh and John Guttag. A framework for understanding sources of harm throughout the machine learning life cycle. In Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pp.\ 1--9, 2021
2021
-
[60]
Fairness through aleatoric uncertainty
Anique Tahir, Lu Cheng, and Huan Liu. Fairness through aleatoric uncertainty. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, CIKM '23, pp.\ 2372–2381, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 979840070...
2023
-
[61]
Neutral rewriter: A rule-based and neural approach to automatic rewriting into gender-neutral alternatives
Eva Vanmassenhove, Chris Emmery, and Dimitar Shterionov. Neutral rewriter: A rule-based and neural approach to automatic rewriting into gender-neutral alternatives. arXiv preprint arXiv:2109.06105, 2021
2021 arXiv
-
[62]
Measuring representational harms in image captioning
Angelina Wang, Solon Barocas, Kristen Laird, and Hanna Wallach. Measuring representational harms in image captioning. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 324--335, 2022
2022
-
[63]
Aleatoric and epistemic discrimination: Fundamental limits of fairness interventions
Hao Wang, Luxi He, Rui Gao, and Flavio Calmon. Aleatoric and epistemic discrimination: Fundamental limits of fairness interventions. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[64]
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. Mind the gap: A balanced corpus of gendered ambiguous pronouns. Transactions of the Association for Computational Linguistics, 6: 0 605--617, 2018
2018
-
[65]
Easy problems that llms get wrong
Sean Williams and James Huckle. Easy problems that llms get wrong. arXiv preprint arXiv:2405.19616, 2024
2024 arXiv
-
[66]
Gender, race, and intersectional bias in resume screening via language model retrieval
Kyra Wilson and Aylin Caliskan. Gender, race, and intersectional bias in resume screening via language model retrieval. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp.\ 1578--1590, 2024
2024
-
[67]
What is your favorite gender, mlm? gender bias evaluation in multilingual masked language models
Jeongrok Yu, Seong Ug Kim, Jacob Choi, and Jinho D Choi. What is your favorite gender, mlm? gender bias evaluation in multilingual masked language models. Information, 15 0 (9): 0 549, 2024
2024
-
[68]
Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, and Krishna P Gummadi. Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment. In Proceedings of the 26th international conference on world wide web, pp.\ 1171-...
2017
-
[69]
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. In Marilyn Walker, Heng Ji, and Amanda Stent (eds.), Proceedings of the 2018 Conference of the North A merican Chapter of the Ass...
2018 doi
-
[70]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[71]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[72]
[pronoun]
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
2022
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.