REVIEW 2 major objections 4 minor 59 references
This paper argues that LLMs measurably change their gender-bias behavior when a prompt signals an evaluation setup — favoring 'they' over 'he' — making many gender-bias benchmarks brittle measures of the task, not the model.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 10:12 UTC pith:SKDPG5DC
load-bearing objection Solid empirical demonstration of prompt sensitivity in gender bias measurement, but the 'testing mode' interpretation overreaches because the manipulation changes the requested output, not just the framing. the 2 major comments →
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: an LLM's measured gender inference is not stable — it shifts when the prompt carries recognizable features of a bias test. Crossing gender salience with instruction presence in a 2×2 design across six models and four tasks, the paper finds mean sensitivity near 0.5 on a zero-to-one scale, with structured shifts in the order they > she > he: evaluation-framed prompts raise 'they' and lower 'he'. Discrete-choice metrics amplify shifts relative to token probabilities. Replicating two published benchmarks, minimal prompt edits changed and sometimes reversed reported bias direction; the paper reads this as a behavioral 'testing mode' — predictable and consistent across models, of u
What carries the argument
The driving apparatus is a 2×2 factorial prompt design: gender salience (explicitly mentioning gender or not) crossed with instruction presence (demanding a response or not), yielding four condition variants per prompt. The target quantity is the pronoun distribution over 'he', 'she', 'they' (aggregated over declined forms), and the sensitivity measure is the Absolute Proportion Difference (APD), an L1 distance between two distributions; APDs are averaged into Gender Salience Effect and Instruction Presence Effect scores, whose mean defines the 'Probability of Pronoun Shift'. This lets the paper attribute shifts to each framing dimension and test whether shifts are structured by pronoun rath
Load-bearing premise
The gender-salient conditions change what the prompt literally asks the model to produce ('the gendered pronoun that comes to mind' versus 'the word'), so the observed shifts may come from the altered request itself rather than from the model recognizing an evaluation setup; if the same task change without evaluation cues produces the same shift, the 'testing mode' interpretation is unsupported.
What would settle it
Hold the requested output constant across conditions — always ask for 'the next word' — while varying only peripheral evaluation signals such as answer options, exam-like formatting, or filler items. If the pronoun distribution does not shift with those peripheral cues, the 'testing mode' reading is refuted and the shifts trace to changed request wording, not evaluation framing. A complementary check: remove all instruction-like formatting from gender-salient prompts; if the neutral-pronoun lift persists with no instruction present, gender salience alone — not evaluation framing — is the drive
If this is right
- Bias scores from existing benchmarks are prompt-conditional: minimal rewordings shift or reverse the reported direction, so single-number bias claims without sensitivity ranges cannot be trusted.
- An apparent trend toward neutrality in LLMs may be compliance with an evaluation frame, not a change in underlying associations — the 'debiasing' the models display under test framing is framing-dependent.
- Discrete-choice metrics exaggerate prompt-induced bias shifts relative to token-probability measures, making the choice of metric a substantive decision about which behavior you mean to report.
- Strategically framed prompts can shift outputs toward fairness goals in practice, but the paper cautions that this is prompt compliance, not genuine debiasing.
- Future benchmark design should avoid heavy instruction and explicit gender references unless the goal is measuring behavior under just such framing; less evaluation-like formats, filler items, and sensitivity ranges are proposed remedies.
Where Pith is reading between the lines
- By the paper's own distributional logic, the same brittleness should appear in other socially charged evaluations (race, age, disability): models have been trained on fairness-test corpora and will reproduce their patterns under test-like prompts. Testing that is a direct extension of the argument.
- The prompt-as-intervention result implies a lightweight deployment-time mitigation for bias, but one that works only when the user happens to phrase prompts in the 'right' way; its value in the wild depends on whether natural-language variation preserves the framing effect.
- If the 'testing mode' reading is right, iterating on static gender-bias benchmarks rewards test-framing compliance — the Goodhart dynamic the paper cites — so keeping bias scores informative over time would require dynamic or hidden evaluation items.
- The they > she > he ordering could reflect 'safe answering' rather than genuine neutrality; checking whether the ordering survives incentives to commit to a gender (few-shot examples, forced informativeness, high temperature) would separate the two.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how prompt framing affects measured gender bias in LLMs. It varies two dimensions—instruction presence and gender salience—across four task formats (completion, association, multiple choice, sentence completion), using six instruction-tuned open-source models. Bias is quantified with token-probability and discrete-choice metrics, summarized through an Absolute Proportion Difference (APD) and associated GenEffect/InstrEffect scores. The authors report that gender salience and, to a lesser extent, instruction presence shift pronoun distributions, with a cross-model tendency toward increased 'they' and decreased 'he.' They also replicate two existing bias benchmarks, apply minimal prompt modifications, and observe shifts or reversals in measured bias, interpreting these as evidence that evaluation-like framing triggers a distinct 'testing mode' behavior. Code and data are released.
Significance. If the central interpretation were established, the paper would be a valuable contribution to the growing literature on the brittleness of LLM bias evaluation and the ecological validity of benchmarks. The study has concrete strengths: it uses multiple models and tasks, reports confidence intervals, releases code and data, and attempts to replicate prior benchmarks. However, the headline claim that evaluation framing per se triggers a distinct behavioral mode is not separable from a direct task-content confound in the current design. The paper can be salvaged either by adding control conditions that hold the requested output type constant while varying framing, or by substantially reframing the conclusions to the narrower, still useful finding that explicit gender-inference instructions shift measured pronoun distributions. The contribution is therefore conditional on revision.
major comments (2)
- [§3, Table 1; §4.3] The Gender Salience manipulation is confounded with the requested output format. In the Association task, Gender+ replaces 'word' with 'gendered pronoun' (and 'The word is' with 'The pronoun is'); in the Completion task, Gender+ adds 'while inferring the described person's gender.' Because token-probability metrics record the first predicted token after the prompt, these edits directly alter the grammatical category and semantic target of the response. The large GenEffect scores in Figures 1–2 could therefore reflect instruction-following/compliance with a changed task rather than sensitivity to an evaluation context. The 'testing mode' interpretation (Abstract; §6.1) and the claim that benchmarks are 'found out' hinge on Gender+ being a pure salience cue. The Limitations admit the manipulation is 'fairly direct' and leave this 'underexplored.' This is load-bearing: the current design ca
- [§6.3, Appendix B] The benchmark modifications in Study One also change the task: appending 'while inferring the described person's gender' to 'Complete the following description:' adds an explicit inference objective, contrary to the B.2 claim that the modification is made 'without altering task semantics.' The P-AT replication is further limited to Flan-T5 variants (B.3), so the statement that the results hold 'for both benchmarks and across tested models' is narrower than the framing suggests. At minimum, the intervention claims should be scaled to the actual manipulations and model coverage.
minor comments (4)
- [Figure 2 caption] Typo: 'instrction' should be 'instruction.'
- [Figure 3] The signed 'sensitivity' plotted here is not the same as the unsigned APD defined in §4.3. Please define explicitly (it appears to be a mean relative probability change) so readers do not conflate the two measures.
- [§5.4, Figure 4] Figure 4 excludes the Association task after filtering, whereas earlier aggregate results (Figure 1) are based on token probabilities. Clarify which conditions and metrics enter each aggregate to avoid apparent inconsistency.
- [B.4, Tables 4–5] The interpretation of AS as 'nearly 60%' is not directly justified; AS is a normalized L1 distance, not a percentage reduction in the original bias score. Please rephrase.
Circularity Check
No significant circularity: the observed prompt sensitivity is a direct experimental contrast and not a quantity derived from its own inputs.
full rationale
The paper's central measurements are direct contrasts between prompt conditions. The Gender Salience Effect Score and Instruction Presence Effect Score are defined as mean L1 (APD) distances between matched prompt variants (Section 4.3). These are operationalizations of the experimental manipulation, not quantities derived from the outcome they purport to explain. No parameter is fitted to one subset of data and then renamed a prediction: the only statistical model is a mixed-effects regression summarizing the pronoun-level trend (Section 5.3), which is a post-hoc summary, not a derivation. There are no self-citations by the authors in the reference list, no imported uniqueness theorem, and no ansatz adopted via citation. The 'testing mode' label is introduced as an interpretation of observed behavioral shifts, and the authors explicitly disclaim an internal-state claim (Section 6.2). The Limitations acknowledge that the prompt manipulation is 'fairly direct' in explicitly mentioning gender; this is a construct-validity confound (the gender-salient conditions also change the requested output, e.g., 'word' vs 'gendered pronoun'), not a circularity, because the APD metric does not assume the shift it measures. The benchmark replications in Section 6.3 apply external benchmarks and compare original vs. modified prompts; the observed changes are empirical, not forced by normalization. Under the stated criteria, no circular step can be exhibited with a specific equation-identity or fitted-input-as-prediction reduction.
Axiom & Free-Parameter Ledger
free parameters (1)
- invalid-completion exclusion threshold =
60%
axioms (5)
- domain assumption Physical attributes such as 'strong', 'slim', 'bald', and 'moustache' activate gender-stereotypical associations in LLMs.
- domain assumption Pronoun choice (he/she/they) is a valid and sufficient operationalization of gender bias for these tasks.
- domain assumption Ten generations with up to 50 tokens under default decoding settings give stable pronoun-distribution estimates.
- domain assumption Instruction-tuned checkpoints and their default decoding settings are representative of LLM behavior.
- standard math L1 distance (APD) over pronoun distributions is a meaningful sensitivity measure.
invented entities (1)
-
'testing mode' behavioral construct
no independent evidence
Cite this review
Pith. "Pith review of Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases." pith.science (2026). https://pith.science/paper/SKDPG5DC
@misc{pith2026250904373,
author = {Pith},
title = {Pith review of: Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases},
year = {2026},
howpublished = {\url{https://pith.science/paper/SKDPG5DC}},
note = {Machine review of arXiv:2509.04373}
}
read the original abstract
As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias. These efforts often rely on evaluation tasks that differ from natural language distributions, as they typically involve carefully constructed task prompts that overtly or covertly signal the presence of gender bias-related content. In this paper, we examine how signaling the evaluative purpose of a task impacts measured gender bias in LLMs. Concretely, we test models under prompt conditions that (1) make the testing context salient, and (2) make gender-focused content salient. We then assess prompt sensitivity across four task formats with both token-probability and discrete-choice metrics. We find that prompts that more clearly align with (gender bias) evaluation framing elicit distinct gender output distributions compared to less evaluation-framed prompts. Discrete-choice metrics further tend to amplify bias relative to probabilistic measures. These findings do not only highlight the brittleness of LLM gender bias evaluations but open a new puzzle for the NLP benchmarking and development community: To what extent can well-controlled testing designs trigger LLM "testing mode" performance, and what does this mean for the ecological validity of future benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024. Do large language models discriminate in hiring decisions on the basis of race, ethnicity, and gender? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 386--397
work page 2024
-
[4]
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths. 2025. Explicitly unbiased large language models still form biased associations. Proceedings of the National Academy of Sciences, 122(8):e2416228122
2025
-
[5]
Mahzarin R Banaji and Curtis D Hardin. 1996. Automatic stereotyping. Psychological science, 7(3):136--141
work page 1996
-
[6]
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings . Advances in neural information processing systems, 29
work page 2016
-
[7]
Samuel Bowman and George Dahl. 2021. What Will it Take to Fix Benchmarking in Natural Language Understanding? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4843--4855
work page 2021
-
[8]
Aylin Caliskan, Pimparkar Parth Ajay, Tessa Charlesworth, Robert Wolfe, and Mahzarin R Banaji. 2022. Gender bias in word embeddings: a comprehensive analysis of frequency, syntax, and semantics. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 156--170
work page 2022
-
[9]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183--186
2017
-
[10]
Anwoy Chatterjee, HSVNS Kowndinya Renduchintala, Sumit Bhatia, and Tanmoy Chakraborty. 2024. Posix: A prompt sensitivity index for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 14550--14565
work page 2024
-
[11]
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. Marked personas: Using natural language prompts to measure stereotypes in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1504--1532
work page 2023
-
[12]
Elliot Creager, J \"o rn-Henrik Jacobsen, and Richard Zemel. 2021. Environment inference for invariant learning. In International Conference on Machine Learning, pages 2189--2200. PMLR
work page 2021
-
[13]
Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, and Tanu Mitra. 2024. “They are uncultured”: Unveiling Covert Harms and Social Threats in LLM Generated Conversations . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20339--20369
work page 2024
-
[14]
Yuhao Dan, Zhikai Lei, Yiyang Gu, Yong Li, Jianghao Yin, Jiaju Lin, Linhao Ye, Zhiyan Tie, Yougen Zhou, Yilei Wang, et al. 2024. Educhat: A large language model-based conversational agent for intelligent education. In China Conference on Knowledge Graph and Semantic Computing, pages 297--308. Springer
work page 2024
-
[15]
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: a case study of semantic representation bias in a high-stakes setting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 120--128
work page 2019
-
[16]
Yitian Ding, Jinman Zhao, Chen Jia, Yining Wang, Zifan Qian, Weizhe Chen, and Xingyu Yue. 2025. Gender bias in large language models across multiple languages: a case study of chatgpt. In Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), pages 552--579
work page 2025
-
[17]
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and mitigating unintended bias in text classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pages 67--73
work page 2018
-
[18]
Probing Explicit and Implicit Gender Bias through LLM Conditional Text Generation
Xiangjue Dong, Yibo Wang, Philip S. Yu, and James Caverlee. 2023. https://arxiv.org/abs/2311.00306 Probing Explicit and Implicit Gender Bias Through LLM Conditional Text Generation . arXiv preprint arXiv:2311.00306
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[19]
Xiangjue Dong, Yibo Wang, Philip S. Yu, and James Caverlee. 2024. https://arxiv.org/abs/2402.11190 Disclosure and mitigation of gender bias in llms . arXiv preprint arXiv:2402.11190
Pith/arXiv arXiv 2024
-
[20]
John C Duchi and Hongseok Namkoong. 2021. Learning models with uniform performance via distributionally robust optimization. The Annals of Statistics, 49(3):1378--1406
work page 2021
-
[21]
Satyam Dwivedi, Sanjukta Ghosh, and Shivam Dwivedi. 2023. https://doi.org/10.21659/rupkatha.v15n4.10 Breaking the bias: Gender fairness in llms using prompt engineering and in-context learning . Rupkatha Journal on Interdisciplinary Studies in Humanities, 15(4)
-
[22]
Chengguang Gan, Qinghao Zhang, and Tatsunori Mori. 2024. Application of llm agents in recruitment: a novel framework for automated resume screening. Journal of Information Processing, 32:881--893
work page 2024
-
[23]
Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Jerry Chun-Wei Lin. 2023. Large language models in education: Vision and opportunities. In 2023 IEEE International Conference on Big Data (BigData), pages 4776--4785. IEEE
work page 2023
-
[24]
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. Journal of Artificial Intelligence Research, 77:103--166
work page 2023
-
[25]
Charles AE Goodhart. 1984. Problems of monetary management: the uk experience. In Monetary theory and practice: The UK experience, pages 91--121. Springer
work page 1984
-
[26]
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. 2024. Alignment faking in large language models. arXiv preprint arXiv:2412.14093
Pith/arXiv arXiv 2024
-
[27]
Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998. Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6):1464
work page 1998
-
[28]
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107--112
2018
-
[29]
Joschka Haltaufderheide and Robert Ranisch. 2024. The Ethics of ChatGPT in Medicine and Healthcare: a Systematic Review on Large Language Models (LLMs) . NPJ digital medicine, 7(1):183
work page 2024
-
[30]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI Generates Covertly Racist Decisions About People Based on Their Dialect . Nature, 633(8028):147--154
work page 2024
-
[31]
Jennifer Hu and Roger Levy. 2023. Prompting is not a substitute for probability measurements in large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5040--5060
2023
-
[32]
Dong Huang, Jie M Zhang, Qingwen Bu, Xiaofei Xie, Junjie Chen, and Heming Cui. 2024. Bias testing and mitigation in llm-based code generation. ACM Transactions on Software Engineering and Methodology
2024
-
[33]
Pavel Izmailov, Polina Kirichenko, Nate Gruver, and Andrew G Wilson. 2022. On feature learning in the presence of spurious correlations. Advances in Neural Information Processing Systems, 35:38516--38532
work page 2022
-
[34]
Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki, and Timothy Baldwin. 2024. Evaluating Gender Bias in Large Language Models via Chain-of-thought Prompting . arXiv preprint arXiv:2401.15585
Pith/arXiv arXiv 2024
-
[35]
Kimmo Karkkainen and Jungseock Joo. 2021. Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1548--1558
work page 2021
-
[36]
Styliani Katsarou, Borja Rodr \' guez-G \'a lvez, and Jesse Shanahan. 2022. Measuring gender bias in contextualized embeddings. In Computer Sciences and Mathematics Forum, page 3. MDPI
work page 2022
-
[37]
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al. 2021. Dynabench: Rethinking benchmarking in nlp. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages...
work page 2021
-
[38]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM Collective Intelligence Conference, pages 12--24
work page 2023
-
[39]
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, J \'e r \'e my Scheurer, Marius Hobbhahn, Alexander Meinke, and Owain Evans. 2024. Me, myself, and ai: The situational awareness dataset (sad) for llms . Advances in Neural Information Processing Systems, 37:64010--64118
work page 2024
-
[40]
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, pages 552--561
work page 2012
-
[41]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pages 3730--3738
work page 2015
-
[42]
Alexander Meinke, Bronson Schoen, J \'e r \'e my Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. 2024. Frontier models are capable of in-context scheming. arXiv preprint arXiv:2412.04984
Pith/arXiv arXiv 2024
-
[43]
Kirsten Morehouse, Siddharth Swaroop, and Weiwei Pan. 2025. Position: Rethinking LLM Bias Probing Using Lessons from the Social Sciences . In Forty-second International Conference on Machine Learning Position Paper Track
work page 2025
-
[44]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring Stereotypical Bias in Pretrained Language Models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356--5371
work page 2021
-
[45]
Joe Needham, Giles Edkins, Govind Pimpale, Henning Bartsch, and Marius Hobbhahn. 2025. Large language models often know when they are being evaluated. arXiv preprint arXiv:2505.23836
Pith/arXiv arXiv 2025
-
[46]
Debora Nozza, Federico Bianchi, Dirk Hovy, et al. 2021. HONEST: Measuring hurtful sentence completion in language models . In Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: Human language technologies. Association for Computational Linguistics
work page 2021
-
[47]
Jane Oakhill, Alan Garnham, and David Reynolds. 2005. Immediate activation of stereotypical gender information. Memory & cognition, 33(6):972--983
work page 2005
-
[48]
Dario Onorati, Elena Sofia Ruzzetti, Davide Venditti, Leonardo Ranaldi, and Fabio Massimo Zanzotto. 2023. Measuring bias in Instruction-Following models with P-AT . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8006--8034
work page 2023
-
[49]
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2019. Distributionally Robust Neural Networks . In International Conference on Learning Representations
work page 2019
-
[50]
Abel Salinas, Parth Shah, Yuzhong Huang, Robert McCormack, and Fred Morstatter. 2023. The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1--15
work page 2023
-
[51]
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting . In International Conference on Learning Representations
work page 2024
-
[52]
Marilyn Strathern. 1997. ‘Improving ratings’: audit in the British University system . European review, 5(3):305--321
work page 1997
-
[53]
Kunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu, Gelei Deng, Shuai Li, Peigui Qi, Weiming Zhang, Tianwei Zhang, and Nenghai Yu. 2024. Gendercare: a comprehensive framework for assessing and reducing gender bias in large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 1196--1210
work page 2024
-
[54]
Kremena Valkanova and Pencho Yordanov. 2024. Irrelevant alternatives bias large language model hiring decisions. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 6899--6912
work page 2024
-
[55]
Kelly is a Warm Person, Joseph is a Role Model
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.243 "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3730--3748. Association for Computat...
-
[56]
Guangyu Wang, Guoxing Yang, Zongxin Du, Longjun Fan, and Xiaohu Li. 2023. Clinicalgpt: Large language models finetuned with diverse medical data and comprehensive evaluation. arXiv preprint arXiv:2306.09968
Pith/arXiv arXiv 2023
-
[57]
Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac. 2024. Implicit bias in large language models: Experimental proof and implications for education. Journal of Research on Technology in Education, pages 1--24
work page 2024
-
[58]
Wikipedia contributors . 2025. https://en.wikipedia.org/wiki/GPTeens Gpteens . Accessed: 2025-05-16
work page 2025
-
[59]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15--20
work page 2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.