REVIEW 3 major objections 3 minor 35 references
Conflicting LLM bias results may be an artifact of audit design, not of the models.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 16:56 UTC pith:LFZ5E2S3
load-bearing objection A genuinely careful audit-design benchmark; the name-cue caveat is real and disclosed, but the central claims hold up. the 3 major comments →
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that audit format changes both the direction and the magnitude of measured demographic bias in LLM allocation decisions. Using 600 human-authored aid appeals calibrated against a large corpus of real crowdfunding campaigns, the benchmark crosses three tasks (Rate, Rank, Allocate), two comparison contexts (single vs. multi-stimulus), and two presentation modes (transparent minimal pairs vs. disguised bundles that co-vary scenario with demographic cues). The paper reports that models advantage ethnic minority claimants on the single-stimulus Rate task, disadvantage Asian and Hispanic claimants on the Rank task, and shift from near-universal equal splitting in transparent A
What carries the argument
The central object is FairFund-Bench, an audit instrument that systematically varies task format, comparison context, and presentation transparency. Its load-bearing component is the contrast between transparent bundles, which present minimal pairs differing only by name, and disguised bundles, which co-vary names with scenarios so that the demographic contrast is not obvious. Position rotations and Graeco-Latin square designs keep the focal contrast statistically identifiable across the bundle set. The paper also uses a five-level causal framing manipulation derived from the CARIN deservingness heuristic to measure how models respond to the controllability of need, and a name-pair validatio
Load-bearing premise
The results depend on the premise that differences in first and last names alone cause models to perceive the intended racial and gender groups, rather than reacting to other traits that happen to correlate with those names, such as class, age, or region.
What would settle it
Repeat the benchmark with group membership signaled by dialect or by explicit identity statements instead of names. If the direction reversal between the rating and ranking tasks, and the transparent–disguised gap in allocation, disappear or substantially change, then the central claims are artifacts of the name cue rather than properties of model bias.
If this is right
- If the central claim is right, single-format audits cannot support credible overall conclusions about whether a model is biased; the same model can appear positively biased, negatively biased, or unbiased depending on the task and presentation.
- Minimal-pair audits understate the demographic disparities models produce in more deployment-like, disguised settings, so bias estimates from such audits should be treated as lower bounds.
- Causal framing of need, not demographics, is the dominant driver of current model allocations, and this pattern is consistent across tasks and models.
- Managers and regulators considering LLMs for welfare, lending, or hiring decisions should expect allocations to track human deservingness heuristics, which may punish stigmatized causes of need, rather than purely demographic fairness.
- The benchmark's four-pillar framework can be adapted to other allocation domains, providing a way to separate a model's substantive bias from its consistency across audit conditions.
Where Pith is reading between the lines
- One implication the paper leaves implicit: the audit-design sensitivity it documents likely extends to other allocation contexts, so any single-format audit of hiring, lending, or housing decisions should be read as design-dependent rather than as a model-level verdict.
- The near-universal equal splitting in transparent bundles, whether caused by audit awareness or by similarity to training evaluations, suggests that current alignment makes models egalitarian mainly when bias is obvious; disguised, realistic prompts may not inherit that guarantee.
- Because framing effects are stable and large while demographic effects are small and direction-reversing, a testable next step is to examine whether explicit allocation criteria (e.g., 'who would benefit most' vs. 'who is most deserving') shift the framing gradient in predictable ways; the paper's own wording-variant data hint that criterion naming enlarges framing effects.
- The finding that small models show lower deservingness alignment than frontier models suggests that scaling or alignment choices affect how strongly models adopt human-like deservingness judgments, a pattern worth tracking as models evolve.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FairFund-Bench is a benchmark for measuring how audit design choices affect apparent demographic and deservingness biases in LLM resource allocation. It varies three tasks (Rate, Rank, Allocate), single vs. multi-stimulus contexts, and transparent vs. disguised presentation, using 600 human-authored aid requests calibrated to a large GoFundMe corpus and 40 validated name pairs. Across 14 LLMs, the paper reports that the direction of demographic bias reverses by task: minorities are advantaged in single-stimulus ratings but some groups are disadvantaged in ranking. On the Allocate task, transparent demographic bundles elicit near-universal equal splitting, while disguised bundles produce 3–4 times larger dollar gaps. Causal framing of need produces much larger and more consistent effects than demographics, in line with human deservingness heuristics. The benchmark scores models on four pillars: demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency, with code and raw responses released.
Significance. If the results hold, this is a valuable methodological contribution. It offers a concrete explanation for contradictory prior LLM bias audits—design choices can flip the sign and magnitude of measured bias—and provides an extensible benchmark for future work. The paper's strengths are real: a careful factorial design with pre-fixed prompts, balanced rotations, temperature=0, explicit parse-failure handling, bootstrap CIs, and full release of raw responses and scoring code. The claim that framing effects dominate demographic effects is also important, as it redirects attention from identity cues to need narratives. The main risks are interpretational: the name-only demographic signal and the two-part transparent/disguised manipulation both bear directly on the central claims.
major comments (3)
- [§6, Appendix D, §4.2] The demographic-bias interpretation rests on the assumption that the 40 name pairs signal race/gender rather than correlated attributes such as class, age, region, or education. Appendix D matches pairs only on perceived competence and hardworkingness, and the manuscript itself concedes in §6 that name-based estimates may not hold under other cues. Because the headline direction-reversal and the transparent–disguised gap are between-name contrasts, they are vulnerable to models responding to non-racial name connotations differently across tasks. This is load-bearing, not a peripheral caveat. Please add a validation step—for example, re-estimating key contrasts with explicit identity statements on a subset, or with a second name set matched on additional perceived attributes—or systematically reframe the claims as effects of name-based demographic cues rather than demographic bias.
- [§3.3, Table 8, §4.2] The transparent–disguised manipulation is two-dimensional: transparent bundles hold scenario fixed, while disguised bundles co-vary the focal axis with scenario (Table 8). The finding that demographic dollar gaps are 3–4 times larger in disguised bundles is therefore attributable either to reduced transparency or to increased stimulus diversity, or both. The equal-splitting discussion in §6 addresses one alternative mechanism but not this design confound. Please add a design cell that orthogonally varies stimulus diversity (e.g., disguised bundles with non-focal scenario variation but overt demographic labels), or temper the causal claim that minimal-pair audits 'understate' disparities to acknowledge that the disguised condition also changes the informational content of the prompt.
- [§4.3, Figure 6] The headline that framing effects exceed demographic disparities by 'roughly an order of magnitude' compares signed framing contrasts (e.g., +$469 structural vs. self-cause) with mean absolute White–non-White gaps from a different bundle type. These metrics are not directly comparable as reported, and the phrase 'the smallest' is ambiguous. Please report a common dispersion measure for both factors—such as mean absolute deviation from the grand mean or variance-explained-style statistics—or explicitly state the comparison base (e.g., transparent versus disguised demographic gaps).
minor comments (3)
- [§4.2] The Hispanic Rank disadvantage is reported as −0.044 [−0.002, 0.090], whose confidence interval includes zero, yet the text groups Hispanic with Asian as 'some of the same groups.' Please qualify this as not statistically distinguishable from zero.
- [Appendix A, Wording variants] The finding that wording variants substantially change framing-effect magnitudes is reported only in the appendix. Given its relevance to the robustness of the §4.3 magnitude claims, a sentence in the main text or Limitations would be helpful.
- [Figure 5] The figure uses one line per model; with 14 models the legend is hard to read. Consider highlighting the median and the named outliers (Grok 4.20, DeepSeek V3.2) rather than plotting every line.
Circularity Check
FairFund-Bench's headline effects are observed model behaviors, not artifacts of the scoring definitions; no circular derivation chain found.
full rationale
I walked the derivation chain from stimulus construction through pillar scoring to the headline findings. The central results—direction reversal between Rate and Rank, larger disguised than transparent Allocate gaps, and framing effects exceeding demographic effects—are all estimated from raw model responses (ratings, ranks, dollar allocations) and are not forced by the scoring definitions. The P2 'deservingness alignment' pillar is explicitly defined as agreement with the CARIN ordering, so calling a high P2 'alignment' is a measurement, not a derivation; moreover, the framing-dominance claim in §4.3 rests on direct dollar contrasts, independent of P2. The transparent-vs-disguised comparison is also empirical: equal splitting is a model behavior, not an imposed code, and disguised bundles balance scenario against the focal axis via Graeco-Latin squares, so the race/gender contrasts are identified rather than constructed. The paper's self-citations (Lukk et al. 2025; Schneiderhan and Lukk 2023) are contextual and non-load-bearing. The limitations it flags—name-only demographic signaling (§6), author wording effects (§6), and the P1/P4 interpretation caveats (§7, Appendix H)—are construct-validity or measurement-interpretation concerns, not circularity, and the paper states them openly. No fitted parameter is later relabeled as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The benchmark is self-contained in the sense that its conclusions are direct observations of model behavior under a clearly specified instrument.
Axiom & Free-Parameter Ledger
free parameters (2)
- Cohen's d standardization denominator =
median of 14 within-model SDs, per pool x task, fixed across lineup
- P1 transparent-Allocate pool exclusion =
transparent race/gender/intersectional Allocate pools removed from P1
axioms (5)
- domain assumption Names from Elder & Hayes (2023) signal race and gender to LLMs without confounded connotations (competence, class, age, region).
- domain assumption The five causal framing conditions form an ordinal Control/deservingness ladder per CARIN and differ only along that dimension.
- domain assumption The CARIN deservingness ordering (van Oorschot & Roosma 2017) is a valid empirical stand-in for human deservingness judgments on these specific stimuli.
- domain assumption Temperature=0 outputs with reasoning suppressed represent the allocation behavior of the models under assessment.
- domain assumption GoFundMe corpus statistics (word counts, narrative features, topic base rates) validate that the hand-written templates resemble real aid requests.
read the original abstract
Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.
Figures
Reference graph
Works this paper leans on
-
[1]
Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024. https://doi.org/10.18653/v1/2024.acl-short.37 Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race , Ethnicity , and Gender ? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics ( Volume 2: Short Papers ) , pag...
-
[2]
Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. 2025. https://doi.org/10.1093/pnasnexus/pgaf089 Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation . PNAS Nexus, 4(3):pgaf089
-
[3]
Lena Armstrong, Abbey Liu, Stephen MacNeil, and Dana \"e Metaxa. 2024. https://doi.org/10.1145/3689904.3694699 The Silicon Ceiling : Auditing GPT 's Race and Gender Biases in Hiring . In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms , Mechanisms , and Optimization , EAAMO '24, pages 1--18, New York, NY, USA. Association for Comp...
arXiv 2024
-
[4]
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L. Griffiths. 2025. https://doi.org/10.1073/pnas.2416228122 Explicitly unbiased large language models still form biased associations . Proceedings of the National Academy of Sciences, 122(8):e2416228122
-
[5]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk , Nelson Elhage, Zac Hatfield-Dodds , Danny Hernandez, Tristan Hume, and 12 others. 2022. https://doi.org/10.48550/arXiv.2204.05862 Training a...
-
[6]
R. A. Bailey. 2008. Design of Comparative Experiments . Cambridge University Press
2008
-
[7]
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: From allocative to representational harms in machine learning. In SIGCIS Conference
2017
-
[9]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language ( Technology ) is Power : A Critical Survey of `` Bias '' in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 5454--5476, Online. Association for Computational Linguistics
-
[10]
Maheep Chaudhary, Ian Su, Nikhil Hooda, Nishith Shankar, Julia Tan, Kevin Zhu, Ryan Lagasse, Vasu Sharma, and Ashwinee Panda. 2025. https://doi.org/10.48550/arXiv.2509.13333 Evaluation Awareness Scales Predictably in Open-Weights Large Language Models . Preprint, arXiv:2509.13333
-
[11]
G. A. Cohen. 1989. On the Currency of Egalitarian Justice . Ethics, 99(4):906--944
1989
-
[12]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. https://doi.org/10.1145/3442188.3445924 BOLD : Dataset and Metrics for Measuring Biases in Open-Ended Language Generation . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT '21, pages 862--872...
arXiv 2021
-
[13]
Elizabeth Mitchell Elder and Matthew Hayes. 2023. https://doi.org/10.1086/723820 Signaling Race , Ethnicity , and Gender with Names : Challenges and Recommendations . The Journal of Politics, 85(2):764--770
-
[14]
Iason Gabriel. 2020. https://doi.org/10.1007/s11023-020-09539-2 Artificial Intelligence , Values , and Alignment . Minds and Machines, 30(3):411--437
-
[15]
S. Michael Gaddis. 2018. https://doi.org/10.1007/978-3-319-71153-9_1 An Introduction to Audit Studies in the Social Sciences . In S. Michael Gaddis, editor, Audit Studies : Behind the Scenes with Theory , Method , and Nuance , pages 3--44. Springer International Publishing, Cham
-
[16]
Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe
Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. 2024. https://doi.org/10.1177/23794607251320229 Auditing large language models for race & gender disparities: Implications for artificial intelligence-based hiring . Behavioral Science & Policy, 10(2):46--55
-
[17]
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. https://doi.org/10.1162/coli_a_00524 Bias and Fairness in Large Language Models : A Survey . Computational Linguistics, 50(3):1097--1179
-
[18]
Bufan Gao and Elisa Kreiss. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.342 Measuring Bias or Measuring the Task : Understanding the Brittle Nature of LLM Gender Biases . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages 6734--6750, Suzhou, China. Association for Computational Linguistics
-
[19]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, and 24 others. 2024. https://doi.org/10.48550/arXiv.2402.00838 OLMo : A...
-
[20]
Maarten Grootendorst. 2022. https://arxiv.org/abs/2203.05794 BERTopic : Neural topic modeling with a class-based TF-IDF procedure . arXiv preprint arXiv:2203.05794
Pith/arXiv arXiv 2022
-
[21]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. https://doi.org/10.1038/s41586-024-07856-5 AI generates covertly racist decisions about people based on their dialect . Nature, 633(8028):147--154
-
[22]
Mich \`e le Lamont and Vir \'a g Moln \'a r. 2002. The Study of Boundaries in the Social Sciences . Annual Review of Sociology, 28:167--195
2002
-
[23]
Louis Lippens. 2024. https://doi.org/10.1016/j.chbah.2024.100054 Computer says `no': Exploring systemic bias in ChatGPT using an audit approach . Computers in Human Behavior: Artificial Humans, 2(1):100054
arXiv 2024
-
[24]
Martin Lukk, Nora Kenworthy, Erik Schneiderhan, and Jeremy Snyder. 2025. https://doi.org/10.1002/nvsm.70041 Disrupting Philanthropy ? A Reality Check for Digital Crowdfunding . Journal of Philanthropy, 30(S1):e70041
-
[25]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. https://doi.org/10.48550/arXiv.2004.09456 StereoSet : Measuring stereotypical bias in pretrained language models . Preprint, arXiv:2004.09456
-
[26]
Joe Needham, Giles Edkins, Govind Pimpale, Henning Bartsch, and Marius Hobbhahn. 2025. https://doi.org/10.48550/arXiv.2505.23836 Large Language Models Often Know When They Are Being Evaluated . Preprint, arXiv:2505.23836
-
[27]
Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé III. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.413 “ You Gotta be a Doctor , Lin ” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 7268--7287. Associati...
-
[28]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. https://doi.org/10.18653/v1/2022.findings-acl.165 BBQ : A hand-built bias benchmark for question answering . In Findings of the Association for Computational Linguistics : ACL 2022 , pages 2086--2105, Dublin, Ireland. Associ...
-
[29]
Alejandro Salinas, Amit Haim, and Julian Nyarko. 2025. https://doi.org/10.48550/arXiv.2402.14875 What's in a Name ? Auditing Large Language Models for Race and Gender Bias . Preprint, arXiv:2402.14875
-
[30]
Erik Schneiderhan and Martin Lukk. 2023. GoFailMe : The Unfulfilled Promise of Digital Crowdfunding . Stanford University Press, Stanford
2023
-
[31]
Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli. 2023. https://doi.org/10.48550/arXiv.2312.03689 Evaluating and Mitigating Discrimination in Language Model Decisions . Preprint, arXiv:2312.03689
-
[32]
Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera , Ana Mar \'i a Mu \ n oz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, and Valentin Hofmann. 2026. https://doi.org/10.48550/arXiv.2601.18486 Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias . Pr...
-
[33]
Wim van Oorschot. 2000. https://doi.org/10.1332/0305573002500811 Who should get what, and why? On deservingness criteria and the conditionality of solidarity among the public . Policy & Politics, 28(1):33--48
-
[34]
Wim van Oorschot and Femke Roosma. 2017. Chapter 1: The Social Legitimacy of Targeted Welfare and Welfare Deservingness . In Wim Van Oorschot, Femke Roosma, Bart Meuleman, and Tim Reeskens, editors, The Social Legitimacy of Targeted Welfare , pages 3--34. Edward Elgar Publishing
2017
-
[35]
Franziska Weeber, Vera Neplenbroek, Jan Batzner, and Sebastian Pad \'o . 2026. https://doi.org/10.18653/v1/2026.acl-long.2079 One Persona , Many Cues , Different Results : How Sociodemographic Cues Impact LLM Personalization . In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pages 44892...
-
[36]
Bernard Weiner. 1985. An Attributional Theory of Achievement Motivation and Emotion . Psychological Review, 92(4):548--573
1985
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.