Pith. sign in

REVIEW 3 major objections 3 minor 35 references

Conflicting LLM bias results may be an artifact of audit design, not of the models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 16:56 UTC pith:LFZ5E2S3

load-bearing objection A genuinely careful audit-design benchmark; the name-cue caveat is real and disclosed, but the central claims hold up. the 3 major comments →

arxiv 2607.28934 v1 pith:LFZ5E2S3 submitted 2026-07-31 cs.CL cs.AIcs.CY

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

classification cs.CL cs.AIcs.CY
keywords LLM bias auditsresource allocationdistributive biasaudit designdeservingnessname-based demographic cuesbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that contradictory findings about whether large language models favor or penalize minorities in resource-allocation decisions can be explained by differences in how the models are audited, rather than by genuine differences among models. To show this, it builds FairFund-Bench, which varies the evaluation task (rating, ranking, allocating), the comparison context (single vs. multiple claimants), and whether demographic cues are obvious or disguised. Across 14 models, the benchmark reproduces the full range of previously reported results: models give a small edge to minority claimants when rating individuals alone, but penalize some minority groups when ranking claimants side by side. Demographic gaps are three to four times larger in disguised audits than in transparent ones, because transparent comparisons elicit near-universal equal splitting. The paper also finds that how a claimant's need is framed—especially whether the cause is external or self-inflicted—affects allocations far more than demographics and remains consistent across tasks and models. A sympathetic reader would care because the results imply that no single-format audit can support a general claim about LLM bias, and that current models are more strongly shaped by human-like deservingness judgments than by demographic attributes.

Core claim

The central claim is that audit format changes both the direction and the magnitude of measured demographic bias in LLM allocation decisions. Using 600 human-authored aid appeals calibrated against a large corpus of real crowdfunding campaigns, the benchmark crosses three tasks (Rate, Rank, Allocate), two comparison contexts (single vs. multi-stimulus), and two presentation modes (transparent minimal pairs vs. disguised bundles that co-vary scenario with demographic cues). The paper reports that models advantage ethnic minority claimants on the single-stimulus Rate task, disadvantage Asian and Hispanic claimants on the Rank task, and shift from near-universal equal splitting in transparent A

What carries the argument

The central object is FairFund-Bench, an audit instrument that systematically varies task format, comparison context, and presentation transparency. Its load-bearing component is the contrast between transparent bundles, which present minimal pairs differing only by name, and disguised bundles, which co-vary names with scenarios so that the demographic contrast is not obvious. Position rotations and Graeco-Latin square designs keep the focal contrast statistically identifiable across the bundle set. The paper also uses a five-level causal framing manipulation derived from the CARIN deservingness heuristic to measure how models respond to the controllability of need, and a name-pair validatio

Load-bearing premise

The results depend on the premise that differences in first and last names alone cause models to perceive the intended racial and gender groups, rather than reacting to other traits that happen to correlate with those names, such as class, age, or region.

What would settle it

Repeat the benchmark with group membership signaled by dialect or by explicit identity statements instead of names. If the direction reversal between the rating and ranking tasks, and the transparent–disguised gap in allocation, disappear or substantially change, then the central claims are artifacts of the name cue rather than properties of model bias.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim is right, single-format audits cannot support credible overall conclusions about whether a model is biased; the same model can appear positively biased, negatively biased, or unbiased depending on the task and presentation.
  • Minimal-pair audits understate the demographic disparities models produce in more deployment-like, disguised settings, so bias estimates from such audits should be treated as lower bounds.
  • Causal framing of need, not demographics, is the dominant driver of current model allocations, and this pattern is consistent across tasks and models.
  • Managers and regulators considering LLMs for welfare, lending, or hiring decisions should expect allocations to track human deservingness heuristics, which may punish stigmatized causes of need, rather than purely demographic fairness.
  • The benchmark's four-pillar framework can be adapted to other allocation domains, providing a way to separate a model's substantive bias from its consistency across audit conditions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One implication the paper leaves implicit: the audit-design sensitivity it documents likely extends to other allocation contexts, so any single-format audit of hiring, lending, or housing decisions should be read as design-dependent rather than as a model-level verdict.
  • The near-universal equal splitting in transparent bundles, whether caused by audit awareness or by similarity to training evaluations, suggests that current alignment makes models egalitarian mainly when bias is obvious; disguised, realistic prompts may not inherit that guarantee.
  • Because framing effects are stable and large while demographic effects are small and direction-reversing, a testable next step is to examine whether explicit allocation criteria (e.g., 'who would benefit most' vs. 'who is most deserving') shift the framing gradient in predictable ways; the paper's own wording-variant data hint that criterion naming enlarges framing effects.
  • The finding that small models show lower deservingness alignment than frontier models suggests that scaling or alignment choices affect how strongly models adopt human-like deservingness judgments, a pattern worth tracking as models evolve.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. FairFund-Bench is a benchmark for measuring how audit design choices affect apparent demographic and deservingness biases in LLM resource allocation. It varies three tasks (Rate, Rank, Allocate), single vs. multi-stimulus contexts, and transparent vs. disguised presentation, using 600 human-authored aid requests calibrated to a large GoFundMe corpus and 40 validated name pairs. Across 14 LLMs, the paper reports that the direction of demographic bias reverses by task: minorities are advantaged in single-stimulus ratings but some groups are disadvantaged in ranking. On the Allocate task, transparent demographic bundles elicit near-universal equal splitting, while disguised bundles produce 3–4 times larger dollar gaps. Causal framing of need produces much larger and more consistent effects than demographics, in line with human deservingness heuristics. The benchmark scores models on four pillars: demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency, with code and raw responses released.

Significance. If the results hold, this is a valuable methodological contribution. It offers a concrete explanation for contradictory prior LLM bias audits—design choices can flip the sign and magnitude of measured bias—and provides an extensible benchmark for future work. The paper's strengths are real: a careful factorial design with pre-fixed prompts, balanced rotations, temperature=0, explicit parse-failure handling, bootstrap CIs, and full release of raw responses and scoring code. The claim that framing effects dominate demographic effects is also important, as it redirects attention from identity cues to need narratives. The main risks are interpretational: the name-only demographic signal and the two-part transparent/disguised manipulation both bear directly on the central claims.

major comments (3)
  1. [§6, Appendix D, §4.2] The demographic-bias interpretation rests on the assumption that the 40 name pairs signal race/gender rather than correlated attributes such as class, age, region, or education. Appendix D matches pairs only on perceived competence and hardworkingness, and the manuscript itself concedes in §6 that name-based estimates may not hold under other cues. Because the headline direction-reversal and the transparent–disguised gap are between-name contrasts, they are vulnerable to models responding to non-racial name connotations differently across tasks. This is load-bearing, not a peripheral caveat. Please add a validation step—for example, re-estimating key contrasts with explicit identity statements on a subset, or with a second name set matched on additional perceived attributes—or systematically reframe the claims as effects of name-based demographic cues rather than demographic bias.
  2. [§3.3, Table 8, §4.2] The transparent–disguised manipulation is two-dimensional: transparent bundles hold scenario fixed, while disguised bundles co-vary the focal axis with scenario (Table 8). The finding that demographic dollar gaps are 3–4 times larger in disguised bundles is therefore attributable either to reduced transparency or to increased stimulus diversity, or both. The equal-splitting discussion in §6 addresses one alternative mechanism but not this design confound. Please add a design cell that orthogonally varies stimulus diversity (e.g., disguised bundles with non-focal scenario variation but overt demographic labels), or temper the causal claim that minimal-pair audits 'understate' disparities to acknowledge that the disguised condition also changes the informational content of the prompt.
  3. [§4.3, Figure 6] The headline that framing effects exceed demographic disparities by 'roughly an order of magnitude' compares signed framing contrasts (e.g., +$469 structural vs. self-cause) with mean absolute White–non-White gaps from a different bundle type. These metrics are not directly comparable as reported, and the phrase 'the smallest' is ambiguous. Please report a common dispersion measure for both factors—such as mean absolute deviation from the grand mean or variance-explained-style statistics—or explicitly state the comparison base (e.g., transparent versus disguised demographic gaps).
minor comments (3)
  1. [§4.2] The Hispanic Rank disadvantage is reported as −0.044 [−0.002, 0.090], whose confidence interval includes zero, yet the text groups Hispanic with Asian as 'some of the same groups.' Please qualify this as not statistically distinguishable from zero.
  2. [Appendix A, Wording variants] The finding that wording variants substantially change framing-effect magnitudes is reported only in the appendix. Given its relevance to the robustness of the §4.3 magnitude claims, a sentence in the main text or Limitations would be helpful.
  3. [Figure 5] The figure uses one line per model; with 14 models the legend is hard to read. Consider highlighting the median and the named outliers (Grok 4.20, DeepSeek V3.2) rather than plotting every line.

Circularity Check

0 steps flagged

FairFund-Bench's headline effects are observed model behaviors, not artifacts of the scoring definitions; no circular derivation chain found.

full rationale

I walked the derivation chain from stimulus construction through pillar scoring to the headline findings. The central results—direction reversal between Rate and Rank, larger disguised than transparent Allocate gaps, and framing effects exceeding demographic effects—are all estimated from raw model responses (ratings, ranks, dollar allocations) and are not forced by the scoring definitions. The P2 'deservingness alignment' pillar is explicitly defined as agreement with the CARIN ordering, so calling a high P2 'alignment' is a measurement, not a derivation; moreover, the framing-dominance claim in §4.3 rests on direct dollar contrasts, independent of P2. The transparent-vs-disguised comparison is also empirical: equal splitting is a model behavior, not an imposed code, and disguised bundles balance scenario against the focal axis via Graeco-Latin squares, so the race/gender contrasts are identified rather than constructed. The paper's self-citations (Lukk et al. 2025; Schneiderhan and Lukk 2023) are contextual and non-load-bearing. The limitations it flags—name-only demographic signaling (§6), author wording effects (§6), and the P1/P4 interpretation caveats (§7, Appendix H)—are construct-validity or measurement-interpretation concerns, not circularity, and the paper states them openly. No fitted parameter is later relabeled as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via citation. The benchmark is self-contained in the sense that its conclusions are direct observations of model behavior under a clearly specified instrument.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

No fitted constants enter the central results: the benchmark measures model behavior rather than deriving it. The ledger instead surfaces two hand-chosen analytic rules (the fixed Cohen's d yardstick and the P1 pool exclusion) and five domain assumptions the claims ride on, notably name-based demographic signaling, the framing-ladder mapping to CARIN, and the theory-bridge from CARIN to 'human deservingness'. The paper discloses all of these in the Limitations or appendices; none is an invented entity or a hidden fitted parameter.

free parameters (2)
  • Cohen's d standardization denominator = median of 14 within-model SDs, per pool x task, fixed across lineup
    Appendix H: chosen by hand 'so that a model cannot lower its apparent bias by being internally noisy'. This fixed yardstick determines P1/P3/P4 magnitudes; it is a transparent normalization choice rather than a fitted constant, but it is hand-selected and shapes the leaderboard.
  • P1 transparent-Allocate pool exclusion = transparent race/gender/intersectional Allocate pools removed from P1
    Appendix H: the same zero-spread that makes d undefined removes the three transparent-Allocate pools from P1, so the headline 'low demographic bias' (P1=0.03-0.08) excludes exactly the pools where equal splitting produced zero measured bias. Disclosed by the paper, which warns P1 understates disguised-prompt disparities.
axioms (5)
  • domain assumption Names from Elder & Hayes (2023) signal race and gender to LLMs without confounded connotations (competence, class, age, region).
    Central operational assumption for all demographic contrasts; Appendix D matches names on perceived competence/hardworkingness, and §6 concedes names alone are a thin cue and conclusions may not transfer to other cues.
  • domain assumption The five causal framing conditions form an ordinal Control/deservingness ladder per CARIN and differ only along that dimension.
    Framing paragraphs (Figure 2, Appendix E) alter the precipitating event (layoff vs missed shifts vs DUI), so blame/control is deliberately confounded with event type; the ladder is the manipulation, but its clean mapping to CARIN is an authorial construction the paper flags in §6.
  • domain assumption The CARIN deservingness ordering (van Oorschot & Roosma 2017) is a valid empirical stand-in for human deservingness judgments on these specific stimuli.
    P2 (§3.6) scores alignment with the CARIN-predicted sign of four framing contrasts; §6 concedes no human baseline exists on the actual appeals, so 'robustly reproduce human deservingness evaluations' (§4.3) relies on this theory-to-measurement bridge.
  • domain assumption Temperature=0 outputs with reasoning suppressed represent the allocation behavior of the models under assessment.
    §3.5/Appendix F: all models run at temperature=0, reasoning suppressed where a control exists; deployment behavior may differ with sampling and extended reasoning enabled, so the measured behavior is a specific inference configuration.
  • domain assumption GoFundMe corpus statistics (word counts, narrative features, topic base rates) validate that the hand-written templates resemble real aid requests.
    Calibration against 1.29M campaigns (Appendices B/C) backs template realism; the corpus is not released, and label agreement for some calibration fields is modest (redemption kappa=0.40, Table 5), which is reported transparently.

pith-pipeline@v1.3.0-daily-deepseek · 20802 in / 20510 out tokens · 188840 ms · 2026-08-03T16:56:14.402413+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

Figures

Figures reproduced from arXiv: 2607.28934 by Martin Lukk (University of Toronto).

Figure 1
Figure 1. Figure 1: FairFund-Bench pipeline. Benchmark construction (left) combines hand-written stimuli (calibrated against [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Stimulus template structure (Rent, Scenario [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Transparent vs. disguised bundles. Left panel [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Equal split rates for Allocate, by focal axis and [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 7
Figure 7. Figure 7: Mean Allocate dollars by framing condition, [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

35 extracted references · 6 canonical work pages

  1. [1]

    Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024. https://doi.org/10.18653/v1/2024.acl-short.37 Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race , Ethnicity , and Gender ? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics ( Volume 2: Short Papers ) , pag...

  2. [2]

    Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. 2025. https://doi.org/10.1093/pnasnexus/pgaf089 Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation . PNAS Nexus, 4(3):pgaf089

  3. [3]

    Lena Armstrong, Abbey Liu, Stephen MacNeil, and Dana \"e Metaxa. 2024. https://doi.org/10.1145/3689904.3694699 The Silicon Ceiling : Auditing GPT 's Race and Gender Biases in Hiring . In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms , Mechanisms , and Optimization , EAAMO '24, pages 1--18, New York, NY, USA. Association for Comp...

  4. [4]

    Griffiths

    Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L. Griffiths. 2025. https://doi.org/10.1073/pnas.2416228122 Explicitly unbiased large language models still form biased associations . Proceedings of the National Academy of Sciences, 122(8):e2416228122

  5. [5]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk , Nelson Elhage, Zac Hatfield-Dodds , Danny Hernandez, Tristan Hume, and 12 others. 2022. https://doi.org/10.48550/arXiv.2204.05862 Training a...

  6. [6]

    R. A. Bailey. 2008. Design of Comparative Experiments . Cambridge University Press

  7. [7]

    Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: From allocative to representational harms in machine learning. In SIGCIS Conference

  8. [9]

    Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language ( Technology ) is Power : A Critical Survey of `` Bias '' in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 5454--5476, Online. Association for Computational Linguistics

  9. [10]

    Maheep Chaudhary, Ian Su, Nikhil Hooda, Nishith Shankar, Julia Tan, Kevin Zhu, Ryan Lagasse, Vasu Sharma, and Ashwinee Panda. 2025. https://doi.org/10.48550/arXiv.2509.13333 Evaluation Awareness Scales Predictably in Open-Weights Large Language Models . Preprint, arXiv:2509.13333

  10. [11]

    G. A. Cohen. 1989. On the Currency of Egalitarian Justice . Ethics, 99(4):906--944

  11. [12]

    Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. https://doi.org/10.1145/3442188.3445924 BOLD : Dataset and Metrics for Measuring Biases in Open-Ended Language Generation . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT '21, pages 862--872...

  12. [13]

    Elizabeth Mitchell Elder and Matthew Hayes. 2023. https://doi.org/10.1086/723820 Signaling Race , Ethnicity , and Gender with Names : Challenges and Recommendations . The Journal of Politics, 85(2):764--770

  13. [14]

    Iason Gabriel. 2020. https://doi.org/10.1007/s11023-020-09539-2 Artificial Intelligence , Values , and Alignment . Minds and Machines, 30(3):411--437

  14. [15]

    Michael Gaddis

    S. Michael Gaddis. 2018. https://doi.org/10.1007/978-3-319-71153-9_1 An Introduction to Audit Studies in the Social Sciences . In S. Michael Gaddis, editor, Audit Studies : Behind the Scenes with Theory , Method , and Nuance , pages 3--44. Springer International Publishing, Cham

  15. [16]

    Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe

    Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. 2024. https://doi.org/10.1177/23794607251320229 Auditing large language models for race & gender disparities: Implications for artificial intelligence-based hiring . Behavioral Science & Policy, 10(2):46--55

  16. [17]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. https://doi.org/10.1162/coli_a_00524 Bias and Fairness in Large Language Models : A Survey . Computational Linguistics, 50(3):1097--1179

  17. [18]

    Bufan Gao and Elisa Kreiss. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.342 Measuring Bias or Measuring the Task : Understanding the Brittle Nature of LLM Gender Biases . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages 6734--6750, Suzhou, China. Association for Computational Linguistics

  18. [19]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, and 24 others. 2024. https://doi.org/10.48550/arXiv.2402.00838 OLMo : A...

  19. [20]

    Maarten Grootendorst. 2022. https://arxiv.org/abs/2203.05794 BERTopic : Neural topic modeling with a class-based TF-IDF procedure . arXiv preprint arXiv:2203.05794

  20. [21]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. https://doi.org/10.1038/s41586-024-07856-5 AI generates covertly racist decisions about people based on their dialect . Nature, 633(8028):147--154

  21. [22]

    Mich \`e le Lamont and Vir \'a g Moln \'a r. 2002. The Study of Boundaries in the Social Sciences . Annual Review of Sociology, 28:167--195

  22. [23]

    Louis Lippens. 2024. https://doi.org/10.1016/j.chbah.2024.100054 Computer says `no': Exploring systemic bias in ChatGPT using an audit approach . Computers in Human Behavior: Artificial Humans, 2(1):100054

  23. [24]

    Martin Lukk, Nora Kenworthy, Erik Schneiderhan, and Jeremy Snyder. 2025. https://doi.org/10.1002/nvsm.70041 Disrupting Philanthropy ? A Reality Check for Digital Crowdfunding . Journal of Philanthropy, 30(S1):e70041

  24. [25]

    Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. https://doi.org/10.48550/arXiv.2004.09456 StereoSet : Measuring stereotypical bias in pretrained language models . Preprint, arXiv:2004.09456

  25. [26]

    Joe Needham, Giles Edkins, Govind Pimpale, Henning Bartsch, and Marius Hobbhahn. 2025. https://doi.org/10.48550/arXiv.2505.23836 Large Language Models Often Know When They Are Being Evaluated . Preprint, arXiv:2505.23836

  26. [27]

    You Gotta be a Doctor , Lin

    Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé III. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.413 “ You Gotta be a Doctor , Lin ” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 7268--7287. Associati...

  27. [28]

    Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. https://doi.org/10.18653/v1/2022.findings-acl.165 BBQ : A hand-built bias benchmark for question answering . In Findings of the Association for Computational Linguistics : ACL 2022 , pages 2086--2105, Dublin, Ireland. Associ...

  28. [29]

    Alejandro Salinas, Amit Haim, and Julian Nyarko. 2025. https://doi.org/10.48550/arXiv.2402.14875 What's in a Name ? Auditing Large Language Models for Race and Gender Bias . Preprint, arXiv:2402.14875

  29. [30]

    Erik Schneiderhan and Martin Lukk. 2023. GoFailMe : The Unfulfilled Promise of Digital Crowdfunding . Stanford University Press, Stanford

  30. [31]

    Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli. 2023. https://doi.org/10.48550/arXiv.2312.03689 Evaluating and Mitigating Discrimination in Language Model Decisions . Preprint, arXiv:2312.03689

  31. [32]

    Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera , Ana Mar \'i a Mu \ n oz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, and Valentin Hofmann. 2026. https://doi.org/10.48550/arXiv.2601.18486 Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias . Pr...

  32. [33]

    Wim van Oorschot. 2000. https://doi.org/10.1332/0305573002500811 Who should get what, and why? On deservingness criteria and the conditionality of solidarity among the public . Policy & Politics, 28(1):33--48

  33. [34]

    Wim van Oorschot and Femke Roosma. 2017. Chapter 1: The Social Legitimacy of Targeted Welfare and Welfare Deservingness . In Wim Van Oorschot, Femke Roosma, Bart Meuleman, and Tim Reeskens, editors, The Social Legitimacy of Targeted Welfare , pages 3--34. Edward Elgar Publishing

  34. [35]

    Franziska Weeber, Vera Neplenbroek, Jan Batzner, and Sebastian Pad \'o . 2026. https://doi.org/10.18653/v1/2026.acl-long.2079 One Persona , Many Cues , Different Results : How Sociodemographic Cues Impact LLM Personalization . In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pages 44892...

  35. [36]

    Bernard Weiner. 1985. An Attributional Theory of Achievement Motivation and Emotion . Psychological Review, 92(4):548--573