Pith. sign in

REVIEW 5 major objections 5 minor 73 references

This paper shows that large language models systematically distort race and gender in generated occupational personas, underrepresenting White and Black workers and overrepresenting Hispanic and Asian workers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 08:19 UTC pith:UFVC37QN

load-bearing objection A large, careful audit of text personas; the cross-model 'shared sources' claim is softer than the abstract suggests, but the measurement itself holds up and deserves referee time. the 5 major comments →

arxiv 2510.21011 v3 pith:UFVC37QN submitted 2025-10-23 cs.HC cs.AIcs.CY

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

classification cs.HC cs.AIcs.CY
keywords representational biaspersona generationoccupational stereotypesrace and genderalgorithmic auditgenerative AIdemographic distortionlarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that large language models do not simply reproduce occupational demographics when asked to generate a persona: they systematically reshape them. Across more than 1.5 million generated profiles for 41 U.S. occupations and four models, White workers are underrepresented by about 31 percentage points and Black workers by about 9, while Hispanic workers are overrepresented by about 17 points and Asian workers by about 12. The models compress each occupation toward a dominant profile rather than representing population-level variation, and they amplify existing gender skews, producing too few women in male-dominated jobs and too many in female-dominated ones. Because the same directions of distortion recur across models from different companies and countries, the authors argue the bias is structural, not model-specific, and that auditing generative systems requires evaluating how synthetic populations reshape demographic visibility.

Core claim

The core discovery is that persona generation imposes directional, group-specific distortions rather than mirroring reality. Using regression of model-generated proportions on real-world labor-market proportions, the paper finds a systematic shift: White and Black workers are suppressed, Hispanic and Asian workers elevated; and a stereotype slope: gender skews are exaggerated in an S-shaped curve, Black representation is flattened to near-erasure in typical occupations, and Asian representation is concentrated where it is already high. Extremes include housekeepers portrayed as nearly 100% Hispanic across all four models and Black workers nearly absent outside a few roles. The patterns recur

What carries the argument

The analytic engine is a shift/exaggeration decomposition: a logistic regression of the model's group share on the real-world group share, with the intercept (alpha) measuring systematic over- or underrepresentation and the slope (beta) measuring amplification or dampening of occupational skews. Estimates are reported on a centered logit scale so intercepts translate into percentage-point shifts at the median occupation. The paper pairs this with a five-way typology of representation—reality, stereotype exaggeration, stereotype reduction, overrepresentation, and underrepresentation—that turns raw divergences into interpretable bias patterns.

Load-bearing premise

The central claim that the biases are structural across models rests on a single fixed prompt being representative; if paraphrasing the prompt or changing system instructions reshuffles the demographic distributions, the cross-model convergence could be an artifact of that shared prompt rather than shared training-data bias.

What would settle it

Generate the same 41 occupations with multiple paraphrased prompts and varied default settings, then check whether White/Black underrepresentation and Hispanic/Asian overrepresentation persist in sign and approximate magnitude; if the pattern flips or disappears under some phrasings, the claim of shared structural bias fails and the observed convergence is prompt-induced.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the models really impose these distortions, any persona-based workflow—UX research, product development, creative writing, simulation—inherits a systematically skewed picture of who works in a given job.
  • The near-erasure of Black workers at typical occupational levels means generated datasets underrepresent Black professionals outside stereotyped roles, which can affect downstream tasks built on synthetic populations.
  • Provider choice materially changes the midpoint and concentration of representation—for example, Gemini lowers women while the other three lift them—so safety branding alone does not predict bias direction or magnitude.
  • Because the same directions recur across all four models, the bias appears to live in shared training data and generation conventions, not in a single vendor's post-training choices.
  • Auditing tools should report representation relative to external baselines and expose alpha/beta summaries, rather than treating single default personas as neutral.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct multi-prompt replication—same occupations, five to ten paraphrased prompts, varied temperature and system instructions—would test whether the shared backbone is training-data bias or shared prompt sensitivity; the paper's own limitation note leaves this open.
  • The same intercept/slope decomposition could be run on other persona fields (names, motivations, biographies, salaries), where the paper suspects additional stereotypes; if the demographic distortions and narrative stereotypes align, the representational harm compounds.
  • If these patterns generalize to other national labor markets, the direction of distortion may be culturally specific—different groups may be over- or underrepresented depending on the benchmark population, so the audit method matters more than the specific numbers.
  • The authors' notion that not all deviations are harmful could be operationalized by separating stereotype-reinforcing distortions (Hispanic housekeepers, Black security guards) from stereotype-challenging ones (women in engineering), which would require a normative judgment beyond the paper's descriptive frame.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This paper audits racial and gender representation in LLM-generated occupational personas. The authors prompted four LLMs (GPT-4, Gemini 2.5, DeepSeek V3.1, Mistral-medium) with a fixed prompt for 41 U.S. occupations, collecting over 1.5 million structured personas, and compare the resulting demographic distributions with 2023 BLS CPS data. They report average percentage-point deviations (e.g., White -31, Black -9, Hispanic +17, Asian +12), fit centered-logit regressions to decompose systematic shifts (intercepts) from stereotype exaggeration/reduction (slopes), and find gender S-curves, Black near-erasure, and overrepresentation of Hispanic/Asian workers. The paper argues that these distortions recur across models and thus reflect shared structural sources of bias, and proposes HCI design implications for making representation visible in persona-generation tools.

Significance. If the results hold, this is a large-scale, externally benchmarked measurement of representational bias in text-based persona generation, extending prior image-search audits to a modality widely used in UX and product design. The study's strengths include the 1.5M+ sample size, a public external benchmark, multiple robustness specifications (trimmed and robust regression), FDR correction, and posted prompt/schema/code for reproducibility. The shift/exaggeration decomposition is a useful framework for characterizing distortions. However, the cross-model convergence inference is confounded by the single shared prompt, and several headline numbers are inconsistent with the paper's own regression tables. These issues are fixable, but they affect the interpretability of the central claims.

major comments (5)
  1. [Abstract; §4.1; Tables 3 and 5] The headline numbers in the abstract and §4.1 (Hispanic +17 pp, Asian +12 pp) do not match the regression tables reported later. Table 3 gives median-centered robust shifts of +9.4 pp and +3.4 pp for Hispanic and Asian, respectively. Table 5 shows per-model Asian shifts of -2.83 (ChatGPT), -3.71 (DeepSeek), +12.62 (Gemini), and -4.12 (Mistral); only Gemini is positive. If the abstract reports simple unweighted occupation-level averages, that estimator should be stated explicitly and reported alongside the regression estimates. As currently written, the paper presents two different quantitative summaries of its central finding without explaining the discrepancy.
  2. [Abstract; §3.3; §4.1; §5.6] The inference that recurrent distortions across GPT-4, Gemini, DeepSeek, and Mistral indicate 'shared structural sources of bias' is not fully warranted by the design. All four models were elicited with the same single prompt template (§3.3), and §5.6 concedes that shared prompt sensitivities cannot be ruled out. The model-level results also show non-uniformity: DeepSeek's Hispanic shift is near zero, and ChatGPT, DeepSeek, and Mistral all have negative Asian shifts (Table 5). The strongest cross-model pattern is Black underrepresentation, not the full set of headline patterns. Please either add a multi-prompt robustness check (e.g., 5-10 paraphrases) or restrict the abstract/conclusions to the specific elicitation regime and to the patterns that are actually uniform across models.
  3. [Abstract; §4.2; Tables 2 and 3] The abstract's characterization that models 'generate demographics with less variation than real-world data, functionally compressing each occupation toward a dominant demographic profile' is contradicted by the reported regression slopes. A logit slope of <1 would indicate compression, but Table 2 reports a positive slope deviation for women (+0.17) and Table 3 reports positive deviations for Asian (+2.10) and White (+1.13), indicating amplification of existing skews. Only Black has a negative slope deviation (-0.89), which reflects near-erasure rather than simple compression. The 'less variation' claim needs to be reconciled with the mostly amplification-type slopes shown in the paper's own regression results.
  4. [§3.2; §3.3; §5.6] The methods state that models could assign one or more racial categories, 'enabling multi-racial identities,' but the BLS benchmark categories are single-race/ethnicity groups, and the Limitations section lists mixed-race individuals as excluded. The paper does not describe how multi-racial persona labels were coded into the White/Black/Asian/Hispanic comparison. If a persona tagged as both Black and White is counted in both groups, the shares are no longer on the same scale as the BLS categories; if such personas were discarded or converted, the rule should be stated. This affects every racial point estimate and should be clarified.
  5. [Tables 2–5] The 'β(logit)' column appears to be mislabeled or misprinted. In Table 2, Women's β(logit) is reported as 13.26, while the text reports a slope deviation of +0.17 (implying β ≈ 1.17). Table 4 reports values such as 27.61 and 6.81 alongside slope deviations of +0.19 and +0.05, which are not on the same scale. As printed, readers cannot verify the slope estimates. Please correct the decimal formatting or column headings and re-run or re-report the table values so that the reported β(logit) values are consistent with the stated slope deviations.
minor comments (5)
  1. [§2.3] The sentence 'Collectively, this evidence shows that representational bias is not confined to images but extends into text...' is duplicated verbatim within the same section.
  2. [References] The reference list contains duplicate or broken numbering: [18] appears for both Cheng et al. and Cheryan et al., and the full reference list is repeated at the end of the manuscript. Please clean up the bibliography.
  3. [Title page/front matter] The manuscript retains ACM template placeholders such as 'Conference acronym ’XX', 'Received 20 February 2007', and the repeated 'This work is currently under peer review' notice. These should be removed or updated.
  4. [Figure 4 caption] The caption says 'ChatGPT = circle' but the legend elsewhere refers to GPT-4; use a single consistent model-name scheme across figures and text.
  5. [§3.5] The paper should state whether the pooled regressions in Tables 2 and 3 account for the repeated occupations across models (e.g., cluster-robust standard errors by occupation). This is especially relevant because DeepSeek has a smaller sample size.

Circularity Check

0 steps flagged

No significant circularity: the audit compares LLM personas to an external BLS benchmark, and its alpha/beta estimates are descriptive fits rather than assumptions used to define the outcome.

full rationale

The paper's central derivation is an empirical measurement: LLM-generated personas are compared against U.S. Bureau of Labor Statistics data, and systematic shifts and stereotype slopes are estimated by regressing model proportions on BLS proportions. The benchmark is external to the models and to the authors, and the regression coefficients (alpha and beta) are descriptive summaries of the observed differences, not fitted parameters that are later renamed as predictions or used to define the outcome. The paper explicitly cautions that BLS is 'a descriptive baseline rather than a normative ideal' (§2.5) and repeats this in §5.6, so the comparison is not circular. The five-model framework is an interpretive taxonomy applied after measurement, not an input that guarantees the findings. No self-citation chain is load-bearing: the prior audit studies cited (Kay et al. 2015, Metaxa et al. 2021) are external works, and the paper's own contributions do not depend on an unverified uniqueness theorem or ansatz from the authors' prior work. The acknowledged limitation that all models received a single fixed prompt and that cross-model agreement 'could also stem from shared prompt sensitivities' (§5.6) is a construct-validity and generalization concern, not a circularity concern, because it does not make the measured deviations equal to the paper's assumptions. The paper is self-contained against an external benchmark, so the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The paper introduces no free parameters, no new entities, and no fitted constants that define the outcome. Its assumptions are all domain assumptions about data representativeness and the elicitation regime. The key fragility is the single-prompt assumption, since the cross-model convergence claim depends on it.

axioms (5)
  • domain assumption BLS Current Population Survey Annual Averages 2023 accurately represents U.S. occupational demographics for the four compared groups.
    The entire benchmark comparison rests on BLS being the ground truth for real-world distributions. Invoked in §3.2.
  • domain assumption LLM 'gender' and 'ethnicity' fields map onto BLS binary-gender and White/Black/Asian/Hispanic categories, including multiracial assignments.
    Model outputs are parsed as demographic categories and compared directly to BLS proportions. Multiracial and intersectional identities are excluded; this comparability assumption is stated in §3.2.
  • domain assumption A single fixed prompt with default API temperature is a representative elicitation of typical LLM persona use.
    All conclusions about systematic bias are measured under one prompt. The authors acknowledge this in §3.3 and §5.6, noting results are not necessarily prompt-invariant.
  • domain assumption The centered-logit linear regression is an adequate model of the relationship between LLM and BLS proportions.
    The shift (alpha) and slope (beta) decomposition assumes a logistic-linear functional form; the authors note alternative specifications could change magnitudes (§5.6).
  • domain assumption Independent context windows and no balancing instructions cause the models to reveal their internal training distributions rather than deliberately balanced outputs.
    Stated in §3.3: 'Requests were issued in independent context windows to prevent models from balancing outputs across prompts.'

pith-pipeline@v1.3.0-alltime-deepseek · 28734 in / 10043 out tokens · 133722 ms · 2026-08-04T08:19:53.078466+00:00 · methodology

0 comments
read the original abstract

As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational biases is critical. We audit over 1.5 million occupational personas generated by four major large language models (GPT-4, Gemini 2.5, DeepSeek V3.1, and Mistral-medium) across 41 U.S. occupations. Comparing these personas against U.S. Bureau of Labor Statistics (BLS) data, we find that models generate demographics with less variation than real-world data, functionally compressing each occupation toward a dominant demographic profile rather than representing population-level variation. A shift/exaggeration decomposition reveals the structure of these distortions: White (-31 percentage points) and Black (-9 pp) workers are consistently underrepresented, while Hispanic (+17 pp) and Asian (+12 pp) workers are overrepresented, with stereotype exaggeration amplifying existing occupational segregation. These distortions are often extreme, including near-total portrayals of housekeepers as Hispanic and the near-erasure of Black workers from many occupations. Because these patterns recur across models with different institutional and cultural origins, they suggest shared structural sources of bias rather than model-specific artifacts. We argue that auditing generative AI requires evaluation frameworks that examine how synthetic populations systematically reshape demographic visibility across social roles.

Figures

Figures reproduced from arXiv: 2510.21011 by Aadi Sudan, Arnav Dixit, David C. Anastasiu, Ilona van der Linden, Kai Lukoff, Sahana Kumar, Smruthi Danda.

Figure 2
Figure 2. Figure 2: Conceptual illustration of how different bias patterns can appear when comparing large language model (LLM) outputs to BLS [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Differences between four large language models’ representations of gender in occupational prompts compared to BLS [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Differences between four large language models’ representations of racial groups in occupational prompts compared to BLS [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Systematic gender representation pooled across all models. Each point is an occupation, plotting BLS percentage of women [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Systematic racial representation pooled across all models. Each point is an occupation, plotting BLS percentages of racial [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Systematic gender representation by model. Each panel plots BLS percentages of women in occupations against LLM outputs, [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

73 extracted references · 14 canonical work pages

  1. [1]

    Rohaid Ali, Oliver Y Tang, Ian D Connolly, Hael A Abdulrazeq, Fatima N Mirza, Rachel K Lim, Benjamin R Johnston, Michael W Groff, Theresa Williamson, Konstantina Svokos, Tiffany J Libby, John H Shin, Ziya L Gokaslan, Curtis E Doberstein, James Zou, and Wael F Asaad. 2023. The face of a surgeon: An analysis of demographic representation in three leading ar...

  2. [2]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. 2019. Guidelines for Human-AI Interaction. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA. doi:10.1145/3290...

  3. [3]

    Oghenemaro Anuyah, Ruyuan Wan, Cornelius Adejoro, Tom Yeh, Ronald Metoyer, and Karla Badillo-Urquiola. 2023. Cultural considerations in AI systems for the global south: A systematic review. InProceedings of the 4th African Human Computer Interaction Conference. ACM, New York, NY, USA, 125–134. doi:10.1145/3628096.3629046

  4. [4]

    Lena Armstrong, Abbey Liu, Stephen MacNeil, and Danaë Metaxa. 2024. The silicon ceiling: Auditing GPT’s race and gender biases in hiring. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. ACM, New York, NY, USA, 1–18. doi:10.1145/3689904.3694699

  5. [5]

    Association for Computing Machinery. 2023. ACM Policy on Authorship. https://www.acm.org/publications/policies/new-acm-policy-on-authorship. Approved by the ACM Publications Board on April 20, 2023. Accessed: 2025-09-11

  6. [6]

    Andrew C Baker, David F Larcker, Charles G McCLURE, Durgesh Saraph, and Edward M Watts. 2024. Diversity washing.Journal of accounting research62, 5 (Dec. 2024), 1661–1709. doi:10.1111/1475-679x.12542

  7. [7]

    Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W Duncan Wadsworth, and Hanna Wallach. 2021. Designing disaggregated evaluations of AI systems: Choices, considerations, and tradeoffs. InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society. ACM, New York, NY, USA. doi:10.1145/34617...

  8. [8]

    Eric P S Baumer. 2017. Toward human-centered algorithm design.Big data & society4, 2 (Dec. 2017), 205395171771885. doi:10.1177/2053951717718854

  9. [9]

    Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society. Series B, Statistical methodology57, 1 (Jan. 1995), 289–300. doi:10.1111/j.2517-6161.1995.tb02031.x

  10. [10]

    Michael Bernstein, Angèle Christin, Jeffrey Hancock, Tatsunori Hashimoto, Chenyan Jia, Michelle Lam, Nicole Meister, Nathaniel Persily, Tiziano Piccardi, Martin Saveski, Jeanne Tsai, Johan Ugander, and Chunchen Xu. 2023. Embedding societal values into social media algorithms.Journal of online trust & safety2, 1 (Sept. 2023). doi:10.54501/jots.v2i1.148

  11. [11]

    Marianne Bertrand and Sendhil Mullainathan. 2004. Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination.American Economic Review94, 4 (Sept. 2004), 991–1013. doi:10.1257/0002828042002561

  12. [12]

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale. In2023 ACM Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, USA, 1493–150...

  13. [13]

    Francine D Blau and Lawrence M Kahn. 2017. The gender wage gap: Extent, trends, and explanations.Journal of Economic Literature55, 3 (Sept. 2017), 789–865. doi:10.1257/jel.20160995

  14. [14]

    Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. InProceedings of the 30th International Conference on Neural Information Processing Systems (NIPS’16). Curran Associates Inc., Red Hook, NY, USA, 4356–4364. doi:10.5555/3157382.3157584

  15. [15]

    Jane Castleman and Aleksandra Korolova. 2025. Aligned Generative Models Exhibit Adultification Bias. https://eclecticai.substack.com/p/aligned- generative-models-exhibit. Accessed: 2025-6-30

  16. [16]

    Christopher N Chapman and Russell P Milham. 2006. The personas’ new clothes: Methodological and practical arguments against a popular method. Proceedings of the Human Factors and Ergonomics Society ... Annual Meeting. Human Factors and Ergonomics Society. Annual Meeting50, 5 (Oct. 2006), 634–636. doi:10.1177/154193120605000503

  17. [17]

    Yuen Chen, Vethavikashini Chithrra Raghuram, Justus Mattern, Rada Mihalcea, and Zhijing Jin. 2025. Causally testing gender bias in LLMs: A case study on occupational bias. InFindings of the Association for Computational Linguistics: NAACL 2025. Association for Computational Linguistics, Stroudsburg, PA, USA, 4984–5004. doi:10.18653/v1/2025.findings-naacl.281

  18. [18]

    Sapna Cheryan, Sianna A Ziegler, Amanda K Montoya, and Lily Jiang. 2017. Why are some STEM fields more gender balanced than others? Psychological bulletin143, 1 (Jan. 2017), 1–35. doi:10.1037/bul0000052

  19. [19]

    2004.The inmates are running the asylum: Why high tech products drive us crazy and how to restore the sanity(2 ed.)

    Alan Cooper. 2004.The inmates are running the asylum: Why high tech products drive us crazy and how to restore the sanity(2 ed.). Sams Publishing, Indianapolis, IN. doi:10.5555/984201

  20. [20]

    Nilanjana Dasgupta. 2011. Ingroup experts and peers as social vaccines who inoculate the self-concept: The stereotype inoculation model. Psychological inquiry22, 4 (Oct. 2011), 231–246. doi:10.1080/1047840x.2011.607313

  21. [21]

    Irune del Rio Gabiola. 2013. Globalizing the care chain: Representations of Latinas in maid in America.Chasqui- revista De Literatura Latinoamericana42 (May 2013), 119–130. https://www.jstor.org/stable/43589516?casa_token=482AB_ _sJUQAAAAA:UqagNKlqQWme3zPexzUCCPkZSIjDZ9IwDYRT9ZwCRVe6CVtQ9nO_Ffpk7JaleY4NBZTvWby0v1xhmfFpX_CahRP4M_ nYsY8D33DBuQGnKtDxkQaURNk ...

  22. [22]

    2023.Data feminism

    Catherine D’Ignazio and Lauren F Klein. 2023.Data feminism. MIT Press, London, England. https://www.researchgate.net/profile/Emily- Yarrow-5/publication/363775476_BOOK_REVIEW_Data_Feminism_By_D’Ignazio_Catherine_and_Klein_Lauren_F_Cambridge_Massachusetts_ and_London_England_The_MIT_Press_2020_ISBN_978-0-262-04400-4/links/63ea85744dcb750da757116a/BOOK-REVI...

  23. [23]

    Upol Ehsan, Pradyumna Tambwekar, Larry Chan, Brent Harrison, and Mark O Riedl. 2019. Automated rationale generation: a technique for explainable AI and its effects on human perceptions. InProceedings of the 24th International Conference on Intelligent User Interfaces. ACM, New York, NY, USA. doi:10.1145/3301275.3302316

  24. [24]

    Motahhare Eslami, Aimee Rickman, Kristen Vaccaro, Amirhossein Aleyasen, Andy Vuong, Karrie Karahalios, Kevin Hamilton, and Christian Sandvig

  25. [25]

    Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. 2024. Bias of AI-generated content: an examination of news produced by large language models.Scientific reports14, 1 (March 2024), 5224. doi:10.1038/s41598-024-55686-2

  26. [26]

    2025.Top Generative AI Chatbots by Market Share – August 2025

    First Page Sage. 2025.Top Generative AI Chatbots by Market Share – August 2025. Technical Report. First Page Sage. https://firstpagesage.com/ reports/top-generative-ai-chatbots/

  27. [27]

    Batya Friedman and Helen Nissenbaum. 1996. Bias in computer systems.ACM transactions on information systems14, 3 (July 1996), 330–347. doi:10.1145/230538.230561

  28. [28]

    2025.AI Safety Index: Summer 2025 Report

    Future of Life Institute, Dylan Hadfield-Menell, Jessica Newman, Tegan Maharaj, Sneha Revanur, Stuart Russell, and David Krueger. 2025.AI Safety Index: Summer 2025 Report. Technical Report. Future of Life Institute. https://futureoflife.org/index

  29. [29]

    Xiao Ge, Chunchen Xu, Daigo Misaki, Hazel Rose Markus, and Jeanne L Tsai. 2024. How culture shapes what people want from AI. InProceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–15. doi:10.1145/3613904.3642660

  30. [30]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets.Commun. ACM64, 12 (Dec. 2021), 86–92. doi:10.1145/3458723

  31. [31]

    R Stuart Geiger, Flynn O’Sullivan, Elsie Wang, and Jonathan Lo. 2025. Asking an AI for salary negotiation advice is a matter of concern: Controlled experimental perturbation of ChatGPT for protected and non-protected group discrimination on a contextual task with no clear ground truth answers.PloS one20, 2 (Feb. 2025), e0318500. doi:10.1371/journal.pone.0318500

  32. [32]

    Anna M Gorska and Dariusz Jemielniak. 2023. The invisible women: uncovering gender bias in AI-generated images of professionals.Feminist media studies23, 8 (Nov. 2023), 4370–4375. doi:10.1080/14680777.2023.2263659

  33. [33]

    Nico Grant. 2024. Google Chatbot’s A.I. Images Put People of Color in Nazi-Era Uniforms.The New York Times(Feb. 2024). https://www.nytimes. com/2024/02/22/technology/google-gemini-german-uniforms.html

  34. [34]

    J Brian Gray. 1994. Regression with graphics: A second course in applied statistics.Technometrics: a journal of statistics for the physical, chemical, and engineering sciences36, 1 (Feb. 1994), 112–113. doi:10.1080/00401706.1994.10485408

  35. [35]

    Douglas Guilbeault, Solène Delecourt, Tasker Hull, Bhargav Srinivasa Desikan, Mark Chu, and Ethan Nadler. 2024. Online images amplify gender bias.Nature626, 8001 (Feb. 2024), 1049–1055. doi:10.1038/s41586-024-07068-x

  36. [36]

    Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA. doi:10.1145/3290605.3300830

  37. [37]

    Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. 2022. Evaluation gaps in machine learning practice. In2022 ACM Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, USA. doi:10.1145/3531146.3533233

  38. [38]

    Hyejun Jeong, Shiqing Ma, and Amir Houmansadr. 2024. Bias similarity across Large Language Models.arXiv [cs.LG](Oct. 2024). arXiv:2410.12010 [cs.LG] doi:10.48550/arXiv.2410.12010

  39. [39]

    Soon-Gyo Jung, Joni Salminen, Kholoud Khalil Aldous, and Bernard J Jansen. 2025. PersonaCraft: Leveraging language models for data-driven persona development.International journal of human-computer studies197, 103445 (March 2025), 103445. doi:10.1016/j.ijhcs.2025.103445

  40. [40]

    Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020. Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA. doi:10.1145/3313831.3376219

  41. [41]

    M Kay, C Matuszek, and S A Munson. 2015. Unequal representation and gender stereotypes in image search results for occupations.Proceedings of the 33rd annual acm(2015). doi:10.1145/2702123.2702520

  42. [42]

    Min Kyung Lee, Daniel Kusbit, Anson Kahng, Ji Tae Kim, Xinran Yuan, Allissa Chan, Daniel See, Ritesh Noothigattu, Siheon Lee, Alexandros Psomas, and Ariel D Procaccia. 2019. WeBuildAI: Participatory framework for algorithmic governance.Proceedings of the ACM on human-computer interaction3, CSCW (Nov. 2019), 1–35. doi:10.1145/3359283

  43. [43]

    Sang Won Lee, Mary Morcos, Dong Won Lee, and Jason Young. 2024. Demographic representation of generative artificial intelligence images of physicians.JAMA network open7, 8 (Aug. 2024), e2425993. doi:10.1001/jamanetworkopen.2024.25993

  44. [44]

    Ang Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng. 2025. LLM Generated Persona is a Promise with a Catch.arXiv [cs.CL](March 2025). arXiv:2503.16527 [cs.CL] doi:10.48550/arXiv.2503.16527

  45. [45]

    Guoying Li. 2011. Robust Regression. InExploring Data Tables, Trends, and Shapes. John Wiley & Sons, Inc., Hoboken, NJ, USA, 281–343. doi:10.1002/9781118150702.ch8 This work is currently under peer review for publication. This version is a preprint and has not undergone formal peer review. MANUSCRIPT 28 van der Linden et al

  46. [46]

    someone like me can be successful

    Penelope Lockwood. 2006. “someone like me can be successful”: Do college students need same-gender role models?Psychology of women quarterly 30, 1 (March 2006), 36–46. doi:10.1111/j.1471-6402.2006.00260.x

  47. [47]

    Li Lucy and David Bamman. 2021. Gender and representation bias in GPT-3 generated stories. InProceedings of the Third Workshop on Narrative Understanding, Nader Akoury, Faeze Brahman, Snigdha Chaturvedi, Elizabeth Clark, Mohit Iyyer, and Lara J Martin (Eds.). Association for Computational Linguistics, Stroudsburg, PA, USA, 48–55. doi:10.18653/v1/2021.nuse-1.5

  48. [48]

    Rohin Manvi, Samar Khanna, Marshall Burke, David Lobell, and Stefano Ermon. 2024. Large Language Models are Geographically Biased.arXiv [cs.CL](Feb. 2024). arXiv:2402.02680 [cs.CL] doi:10.48550/arXiv.2402.02680

  49. [49]

    Tara Matthews, Tejinder Judge, and Steve Whittaker. 2012. How do designers and user experience professionals actually perceive and use personas?. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1219–1228. doi:10.1145/2207676.2208573

  50. [50]

    Danaë Metaxa, Michelle A Gan, Su Goh, Jeff Hancock, and James A Landay. 2021. An image of society: Gender and racial representation and impact in image search results for occupations.Proceedings of the ACM on human-computer interaction5, CSCW1 (April 2021), 1–23. doi:10.1145/3449100

  51. [51]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. InProceedings of the Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, USA, 220–229. doi:10.1145/3287560.3287596

  52. [52]

    2004.Engaging Personas and Narrative Scenarios

    Lene Nielsen. 2004.Engaging Personas and Narrative Scenarios. Working Paper 6448. Copenhagen Business School, København, Denmark. https://research.cbs.dk/en/publications/engaging-personas-and-narrative-scenarios Final published version available via CBS Research Portal

  53. [53]

    Ihudiya Finda Ogbonnaya-Ogburu, Angela D R Smith, Alexandra To, and Kentaro Toyama. 2020. Critical Race Theory for HCI. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–16. doi:10.1145/3313831.3376392

  54. [54]

    Maria Olsson and Sarah E Martiny. 2018. Does exposure to counterstereotypical role models influence girls’ and women’s gender stereotypes and career choices? A review of social psychological research.Frontiers in psychology9 (Dec. 2018), 2264. doi:10.3389/fpsyg.2018.02264

  55. [55]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA. doi:10.1145/3586183.3606763

  56. [56]

    Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Wortman Vaughan, and Hanna Wallach. 2021. Manipulating and measuring model interpretability. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA. doi:10.1145/3411764.3445315

  57. [57]

    Inioluwa Deborah Raji and Joy Buolamwini. 2019. Actionable auditing: Investigating the impact of publicly naming biased performance results of commercial AI products. InProceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. ACM, New York, NY, USA. doi:10.1145/3306618. 3314244

  58. [58]

    Joni Salminen, João M Santos, Soon-Gyo Jung, and Bernard J Jansen. 2024. Picturing the fictitious person: An exploratory study on the effect of images on user perceptions of AI-generated personas.Computers in Human Behavior: Artificial Humans2, 1 (Jan. 2024), 100052. doi:10.1016/j.chbah. 2024.100052

  59. [59]

    Vanessa Sattele and Centro de Investigaciones de Diseño Industrial, UNAM. 2024. Generating user personas with AI: Reflecting on its implications for design. InProceedings of DRS. Design Research Society. doi:10.21606/drs.2024.1024

  60. [60]

    Morgan Klaus Scheuerman, Jacob M Paul, and Jed R Brubaker. 2019. How computers see gender: An evaluation of gender classification in commercial facial analysis services.Proceedings of the ACM on human-computer interaction3, CSCW (Nov. 2019), 1–33. doi:10.1145/3359246

  61. [61]

    Morgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, and Jed R Brubaker. 2020. How we’ve taught algorithms to see identity: Constructing race and gender in image databases for facial analysis.Proceedings of the ACM on human-computer interaction4, CSCW1 (May 2020), 1–35. doi:10.1145/3392866

  62. [62]

    Xiaoxiao Shang, Zhiyuan Peng, Qiming Yuan, Sabiq Khan, Lauren Xie, Yi Fang, and Subramaniam Vincent. 2022. DIANES: A DEI Audit Toolkit for News Sources. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22). Association for Computing Machinery, New York, NY, USA, 3312–3317. doi:10.114...

  63. [63]

    Nancy Signorielli and Aaron Bacue. 1999. Recognition and respect: A content analysis of prime-time television characters across three decades.Sex roles40, 7-8 (April 1999), 527–544. doi:10.1023/a:1018883912900

  64. [64]

    Vivek K Singh, Mary Chayko, Raj Inamdar, and Diana Floegel. 2020. Female librarians and male computer programmers? Gender bias in occupational images on digital media platforms.Journal of the Association for Information Science and Technology71, 11 (Nov. 2020), 1281–1294. doi:10.1002/asi.24335

  65. [65]

    Aleksandra Sorokovikova, Pavel Chizhov, Iuliia Eremenko, and Ivan P Yamshchikov. 2025. Surface fairness, deep bias: A comparative study of bias in language models. InProceedings of the 6th Workshop on Gender Bias in Natural Language Processing (GeBNLP). Association for Computational Linguistics, Stroudsburg, PA, USA, 206–227. doi:10.18653/v1/2025.gebnlp-1.20

  66. [66]

    C Steele and Joshua Aronson. 1995. Stereotype threat and the intellectual test performance of African Americans.Journal of personality and social psychology69, 5 (Nov. 1995), 797–811. doi:10.1037/0022-3514.69.5.797

  67. [67]

    Jane G Stout, Nilanjana Dasgupta, Matthew Hunsinger, and Melissa A McManus. 2011. STEMing the tide: using ingroup experts to inoculate women’s self-concept in science, technology, engineering, and mathematics (STEM).Journal of personality and social psychology100, 2 (Feb. 2011), 255–270. doi:10.1037/a0021385 This work is currently under peer review for pu...

  68. [68]

    Bureau of Labor Statistics

    U.S. Bureau of Labor Statistics. 2023. Current Population Survey: Annual Averages, 2023. https://www.bls.gov/cps/cps_aa2023.htm. Accessed: March 7, 2025

  69. [69]

    Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao ’kenneth Huang, and Shomir Wilson. 2023. Nationality Bias in Text Generation.arXiv [cs.CL](Feb. 2023). arXiv:2302.02463 [cs.CL] doi:10.48550/arXiv.2302.02463

  70. [70]

    James Vincent. 2025. We tried out DeepSeek. It works well — until we asked it about Tiananmen Square and Taiwan. https: //www.theguardian.com/technology/2025/jan/28/we-tried-out-deepseek-it-works-well-until-we-asked-it-about-tiananmen-square-and- taiwan?utm_source=chatgpt.com. Accessed: 2025-9-10

  71. [71]

    Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. 2024. Survey of bias in Text-to-Image generation: Definition, evaluation, and mitigation.arXiv [cs.CV](April 2024). arXiv:2404.01030 [cs.CV] doi:10.48550/arXiv.2404.01030

  72. [72]

    Ruta Yemane and Mariña Fernández-Reino. 2021. Latinos in the United States and in Spain: the impact of ethnic group stereotypes on labour market outcomes.Journal of ethnic and migration studies47, 6 (April 2021), 1240–1260. doi:10.1080/1369183x.2019.1622806

  73. [73]

    Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the effect of accuracy on trust in machine learning models. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA. doi:10.1145/3290605.3300509 Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 This work is curre...