Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

People adhere to biased AI hiring recommendations up to 90 percent of the time, a 528-person experiment finds.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

People largely adopt racially biased AI hiring recommendations, selecting the AI-favored group up to 90% of the time; prior IAT exposure may reduce stereotype-congruent choices by about 13%.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Solid large-N evidence that biased AI hiring recommendations shift human decisions, but the 'close adherence' claim overreads a non-commensurate probability comparison. the 3 major comments →

arxiv 2509.04404 v2 pith:JDKC6YBZ submitted 2025-09-04 cs.CY cs.AIcs.CLcs.HC

No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy

classification cs.CY cs.AIcs.CLcs.HC
keywords AI biashuman-in-the-loopresume screeninghiring discriminationimplicit association testhuman autonomyLLM recommendationshuman-AI decision making
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the human-in-the-loop model of AI hiring oversight actually works: when an AI system's recommendations favor one racial group, do the human reviewers catch and correct the bias? The answer, from a resume-screening experiment with 528 participants and 1,526 decisions across three racial comparisons (White vs. Black, White vs. Asian, White vs. Hispanic), is no. People who screened resumes without AI advice, or with racially neutral advice, chose White and non-White candidates at equal rates. But when the simulated AI favored a particular group, participants favored that group up to 90 percent of the time, and in most biased conditions their selection rates were statistically indistinguishable from the AI's recommendation rates. The authors conclude that biased AI recommendations propagate through human reviewers and that human-in-the-loop oversight, as currently practiced, does not protect against them.

Core claim

Biased AI hiring recommendations do not merely accompany human judgment—they override it. Across 16 occupations and three racial comparisons, participants chose the AI-favored group at the AI's own recommendation rate, up to 90 percent in extreme conditions. This held for stereotype-congruent and stereotype-incongruent biases, showing people do not resist recommendations that contradict common stereotypes. Completing a race-status implicit association test before the task raised selection of stereotype-incongruent candidates by about 13 percent. Participants' own IAT scores, explicit beliefs, hiring experience, and AI familiarity did not predict decisions; perceived AI quality and importance

What carries the argument

The simulated AI recommendation is the central mechanism, varied along two axes: direction (congruent or incongruent with US race-status stereotypes) and magnitude (moderate or severe). Moderate levels were calibrated by a retrieval simulation in which three text-embedding models ranked resumes by cosine similarity to job descriptions and kept the top 10 percent, following the authors' earlier measurement of real-world LLM resume-screening bias; severe levels recommended every candidate of one race and none of the other. A binomial logistic mixed model compares the probability of a participant choosing a majority-White slate with the rate at which the simulated AI recommended White candidate

Load-bearing premise

The experiment assumes the simulated AI recommendation rates approximate the bias of real deployed hiring systems: Moderate conditions are averages from three embedding models in a top-10-percent retrieval simulation, so if actual LLM screening tools rank, present, or bias candidates differently, the measured human adherence rates may not generalize.

What would settle it

Run the same screening protocol with a real LLM-based hiring tool whose racial recommendation rates are independently measured, instead of the simulated flags used here; if reviewers' choices diverge from the tool's recommendation rates—or if telling reviewers the AI may be biased eliminates the adherence—the propagation claim is bounded. A cheaper check: rerun the protocol without the four-minute time limit and see whether adherence depends on time pressure.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Human-in-the-loop oversight, as currently implemented, will not reliably neutralize biased AI hiring recommendations; the human reviewer tends to reproduce the AI's racial preferences rather than correct them.
  • Hiring outcomes hinge on the direction of AI bias: stereotype-congruent recommendations can amplify existing racial disparities in high-status jobs, while incongruent recommendations could reduce or reverse them, so the same system can either worsen or improve inequality depending on context.
  • Bias-awareness exercises in the style of IATs can measurably increase selection of stereotype-incongruent candidates (about 13 percent), offering a concrete design lever for AI-HITL systems.
  • Because even participants who rated AI recommendations as poor quality or unimportant shifted their behavior under some bias conditions, training that simply teaches skepticism of AI is insufficient; calibrating judgments of AI performance is the harder task.
  • Regulatory and organizational policy that treats human review as a built-in safeguard against AI bias in hiring rests on an unsafe assumption and should be revised to account for demonstrated bias propagation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The bias-propagation mechanism—an AI recommendation anchoring a time-pressured human choice—is not specific to resumes; if it generalizes, similar effects should appear in other high-stakes human-AI collaborations such as healthcare triage, credit decisions, or academic review, where the same anchoring conditions hold.
  • Because IAT scores themselves did not predict decisions but IAT order did, the intervention probably works through priming awareness of stereotypes rather than measuring a stable trait; a cheaper test would be whether a simple written reminder about race-status stereotypes produces the same 13 percent shift without running a full IAT.
  • The job-status asymmetry suggests a prioritization rule for mitigation: bias propagation should be strongest where the AI's preference aligns with a pre-existing cultural schema (White candidates for high-status work), so interventions and audits should target schema-congruent biases first.
  • The experiment's four-minute time limit was chosen to mimic real screening pressure; an untimed replication would reveal whether adherence to biased recommendations is a time-pressure phenomenon or a more general deference to machine advice.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a large-scale human experiment (N=528; 1,526 analyzed resume-screening scenarios) in which participants selected candidates for high- and low-status occupations while receiving simulated AI recommendations with varying levels of racial bias. The authors find that biased AI recommendations shift human choices relative to no-AI and neutral-AI baselines, and that the probability of a participant making a majority-White selection often approaches the AI's rate of recommending White candidates, especially in the most severe bias conditions. They also report that completing an IAT before the decision task is associated with a 13% increase in selecting stereotype-incongruent candidates, and that self-reported perceptions of AI recommendation quality and importance do not reliably protect against biased recommendations. The paper concludes that human-in-the-loop oversight, as currently practiced, does not reliably neutralize AI bias in resume screening.

Significance. If the main claim holds, the paper makes an important contribution to the human-AI decision-making and algorithmic fairness literatures: it provides the largest human-subjects evidence to date that biased AI recommendations can propagate into human hiring decisions, and it examines moderators (IATs, AI perceptions, prior experience) that have rarely been tested in AI-HITL contexts. The study is well-powered, uses quality-controlled stimuli, includes both realistic and counterfactual bias levels, and the authors share anonymized data and analysis code, which strengthens reproducibility. The directional finding—biased AI shifts decisions relative to no-AI and neutral baselines—is credible and important even if, as discussed below, the paper's stronger 'close adherence' claim rests on a non-commensurate comparison.

major comments (3)
  1. [Table 3 and Section 4.1] The claim that participants' preference rates 'did not significantly differ' from AI recommendation rates in most biased conditions compares two non-commensurate quantities. The human outcome modeled in Table 3 is the probability of a majority-White response (selecting two or more White candidates out of three), while the AI rate in Table 2 is the proportion of AI recommendations that are White. If a human selected White candidates at the same per-recommendation rate p as the AI, the expected majority-White probability would be f(p)=3p^2(1-p)+p^3, not p. For the High Congruent/Moderate condition (average AI rate p≈.737), the equality benchmark is f(p)≈.876, whereas the reported human probability is .750. The reported z-tests therefore do not test the paper's 'close adherence' hypothesis. The analyses should be re-run at the individual-candidate choice level (e.g., modeling each of the th
  2. [Section 4.1, Figure 4, and Abstract] The paper states that completing an IAT before the decision task increases selection of stereotype-incongruent candidates by 13%, and this is presented as a key contribution in the abstract. However, the immediately preceding sentence reports that no post-hoc pairwise comparisons for the Task Order × Job Status interaction were significant. The significant omnibus interaction alone does not justify the causal-sounding 'can increase by 13%' claim, especially without multiplicity-corrected pairwise tests or a direct test of the specific contrasts on which the 13% figure is based. The authors should either provide the relevant pairwise tests, report the result as a non-significant trend, or weaken the abstract and conclusions accordingly. This is a secondary claim relative to the main propagation result, but it is one of the stated main contributions and currently overstates the evidence.
  3. [Appendix F and Section 3.1] The simulated AI bias levels for the Moderate conditions are averages over three embedding models, multiple instruction paraphrases, and occupation groups using a top-10% cosine-similarity threshold. These values are reasonable as inputs, and the paper is transparent about their origin, but the phrase 'approximates factual ... estimates of racial bias in real-world AI systems' overstates their status. The comparison in Table 3 also treats these simulated rates as fixed constants without incorporating uncertainty from the retrieval simulation. At minimum, the authors should state more carefully in the main text that these are simulation-based estimates, not measurements from deployed hiring systems, and ideally provide a sensitivity analysis showing how the 'close adherence' conclusion changes across the range of bias magnitudes (e.g., the per-race values in Table 2 rather than only their
minor comments (5)
  1. [Table 3] There are apparent sign and formatting errors in the Δ AI Rec column. For Low Congruent/Severe, Prob=.138 and AI Rec=0.000, so the difference should be +.138, not −.138. For High Congruent/Severe, '−0.96**' appears to be a decimal-point typo for '−.096**'.
  2. [Figure 3 caption] Typo: 'Particpants' should be 'Participants'.
  3. [Appendix E, Figure 7 caption] Typo: 'recieved' should be 'received'.
  4. [Section 3.4 and Tables 7-8] The variable names in the model output (e.g., 'biasSim-Cong-New', 'biasExt-Cong', 'Group recode', 'I recode') are not defined in the main text. Please add a short legend or rename them for readability, since these terms are used in the appendix tables but not explained there either.
  5. [Throughout] The text contains several typographical artifacts, such as 'ANOV A' instead of 'ANOVA' (Section 3.4 and Appendix N.2) and inconsistent spacing in 'AI Recommendation' headings. A careful proofreading pass is needed.

Circularity Check

0 steps flagged

No significant circularity: the AI bias levels are fixed inputs from prior simulation, and the human outcomes are measured independently.

full rationale

The paper's claimed derivation is not circular. The AI recommendation rates in Table 2 are fixed before the human experiment: they are estimated in Appendix F by embedding-based retrieval simulations following Wilson and Caliskan (2024), and then used as experimental conditions (Moderate and Severe bias). No parameter of the human-response model (BLMM in Section 3.4) is fitted to produce these rates, and the human outcome data could have contradicted the propagation hypothesis (e.g., participants could ignore the recommendations and remain at baseline). The central claim that biased AI recommendations shift human decisions is therefore an empirical result about newly collected behavioral data, not an algebraic consequence of the inputs. The self-citation to Wilson and Caliskan (2024) supplies the simulation procedure and motivating evidence of LLM resume-screening bias, but it is not used as the proof of human bias propagation; multiple independent references also support the existence of AI hiring bias. The comparison in Table 3 between the human predicted probability of a majority-White response and the AI's recommendation rate may raise a statistical-validity concern about commensurability of quantities, but that is a methodological issue, not circularity, because the two quantities are not made equal by construction. Consequently, no circular step meets the evidentiary standard of the review.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

No free parameters are fitted to the human outcome data. The listed hand-chosen thresholds and groupings are inputs to the AI-bias simulation, not to the response model. The core experimental outcome is measured, and no new entities are postulated.

free parameters (2)
  • Top-10% retrieval threshold = 10%
    Hand-chosen cutoff in the AI resume-screening simulation (Appendix F) to define 'selected' resumes; directly determines the Moderate AI bias magnitudes in Table 2.
  • Occupation demographic grouping average = average of representative and deviating occupation group averages
    Occupations are hand-grouped into 'approximately US population demographics' versus 'deviating' categories (Appendix A, F.3); the Moderate bias values are computed separately within each group and then averaged for the main experiment.
axioms (4)
  • domain assumption Resume names and affinity-group memberships signal racial identity to participants as intended
    Section 3.1, Table 1. The experiment relies on participants inferring race from names and organizations; if the cues are ambiguous or unnoticed, the race manipulation fails.
  • domain assumption Simulated AI recommendation rates approximate real-world LLM resume-screening bias
    Section 3.1 and Appendix F. Bias magnitudes are estimated from retrieval simulations with three embedding models using a top-10% cutoff; the external validity of the human experiment depends on this approximation.
  • domain assumption The IAT scoring algorithm and status-race stimuli measure the intended implicit associations
    Section 3.4, Appendix G. The Task Order effect interpretation depends on the IAT actually triggering awareness of race-status stereotypes.
  • domain assumption Occupation status is adequately operationalized by average salary ($30k-$35k vs $110k-$135k)
    Section 3.1 and Appendix A. The congruent/incongruent framing of AI bias relies on these occupations being perceived as high or low status.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy." pith.science (2026). https://pith.science/paper/JDKC6YBZ

@misc{pith2026250904404,
  author       = {Pith},
  title        = {Pith review of: No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JDKC6YBZ}},
  note         = {Machine review of arXiv:2509.04404}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this study, we conduct a resume-screening experiment (N=528) where people collaborate with simulated AI models exhibiting race-based preferences (bias) to evaluate candidates for 16 high and low status occupations. Simulated AI bias approximates factual and counterfactual estimates of racial bias in real-world AI systems. We investigate people's preferences for White, Black, Hispanic, and Asian candidates (represented through names and affinity groups on quality-controlled resumes) across 1,526 scenarios and measure their unconscious associations between race and status using implicit association tests (IATs), which predict discriminatory hiring decisions but have not been investigated in human-AI collaboration. When making decisions without AI or with AI that exhibits no race-based preferences, people select all candidates at equal rates. However, when interacting with AI favoring a particular group, people also favor those candidates up to 90% of the time, indicating a significant behavioral shift. The likelihood of selecting candidates whose identities do not align with common race-status stereotypes can increase by 13% if people complete an IAT before conducting resume screening. Finally, even if people think AI recommendations are low quality or not important, their decisions are still vulnerable to AI bias under certain circumstances. This work has implications for people's autonomy in AI-HITL scenarios, AI and work, design and evaluation of AI hiring systems, and strategies for mitigating bias in collaborative decision-making tasks. In particular, organizational and regulatory policy should acknowledge the complex nature of AI-HITL decision making when implementing these systems, educating people who use them, and determining which are subject to oversight.

Figures

Figures reproduced from arXiv: 2509.04404 by Anna-Maria Gueorguieva, Aylin Caliskan, Kyra Wilson, Mattea Sim.

Figure 1
Figure 1. Figure 1: Predicted probability of preference for White can [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: An example of the interface 575 participants [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Predicted probability of participants preferring [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Predicted probability of participants preferring [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The difference in predicted probability of preferring White candidates between conditions with AI recommendations [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: After the fourth, eight, twelfth, and final validation [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: An example of the interface subjects saw when [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: The distribution of ratings given to resumes within [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The relationship between the three MTEs used [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Illustration of the resume screening as document retrieval framework. Task instructions are appended to job descrip [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Resumes with white names are significantly pre [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 12
Figure 12. Figure 12: Resumes with white names are significantly pre [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Pictures used to represent racial groups in white [PITH_FULL_IMAGE:figures/full_fig_p018_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Predicted probability of favoring white candidates in resume screening task split by response to AI recommendation [PITH_FULL_IMAGE:figures/full_fig_p019_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Predicted probability of favoring white candidates in resume screening task split by response to AI recommendation [PITH_FULL_IMAGE:figures/full_fig_p020_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: An example of the interface subjects saw when [PITH_FULL_IMAGE:figures/full_fig_p021_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Distribution of IAT scores for each Race condi￾tion. Positive values indicate associations between white and high status; negative values indicate associations between non-white and high status. -4 0 4 8 12 White vs. Black White vs. Asian White vs. Hispanic Race explicit_scale [PITH_FULL_IMAGE:figures/full_fig_p022_17.png] view at source ↗
Figure 20
Figure 20. Figure 20: Number of participant responses by each answer [PITH_FULL_IMAGE:figures/full_fig_p022_20.png] view at source ↗
Figure 23
Figure 23. Figure 23: The strength of association between the categori [PITH_FULL_IMAGE:figures/full_fig_p023_23.png] view at source ↗
Figure 22
Figure 22. Figure 22: Number of participant responses by each answer [PITH_FULL_IMAGE:figures/full_fig_p023_22.png] view at source ↗
Figure 24
Figure 24. Figure 24: The importance of each variable as a percentage [PITH_FULL_IMAGE:figures/full_fig_p024_24.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Resume-ing Control: (Mis)Perceptions of Agency Around GenAI Use in Recruiting Workflows

    cs.CY 2026-04 unverdicted novelty 5.0

    Recruiters perceive themselves as retaining agency over GenAI in hiring pipelines, yet GenAI invisibly architects core evaluation inputs, producing only marginal efficiency gains at the cost of deskilling.

Reference graph

Works this paper leans on

81 extracted references · 78 canonical work pages · cited by 1 Pith paper

  1. [1]

    Agerstr \"o m, J.; and Rooth, D.-O. 2011. The role of automatic obesity stereotypes in real hiring discrimination. Journal of Applied Psychology, 96(4): 790

  2. [2]

    Agresti, A.; and Tarantola, C. 2018. Simple ways to interpret effects in modeling ordinal categorical data. Statistica Neerlandica, 72(3): 210--223

  3. [3]

    J.; and van den Hoven, J

    Aizenberg, E.; Dennis, M. J.; and van den Hoven, J. 2025. Examining the assumptions of AI hiring assessments and their impact on job seekers’ autonomy over self-representation. AI & society, 40(2): 919--927

  4. [4]

    Armstrong, L.; Liu, A.; MacNeil, S.; and Metaxa, D. 2024. The Silicon Ceiling: Auditing GPT’s Race and Gender Biases in Hiring. In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, 1--18

  5. [5]

    Bertrand, M.; and Mullainathan, S. 2004. Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American economic review, 94(4): 991--1013

  6. [6]

    Binns, R. 2022. Human Judgment in algorithmic loops: Individual justice and automated decision-making. Regulation & governance, 16(1): 197--211

  7. [7]

    Brambilla, M.; Sacchi, S.; Castellini, F.; and Riva, P. 2010. The effects of status on perceived warmth and competence. Social Psychology

  8. [8]

    Bursell, M.; and Roumbanis, L. 2024. After the algorithms: A study of meta-algorithmic judgments and diversity in the hiring process at a large multisite company. Big Data & Society, 11(1): 20539517231221758

  9. [9]

    J.; and Narayanan, A

    Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334): 183--186

  10. [10]

    Cao, S.; and Huang, C.-M. 2022. Understanding user reliance on AI in assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2): 1--23

  11. [11]

    P.; Pogacar, R.; Pullig, C.; Kouril, M.; Aguilar, S.; LaBouff, J.; Isenberg, N.; and Chakroff, A

    Carpenter, T. P.; Pogacar, R.; Pullig, C.; Kouril, M.; Aguilar, S.; LaBouff, J.; Isenberg, N.; and Chakroff, A. 2019. Survey-software implicit association tests: A methodological and empirical analysis. Behavior research methods, 51: 2194--2208

  12. [12]

    Chan, E. 2024. 2024 hiring trends survey: What makes a great job candidate?

  13. [13]

    E.; and Banaji, M

    Charlesworth, T. E.; and Banaji, M. R. 2022. Patterns of implicit and explicit attitudes: IV. Change and stability from 2007 to 2020. Psychological Science, 33(9): 1347--1371

  14. [14]

    V.; Wortman Vaughan, J.; and Bansal, G

    Chen, V.; Liao, Q. V.; Wortman Vaughan, J.; and Bansal, G. 2023. Understanding the role of human intuition on reliance in human-AI decision-making with explanations. Proceedings of the ACM on Human-computer Interaction, 7(CSCW2): 1--32

  15. [15]

    Cheng, M.; Durmus, E.; and Jurafsky, D. 2023. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1504--1532

  16. [16]

    Cohen, J. 2016. A power primer. Quantitative Methods in Psychology

  17. [17]

    Dastin, J. 2018. Insight - Amazon scraps secret AI recruiting tool that showed bias against women. https://www.reuters.com/article/idUSKCN1MK0AG/. [Accessed 28-04-2024]

  18. [18]

    Diaz, I.; Hubbard, A.; Decker, A.; and Cohen, M. 2015. Variable importance and prediction methods for longitudinal problems with missing variables. PloS one, 10(3): e0120031

  19. [19]

    M.; and Hayes, M

    Elder, E. M.; and Hayes, M. 2023. Signaling race, ethnicity, and gender with names: Challenges and recommendations. The Journal of Politics, 85(2): 764--770

  20. [20]

    EU AI Act. 2024. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artific...

  21. [21]

    J.; Graus, D.; Hacker, P.; Saldivar, J.; Zuiderveen Borgesius, F.; and Biega, A

    Fabris, A.; Baranowska, N.; Dennis, M. J.; Graus, D.; Hacker, P.; Saldivar, J.; Zuiderveen Borgesius, F.; and Biega, A. J. 2025. Fairness and bias in algorithmic hiring: A multidisciplinary survey. ACM Transactions on Intelligent Systems and Technology, 16(1): 1--54

  22. [22]

    T.; Cuddy, A

    Fiske, S. T.; Cuddy, A. J.; Glick, P.; and Xu, J. 2018. A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. In Social cognition, 162--214. Routledge

  23. [23]

    Fourrier, C.; Habib, N.; Lozovskaya, A.; Szafer, K.; and Wolf, T. 2024. Open LLM Leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard

  24. [24]

    Gambarota, F.; and Alto \`e , G. 2024. Ordinal regression models made easy: A tutorial on parameter interpretation, data simulation and power analysis. International Journal of Psychology, 59(6): 1263--1292

  25. [25]

    Gautam, V.; Subramonian, A.; Lauscher, A.; and Keyes, O. 2024. Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP. In The 5th Workshop on Gender Bias in Natural Language Processing, 323

  26. [26]

    Glazko, K.; Mohammed, Y.; Kosa, B.; Potluri, V.; and Mankoff, J. 2024. Identifying and improving disability bias in GPT-based resume screening. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 687--700

  27. [27]

    F.; Liu, W.; Shirase, L.; Tomczak, D

    Gonzalez, M. F.; Liu, W.; Shirase, L.; Tomczak, D. L.; Lobbe, C. E.; Justenhoven, R.; and Martin, N. R. 2022. Allying with AI? Reactions toward human-based, AI/ML-based, and augmented hiring processes. Computers in Human Behavior, 130: 107179

  28. [28]

    G.; McGhee, D

    Greenwald, A. G.; McGhee, D. E.; and Schwartz, J. L. 1998. Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6): 1464

  29. [29]

    G.; Nosek, B

    Greenwald, A. G.; Nosek, B. A.; and Banaji, M. R. 2003. Understanding and using the implicit association test: I. An improved scoring algorithm. Journal of personality and social psychology, 85(2): 197

  30. [30]

    Heinze, G.; Wallisch, C.; and Dunkler, D. 2018. Variable selection--a review and recommendations for the practicing statistician. Biometrical journal, 60(3): 431--449

  31. [31]

    Helwig, N. E. 2025. Versatile descent algorithms for group regularization and variable selection in generalized linear models. Journal of Computational and Graphical Statistics, 34(1): 239--252

  32. [32]

    HireVue. 2017. Unilever Finds Top Talent Faster With Hirevue Assessments

  33. [33]

    Hofmann, W.; Gawronski, B.; Gschwendner, T.; Le, H.; and Schmitt, M. 2005. A meta-analysis on the correlation between the Implicit Association Test and explicit self-report measures. Personality and social psychology bulletin, 31(10): 1369--1385

  34. [34]

    Holm, S. 1979. A simple sequentially rejective multiple test procedure. Scandinavian journal of statistics, 65--70

  35. [35]

    Huang, Y.; Tibbe, T.; Tang, A.; and Montoya, A. 2023. Lasso and group lasso with categorical predictors: Impact of coding strategy on variable selection and prediction. Journal of Behavioral Data Science, 3(2): 15--42

  36. [36]

    Jakesch, M.; Bhat, A.; Buschek, D.; Zalmanson, L.; and Naaman, M. 2023. Co-writing with opinionated language models affects users’ views. In Proceedings of the 2023 CHI conference on human factors in computing systems, 1--15

  37. [37]

    Kahneman, D. 2011. Thinking, fast and slow. macmillan

  38. [38]

    Kim, D.; Vegt, N.; Visch, V.; and Bos-De Vos, M. 2024. How Much Decision Power Should (A) I Have?: Investigating Patients’ Preferences Towards AI Autonomy in Healthcare Decision Making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1--17

  39. [39]

    u per, A.; and Kr \

    K \"u per, A.; and Kr \"a mer, N. 2025. Psychological traits and appropriate reliance: Factors shaping trust in AI. International Journal of Human--Computer Interaction, 41(7): 4115--4131

  40. [40]

    Lacroux, A.; and Martin-Lacroux, C. 2022. Should I trust the artificial intelligence to recruit? Recruiters’ perceptions and behavior when faced with algorithm-based recommendation systems during resume screening. Frontiers in Psychology, 13: 895997

  41. [41]

    H.; Sarkar, A.; Tankelevitch, L.; Drosos, I.; Rintel, S.; Banks, R.; and Wilson, N

    Lee, H.-P. H.; Sarkar, A.; Tankelevitch, L.; Drosos, I.; Rintel, S.; Banks, R.; and Wilson, N. 2025. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers

  42. [42]

    H.; and Chew, C

    Lee, M. H.; and Chew, C. J. 2023. Understanding the effect of counterfactual explanations on trust and reliance on ai for human-ai collaborative clinical decision making. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2): 1--22

  43. [43]

    Li, L.; Lassiter, T.; Oh, J.; and Lee, M. K. 2021. Algorithmic hiring in practice: Recruiter and HR Professional's perspectives on AI use in hiring. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 166--176

  44. [44]

    Matuschek, H.; Kliegl, R.; Vasishth, S.; Baayen, H.; and Bates, D. 2017. Balancing Type I error and power in linear mixed models. Journal of memory and language, 94: 305--315

  45. [45]

    McNeish, D. M. 2015. Using lasso for predictor selection and to assuage overfitting: A method long overlooked in behavioral sciences. Multivariate behavioral research, 50(5): 471--484

  46. [46]

    Melamed, D.; Barry, L.; Montgomery, B.; and Okuwobi, O. F. 2020. Measuring racial status beliefs with implicit associations. American sociological review, 85(6): 1123--1131

  47. [47]

    W.; Barry, L.; Montgomery, B.; and Okuwobi, O

    Melamed, D.; Munn, C. W.; Barry, L.; Montgomery, B.; and Okuwobi, O. F. 2019. Status characteristics, implicit bias, and the production of racial inequality. American Sociological Review, 84(6): 1013--1036

  48. [48]

    Meng, R.; Liu, Y.; Rayhan Joty, S.; Xiong, C.; Zhou, Y.; and Yavuz, S. 2024. SFR-Embedding-Mistral:Enhance Text Retrieval with Transfer Learning. Salesforce AI Research Blog

  49. [49]

    Montgomery, B.; Park, H.; Barry Burrill, L.; and Melamed, D. 2024. Measuring gender status beliefs. Socius, 10: 23780231241245845

  50. [50]

    Muennighoff, N.; Su, H.; Wang, L.; Yang, N.; Wei, F.; Yu, T.; Singh, A.; and Kiela, D. 2024. Generative representational instruction tuning. arXiv preprint arXiv:2402.09906

  51. [51]

    Peng, A.; Nushi, B.; Kiciman, E.; Inkpen, K.; and Kamar, E. 2022. Investigations of performance and bias in human-AI teamwork in hiring. In Proceedings of the AAAI conference on artificial intelligence, volume 36, 12089--12097

  52. [52]

    Prunkl, C. 2024. Human autonomy at risk? An analysis of the challenges from AI. Minds and Machines, 34(3): 26

  53. [53]

    Quillian, L.; and Lee, J. J. 2023. Trends in racial and ethnic discrimination in hiring in six Western countries. Proceedings of the National Academy of Sciences, 120(6): e2212875120

  54. [54]

    Raghavan, M.; Barocas, S.; Kleinberg, J.; and Levy, K. 2020. Mitigating bias in algorithmic hiring: Evaluating claims and practices. In Proceedings of the 2020 conference on fairness, accountability, and transparency, 469--481

  55. [55]

    Resume Builder 2024. 2024. 7 in 10 Companies Will Use AI in the Hiring Process in 2025, Despite Most Saying It’s Biased

  56. [56]

    Reuben, E.; Sapienza, P.; and Zingales, L. 2014. How stereotypes impair women’s careers in science. Proceedings of the National Academy of Sciences, 111(12): 4403--4408

  57. [57]

    Rosenthal, J. A. 1996. Qualitative descriptors of strength of association and effect size. Journal of social service Research, 21(4): 37--59

  58. [58]

    M.; and Sach, A

    Rosenthal-von der P \"u tten, A. M.; and Sach, A. 2024. Michael is better than Mehmet: exploring the perils of algorithmic biases and selective adherence to advice from automated decision support systems in hiring. Frontiers in Psychology, 15: 1416504

  59. [59]

    A.; and Ashmore, R

    Rudman, L. A.; and Ashmore, R. D. 2007. Discrimination and the implicit association test. Group Processes & Intergroup Relations, 10(3): 359--372

  60. [60]

    S \'a nchez-Monedero, J.; Dencik, L.; and Edwards, L. 2020. What does it mean to'solve'the problem of discrimination in hiring? Social, technical and legal perspectives from the UK on automated hiring systems. In Proceedings of the 2020 conference on fairness, accountability, and transparency, 458--468

  61. [61]

    Schoeffer, J.; De-Arteaga, M.; and Kuehl, N. 2024. Explanations, fairness, and appropriate reliance in human-AI decision-making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1--18

  62. [62]

    K.; and Zonderman, A

    Shaked, D.; Williams, M.; Evans, M. K.; and Zonderman, A. B. 2016. Indicators of subjective social status: Differential associations across race and sex. SSM-population health, 2: 700--707

  63. [63]

    M.; and Watts, D

    Sharma, A.; Hofman, J. M.; and Watts, D. J. 2015. Estimating the causal impact of recommendation systems from observational data. In Proceedings of the Sixteenth ACM Conference on Economics and Computation, 453--470

  64. [64]

    Shtaynberger, J.; and Bar, H. 2023. Equivalence Testing

  65. [65]

    Solsman, J. E. 2018. YouTube's AI is the puppet master over most of what you watch. CNET

  66. [66]

    Spatola, N. 2024. The efficiency-accountability tradeoff in AI integration: Effects on human performance and over-reliance. Computers in Human Behavior: Artificial Humans, 2(2): 100099

  67. [67]

    Tabassi, E. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0)

  68. [68]

    Tonidandel, S.; and LeBreton, J. M. 2011. Relative importance analysis: A useful supplement to regression analysis. Journal of Business and Psychology, 26: 1--9

  69. [69]

    T.; Hooker, G.; Ellner, S

    Tredennick, A. T.; Hooker, G.; Ellner, S. P.; and Adler, P. B. 2021. A practical guide to selecting models for exploration, inference, and prediction in ecology. Ecology, 102(6): e03336

  70. [70]

    Valentino, L. 2022. Constructing the racial hierarchy of labor: the role of race in occupational prestige judgments. Sociological Inquiry, 92(2): 647--673

  71. [71]

    S.; and Krishna, R

    Vasconcelos, H.; J \"o rke, M.; Grunde-McLaughlin, M.; Gerstenberg, T.; Bernstein, M. S.; and Krishna, R. 2023. Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1): 1--38

  72. [72]

    Wang, L.; Yang, N.; Huang, X.; Yang, L.; Majumder, R.; and Wei, F. 2023. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368

  73. [73]

    Weber, L. 2024. New York City Passed an AI Hiring Law. So Far, Few Companies Are Following It. The Wall Street Journal

  74. [74]

    Weerts, H.; Kelly-Lyth, A.; Binns, R.; and Adams-Prassl, J. 2024. Unlawful Proxy Discrimination: A Framework for Challenging Inherently Discriminatory Algorithms. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 1850--1860

  75. [75]

    Wilkens, U.; Lutzeyer, I.; Zheng, C.; Beser, A.; and Prilla, M. 2025. Augmenting diversity in hiring decisions with artificial intelligence tools. The International Journal of Human Resource Management, 1--38

  76. [76]

    Williamson, S.; and Foley, M. 2018. Unconscious bias training: The ‘silver bullet’for gender equity? Australian Journal of Public Administration, 77(3): 355--359

  77. [77]

    Wilson, K.; and Caliskan, A. 2024. Gender, race, and intersectional bias in resume screening via language model retrieval. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 1578--1590

  78. [78]

    Wilson, K.; and Caliskan, A. 2025. Gender, race, and intersectional bias in AI resume screening via language model retrieval. Brookings

  79. [79]

    Yang, D.; Hovy, D.; Jurgens, D.; and Plank, B. 2025. Socially Aware Language Technologies: Perspectives and Practices. Computational Linguistics, 1--15

  80. [80]

    J.; Chahal, R.; Gotlib, I

    Zhou, D. J.; Chahal, R.; Gotlib, I. H.; and Liu, S. 2024. Comparison of lasso and stepwise regression in psychological data. Methodology, 20(2): 121--143

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.