{"id":"840f0216-80ed-4481-9b39-8480e67481ed","arxiv_id":"2606.28978","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Audit of 14 LLMs finds generational reversal from pro-White (+2.12 pp) to null or pro-Black (up to -3.01 pp) bias in resume callbacks using paired methodology.","lead":"This paper audits 14 LLMs with paired resumes and finds only the 2023 model shows pro-White bias while 2024+ models show null or pro-Black bias. Smart generalists should read it to see how AI fairness properties appear to have shifted across recent model generations.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Paired-resume callback gap may not measure the same construct when extracted from LLM prompt responses rather than employer decisions","rationale":"The reader's weakest_assumption is precisely the load-bearing assumption required to interpret LLM output differences as evidence of hiring discrimination comparable to the cited field experiments. Because the review was abstract-only, the full methods section would be needed to assess whether the authors performed any robustness checks on prompt wording or human-LLM correspondence; absent those, the concern stands and the UNVERDICTED status is appropriate.","tokens_in":1639,"tokens_out":426,"duration_ms":27052,"concrete_test":"Take the exact resume pairs and names from the paper, re-prompt each of the 14 models with three prompt variants: (a) the paper's original screening prompt, (b) an explicit \"act as a hiring manager and decide callbacks\" prompt with no fairness language, and (c) the same as (b) but with added \"ignore names and focus only on qualifications.\" Recompute the callback gaps; if the 2023-vs-2024 pattern disappears or reverses under (b) or (c), the headline generational claim is sensitive to prompt framing and the link to field-experiment discrimination is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim treats the difference in LLM output (e.g., interview recommendation or rating) across resume pairs that differ only in name as directly comparable to the callback gap measured in Kline, Rose & Walters field experiments. For the generational reversal result to support statements about \"algorithmic hiring bias,\" this equivalence must hold. It is least secure because (1) LLM outputs are generated from an explicit prompt whose wording, system instructions, and refusal/alignment training can alter measured gaps independently of any underlying preference, and (2) the original methodology relies on real decision-makers who incur costs and face legal constraints; an LLM does not. The abstract provides no evidence that the authors verified the mapping (e.g., via prompt ablation or comparison to human raters on identical pairs).","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper audits 14 mainstream LLMs for hiring discrimination on race and gender using the paired-resume callback methodology of Kline, Rose, and Walters (2022). It reports that the single 2023-vintage model reproduces the pro-White callback gap observed in human field experiments (+2.12 pp, p<0.01), while all 2024+ models exhibit either a null gap or a significant pro-Black reversal (up to -3.01 pp); an analogous generational reversal appears on the gender axis. Results rest on 24,024 paired evaluations per model.","tokens_in":1801,"tokens_out":608,"duration_ms":14725,"significance":"If the mapping from LLM output differences to the field-experiment callback gap is valid, the documented reversal would indicate that post-2023 alignment and safety training have materially altered the direction of demographic bias in LLM-based resume screening. This would be a substantive empirical contribution to the literature on algorithmic fairness and the effects of RLHF-style training, with direct implications for deployment of LLMs in employment contexts.","major_comments":[{"comment":"Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings.","section":"Methods"},{"comment":"§3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct.","section":"Results / Discussion"},{"comment":"Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly.","section":"Results"}],"minor_comments":[{"comment":"Abstract and §1: the phrase “algorithmic hiring bias” is used without distinguishing between bias in the LLM output distribution and bias in downstream hiring decisions; a brief clarifying sentence would improve precision.","section":"Abstract / Introduction"},{"comment":"The total of 24,024 paired postings per model is stated but the exact breakdown (number of occupations, name pairs, resume templates) is not summarized in a table; adding such a table would aid reproducibility.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important aspects of transparency and robustness. We address each major comment below and indicate planned revisions.","responses":[{"response":"We agree that these details are essential for reproducibility. In the revised manuscript we will add a new subsection to the Methods that provides the complete prompt templates, system instructions, temperature settings (set to 0 wherever supported for determinism), and the exact protocol used to extract callback decisions and handle any refusals or non-callback outputs.","revision_made":"yes","referee_comment":"[Methods] Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings."},{"response":"The paper applies the identical paired-resume callback definition and name-swapping procedure as Kline, Rose, and Walters (2022), so the measured gap is the same operational quantity. The contribution is to show how this quantity changes across LLM generations. We will expand the discussion to clarify the mapping, note that LLM outputs are not employer decisions, and address limitations including cost and legal constraints on deployment; however, new human-rater calibration or ablations are outside the scope of the current study design.","revision_made":"partial","referee_comment":"[Results / Discussion] §3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct."},{"response":"We agree that disaggregated results would strengthen the presentation. The revised manuscript will include supplementary tables reporting callback gaps broken down by resume template, occupation category, and name set for each model, confirming that the generational reversal pattern is not confined to particular subsets.","revision_made":"yes","referee_comment":"[Results] Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly."}],"tokens_in":1368,"tokens_out":547,"duration_ms":24683,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The central observation is straightforward: using the Kline-Rose-Walters paired-resume setup on 14 LLMs, the single 2023 model reproduces the pro-White callback gap seen in field data, while every later model shows either no gap or a reversal favoring Black resumes (up to 3 pp). The gender results follow the same generational split. They ran this at scale with 24k pairs per model and report the usual significance numbers.\n\nWhat stands out as new is simply documenting that shift across release years with a fixed method. The execution looks clean on the numbers side, and they cite the source methodology without overclaiming.\n\nThe soft spot is the leap from LLM output differences to statements about algorithmic hiring bias. Real callback gaps come from decision-makers who bear costs and legal exposure; LLM responses come from prompted text generation shaped by alignment and system instructions. Without prompt ablations, refusal checks, or side-by-side human comparisons on the same pairs, it is unclear how much of the reversal is an artifact of how the models were trained rather than a stable preference. The abstract gives effect sizes but the full methods section would need to show those controls to make the generational claim robust.\n\nThis is useful for groups studying post-training effects on model behavior or fairness audits in employment tools. It is worth sending to referees because the empirical pattern is easy to replicate and the topic matters for deployment, even if the interpretation will require more work on the mapping to actual decisions.","headline":"The paper shows a reversal in LLM resume bias from pro-White in the 2023 model to null or pro-Black in 2024+ ones via paired resumes, but the direct link to real hiring discrimination remains unproven.","tokens_in":2261,"tokens_out":389,"would_cite":false,"duration_ms":16256,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Only the 2023 LLM shows pro-White hiring bias matching human experiments, while 2024 and later models show null or pro-Black gaps.","keywords":["large language models","hiring discrimination","racial bias","gender bias","resume screening","algorithmic bias","paired audit"],"falsifier":"Repeating the exact paired-resume tests on the same 14 models but with a fresh set of job postings or name variants and finding no generational difference in the direction of bias.","tokens_in":2550,"feed_emoji":"","tokens_out":725,"duration_ms":33273,"temperature":0.7,"pith_summary":"This paper audits fourteen mainstream LLMs for hiring discrimination by applying the paired-resume method, feeding each model thousands of resume pairs that differ only in names signaling race or gender. The single 2023-vintage model produces a statistically significant pro-White callback gap of 2.12 percentage points, replicating the direction and rough size of gaps found in real-world employer experiments. Every model released in 2024 or after instead shows either no detectable gap or a significant pro-Black reversal reaching 3.01 percentage points, with the identical generational pattern appearing on the gender axis. The tests cover 24,024 paired evaluations per model. A sympathetic reader would care because LLMs are increasingly used to screen job applications, so any shift in their bias direction would directly affect who receives callbacks.","feed_headline":"Newer LLMs drop pro-White bias in resume screening","feed_subtitle":"2023 model matches human experiments with +2.12 pp White preference; 2024+ show null or Black preference","key_machinery":"Paired-resume methodology applied to LLM outputs, comparing callback rates on resumes identical except for names that signal race or gender.","core_discovery":"The paper establishes that the direction of bias in LLM resume screening reversed across model generations. The 2023 model reproduces the pro-White callback gap of +2.12 pp documented in field experiments on labor market discrimination, significant at the 1% level. Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal of up to -3.01 pp. The same pattern holds on the gender axis. These results are based on 24,024 paired postings per model across the 14 models tested.","pith_inferences":["If the reversal stems from post-2023 alignment techniques, those techniques may have over-corrected on demographic cues.","The pattern suggests that bias in LLM decision systems may require repeated audits with each new model release rather than a one-time check.","The findings could connect to broader questions about how training data updates affect fairness in other high-stakes LLM applications."],"forward_implications":["Newer LLMs could produce different or reversed demographic hiring outcomes when used for automated screening.","The bias reversal coincides with model release year, indicating that training or alignment changes alter how demographic signals are processed.","The generational shift appears on both the racial and gender axes.","Large-scale testing across 24,024 pairs per model supports the claim of a consistent reversal after 2023."],"fun_headline_variants":["LLM bias flips from pro-White to pro-Black after 2023","2023 models copy human bias, 2024+ reverse it","Resume screening bias reverses in recent LLMs","Hiring discrimination direction changes in LLM generations","Post-2023 LLMs eliminate pro-White resume bias"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The paired resumes that differ only in the racial or gender signal from names produce a valid measure of the LLM's hiring discrimination in the same way they measure human discrimination.","fun_headline_variants_meta":{"raw":{"variants":["LLM bias flips from pro-White to pro-Black after 2023","2023 models copy human bias, 2024+ reverse it","Resume screening bias reverses in recent LLMs","Hiring discrimination direction changes in LLM generations","Post-2023 LLMs eliminate pro-White resume bias"]},"model":"grok-4.3","cost_usd":0.012419,"raw_usage":{"total_tokens":5374,"prompt_tokens":598,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":124187000,"prompt_tokens_details":{"text_tokens":598,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4709,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":598,"tokens_out":67,"duration_ms":43363,"temperature":1.0,"reasoning_tokens":4709,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T09:45:16.818034+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the exact paired-resume tests on the same 14 models but with a fresh set of job postings or name variants and finding no generational difference in the direction of bias.","supporting_citations":[],"review_version":1}