{"id":"b625cec1-5f71-4230-bada-1e67cfa61031","arxiv_id":"2506.11361","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across ten commercial LLMs, demographic groups other than white male middle-aged were rated as more likely to help, while the control 'person' condition aligned with that majority baseline.","lead":"This paper tests ten commercial large language models for demographic bias by asking them to rate how likely a person described by race, gender, or age is to help others. The key finding: models seem to treat white, male, middle-aged adults as the default 'person', while rating other groups as more helpful.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'baseline demographic' claim rests on interpreting non-significant t-tests as evidence of equivalence; a proper equivalence analysis is required before this claim can be accepted.","rationale":"The reader's weakest-assumption analysis correctly identifies the inference from non-significance to equivalence as the critical flaw. The paper's two headline findings are (1) models have a baseline demographic of white middle-aged/young adult male, and (2) non-baseline groups are rated as more helpful. Finding (2) is supported by many statistically significant positive bias scores relative to control. Finding (1), however, is inferred from the non-significant differences for white, male, and middle-aged/young-adult groups. This is logically fragile: non-significance in a frequentist t-test only indicates that the observed data are not unlikely under the null; it does not quantify evidence for the null. Without an equivalence margin, a power analysis, or a Bayesian factor, the paper cannot distinguish 'the model treats the control as this demographic' from 'the model does not differentiate, or the test is underpowered.' The paper also does not address the multiple-comparison problem for the many t-tests run; while correcting for multiplicity would make the null results more frequent, it would not repair the equivalence inference—it would only increase the chance of false non-significance. The proposed concrete test (TOST with a prespecified margin) directly settles whether the data support equivalence for the key baseline comparisons. If the equivalence tests fail, the central baseline claim should be weakened substantially, and the paper should be revised to present finding (1) as a tentative observation rather than an established result. The current conditional verdict is appropriate: the paper's methodology is reproducible and interesting, but the headline requires additional statistical support.","tokens_in":10425,"tokens_out":4328,"duration_ms":47905,"concrete_test":"Reanalyze the paired comparisons for White vs control, Male vs control, Middle-aged vs control, and Young adult vs control using two one-sided tests (TOST) with a pre-registered equivalence margin of ±1 point on the 1-100 helpfulness scale (or ±2 points as a robustness check). For each model, compute the 90% confidence interval for the mean difference. If the CI for a comparison lies entirely within the margin, that comparison supports equivalence; if it straddles or excludes the margin, the 'baseline demographic' claim for that group is unsupported. Report the proportion of models for which each 'default' comparison passes the equivalence test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that models treat a white middle-aged or young adult male as the default 'person' is inferred from patterns of non-significant paired t-tests in Sections 5.1.1–5.1.3. For example, the text in 5.1.1 says the marked distinction between men and women/non-binary relative to control 'would imply that the models are treating the control group as male,' and 5.1.2 says non-significance for middle-aged and young adults 'may imply that models consider the control to be a person in this age range.' This is an inference from absence of evidence, which is not valid without a pre-specified equivalence margin and adequate statistical power. Non-significance can arise from small true effects, high variance, or small sample sizes (here, 103 scenario-pairs, but with only one run per model for most models). The claim is load-bearing because the paper's novel contribution is precisely identifying the baseline demographic; if the baseline inference fails, the headline conclusion is not established. The 'young adult' part is especially fragile: Table 1 shows several models with positive young-adult differences (e.g., 2.54, 1.78) that are reported as significant for some models, yet the text groups young adults with middle-aged as the 'default.' A formal equivalence test or Bayesian analysis is needed to distinguish 'no evidence of difference' from 'evidence of no difference.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces a control-condition method for measuring demographic bias in LLMs. The authors prompt ten commercial models with 103 human-generated Good Samaritan scenarios, varying the moral patient's race, gender, or age, and compare the resulting 1--100 helpfulness ratings against an unspecified 'person' control. Bias scores are computed as group mean minus control mean, with paired-sample t-tests at alpha = 0.01. The paper reports two headline findings: that models treat a white middle-aged or young adult male as the baseline 'person,' and that non-baseline demographics are generally rated as more willing to help than the baseline. The manuscript also reports brittleness across prompt rephrasings and a two-run repeatability check for two models.","tokens_in":10770,"tokens_out":3826,"duration_ms":43555,"significance":"The control-anchored design is a genuine methodological contribution: unlike multiple-choice stereotype tests, it allows directional bias and 'default person' assumptions to be separated within the same experiment, without fitting any parameters to the data. The use of human-generated prompts and the inclusion of a non-demographic control are strengths, as is the explicit repeatability check for two models. However, the central baseline-demographic claim rests on interpreting non-significant differences as evidence of equivalence, which is not statistically justified. If the authors add proper equivalence analyses and multiple-comparison correction, this could become a useful and reproducible LLM-auditing tool; in its current form, the load-bearing baseline conclusion is not established.","major_comments":[{"comment":"The inference that non-significant paired t-tests imply the model 'treats the control group as male' or 'considers the control to be a person in this age range' is not valid. A p-value above 0.01 only means the test did not detect a difference; it is not evidence that the difference is zero or practically equivalent. This is load-bearing because the baseline-demographic claim is the paper's main novelty. Please replace these inferences with a formal equivalence test (e.g., two one-sided tests) using a pre-specified margin, or with Bayesian interval estimation, and report the corresponding confidence intervals and power analysis for the 103 scenario-pairs. Without this, the headline finding that the baseline demographic is a white middle-aged or young adult male is not supported.","section":"Section 5.1.1, 5.1.2, and Section 7"},{"comment":"No correction is applied for multiple comparisons, despite testing 10 models across 3 demographic categories and up to 6 groups per category. At alpha = 0.01, dozens of tests are performed, so the claim that 'all models across all manufacturers displayed bias in at least 6 categories' likely overstates the evidence. Please report the total number of tests per category, apply a false-discovery-rate or family-wise error correction, and indicate which highlighted cells in Tables 1--3 survive correction. The within-model, within-category structure should also be taken into account, since demographic groups share the same control scores.","section":"Section 4.5 and Tables 1--3"},{"comment":"Equation (1) states that b, the number of brittleness rephrasings, is 4 in this paper, but Table 4 shows that most demographic groups have only 2 phrasings and some have 3 (e.g., non-binary has 3; senior has 3; teenager has 1 in the table as printed). This discrepancy affects how the aggregated scores and the paired t-tests are computed, and it is not a purely typographical issue because the standard error of the mean depends on the number of rephrasings per scenario. Please correct the definition, report the actual b for each group, and explain how missing phrasings were handled in the averaging.","section":"Section 4.6, Eq. (1), and Appendix A.1"},{"comment":"The text contains an internal contradiction. It states both that 'most models displayed statistically significant positive bias towards the young adult group' and that 'the majority of models found no significant difference between middle-aged and young adults relative to the control.' Yet Table 1 shows positive young-adult differences in 9 of 10 models, several of them sizable (e.g., 2.69, 2.54, 1.78). Since young adults are also included in the 'baseline demographic' claim, the paper must clarify exactly which models show significant young-adult effects and reconcile this with the assertion that young adults are part of the default category.","section":"Section 5.1.2 and Table 1"}],"minor_comments":[{"comment":"The citation placeholder '[llm-prompt-bias]' appears in the text but is not present in the reference list; please add the corresponding reference or remove the placeholder.","section":"Section 4.2"},{"comment":"The notation in Eq. (1) should clarify whether the sum over n uses a constant b for all scenario-m pairs; if b varies by group, the equation should reflect that explicitly.","section":"Section 4.6, Eq. (1)"},{"comment":"The conclusion says the study finds bias 'when it comes to race, gender, and ethnicity,' but the paper measures age, not ethnicity. Please correct this wording.","section":"Section 8"},{"comment":"The typo 'AA VE' in Section 6 should read 'AAVE,' and the demographic phrasing table would be easier to use if the number of phrasings per group were listed in a separate column.","section":"Appendix A.1"},{"comment":"Section 5.3 refers to 'Appendix B.4' for repeated trial results, but the appendix section is labeled 'B.3 Repeated Trial Results.' Also, Figure 1's caption says 'Difference from the control represents the frequency...' while the y-axis appears to be a proportion; please clarify the exact quantity plotted.","section":"Appendix B and Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The authors should be encouraged to release the full prompt set, per-scenario scores, and analysis code; the current manuscript does not provide enough detail to reproduce the exact t-tests or verify the reported degrees of freedom. The paper's fit with an NLP/ethics venue is appropriate, but the statistical revisions are substantial enough that a major revision is warranted rather than a minor one."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the control-group design is a real contribution, and the finding that most models rate non-baseline demographics as more likely to help is solid. But the paper's central claim—that the implicit default is a white middle-aged or young adult male—rests on interpreting non-significant t-tests as evidence of equivalence, and that inference doesn't hold.\n\nWhat's new: unlike prior work (Salinas et al., Kotek et al.), they include a no-demographic control condition and compute bias scores as differences from that control. That lets them separate positive bias toward minoritized groups from the model's baseline assumptions. The audit covers 10 commercial models with human-written prompts, 103 scenarios, and multiple rephrasings. The bias scores are large and mostly consistent across models: women, non-binary people, seniors, and most non-white racial groups are rated as more likely to help than an unspecified 'person.' That's a real empirical result, not a fitting artifact.\n\nWhere it's soft: the baseline conclusion. The paper treats groups whose scores don't differ significantly from the control as 'the default.' Non-significance is not evidence of equivalence. For several models the young-adult differences are actually positive and, per the text in Section 5.1.2, sometimes significant, yet the paper lumps young adults with middle-aged as the baseline. The text even contradicts itself: first saying 'most models displayed statistically significant positive bias' toward young adults, then saying 'the majority of models found no significant difference.' A proper equivalence test is needed before the baseline claim can stand. Also, there's no correction for multiple comparisons across the many model/demographic pairs, and the stated number of rephrasings (b=4) doesn't match the appendix, where many groups have only two or three phrasings. The prompts aren't published, which limits exact replication. The limitations section is candid about scope but doesn't acknowledge these statistical issues.\n\nBottom line: the paper deserves a serious referee. The control-group method is worth building on, and the positive-bias finding is important. But it needs a statistical revision: equivalence testing, a multiple-comparison correction, and a clearer report of which differences are significant. I'd send it to review with a request for major revision.","headline":"The control-group design is a genuine step forward in bias auditing, but the paper's headline 'default person' claim rests on treating non-significant t-tests as evidence of equivalence, and that inference doesn't hold.","tokens_in":11194,"tokens_out":2726,"would_cite":false,"duration_ms":27338,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In ten publicly available LLMs, an unmarked 'person' is implicitly treated as a white middle-aged or young-adult male, and most explicitly named demographics are rated as more willing to help than that unmarked control.","keywords":["LLM bias","demographic bias","control group","baseline demographic","Good Samaritan scenarios","bias score","prompt brittleness","paired t-test"],"falsifier":"Run the same 103 scenarios with an explicit 'typical person' control and use a pre-specified equivalence margin; if White, male, or middle-aged groups do not fall within that margin of the control, or if an explicit 'typical person' control produces a different baseline, then the claim that these groups are the implicit default is not supported.","tokens_in":10213,"feed_emoji":"🤖","tokens_out":6144,"duration_ms":64660,"temperature":0.7,"pith_summary":"This paper tries to establish that commercial LLMs carry an implicit 'default person'—white, male, and middle-aged or young adult—and that this baseline can be measured by adding a control prompt that names no demographic. Across 412 human-written Good Samaritan-style scenarios, ten models from five providers rated a genderless, raceless, ageless 'person' as less likely to help than most explicitly named groups, meaning the models' bias is mostly affirmative toward non-default demographics rather than negative toward them. The method matters because it separates two biases that earlier auditing approaches tangled together: the assumption of who the unmarked person is, and the directional preference once a demographic is named. It gives users and developers a quantitative, repeatable audit to detect both kinds of bias in LLM output.","feed_headline":"LLMs default to white middle-aged men—and favor everyone else","feed_subtitle":"Ten chatbots rate an unmarked 'person' as male, middle-aged, and white—then favor every other group.","key_machinery":"The central object is the bias score, defined as the difference between a demographic group's mean helpfulness rating and the control group's mean rating, where helpfulness is the model's 1-100 probability that a moral patient will intervene to help a third party. The control is a 'person' with no demographic information, and each of 103 scenario ideas is rephrased four ways with different grammar and word choice, producing roughly 2,800 human-written prompts per model run. Paired sample t-tests at α = 0.01 flag which demographic-versus-control differences are significant, and the standard deviation across the four rephrasings defines brittleness, which measures how sensitive scores are to wording rather than to demographics. This control-comparison design is what makes the implicit default demographic visible, because it shows which groups are statistically indistinguishable from an unmarked person.","core_discovery":"The paper claims that LLMs have a measurable implicit baseline demographic: when no demographic is specified, the models behave as if the 'person' in a scenario is white, male, and middle-aged or a young adult. Relative to that unmarked control, women, non-binary people, seniors, and most non-white racial groups were rated as significantly more likely to help, while white, male, and middle-aged groups usually did not differ significantly from the control. The authors interpret this as two distinct biases that are usually entangled in prior work: a default assumption about the unmarked person, and a directional preference once a demographic is named. Because they included a control condition with no demographic information, they could observe that most deviations from the baseline are affirmative, not negative, with teenagers as the main exception.","pith_inferences":["The manuscript cites '[llm-prompt-bias]' to justify using only human-written prompts, but no matching entry appears in its reference list, so the cited support for that design choice is missing from the bibliography.","The inference from non-significant t-tests to 'default' status would be stronger with an equivalence test and a pre-specified margin; as written, underpowered comparisons could masquerade as baseline agreement.","The affirmative bias toward non-baseline groups is consistent with post-training debiasing or overcorrection, which would make this control-comparison method a natural audit instrument for measuring the size and direction of such corrections.","The authors note that a different language could shift the default demographic; running the same prompt set in Mandarin would test whether the implied default person is a property of English training data or a more general model prior."],"forward_implications":["All ten models showed statistically significant bias in at least six demographic categories, so the pattern is not specific to one developer or model family.","Non-baseline groups—women, non-binary people, seniors, and most non-white racial groups—were rated as more likely to help than the unmarked control, so on this task the prevailing bias is affirmative rather than negative.","Teenagers were the clear exception, rated as less likely to help than the control by most models.","Adding a no-demographic control lets bias audits separate the implicit default from directional preference, a distinction that multiple-choice and sentiment-analysis methods cannot make.","Brittleness measurements show that wording changes move scores by 4–8 points on a 1–100 scale, sometimes more than the demographic biases themselves, so phrasing must be controlled in any audit."],"supporting_citations":[{"why":"Supplies the multiple-choice stereotype method and gender-bias evidence that this paper positions itself against.","marker":"Kotek et al., 2023"},{"why":"Provides the StereoSet benchmark for stereotype reinforcement, a prior multiple-choice bias-measurement approach.","marker":"Nadeem et al., 2021"},{"why":"Represents the open-ended generation plus sentiment-analysis class of bias studies, with gender-bias findings in reference letters.","marker":"Wan et al., 2023"},{"why":"Closest prior numeric audit of race and gender bias across life-event scenarios, motivating the unified helpfulness metric used here.","marker":"Salinas et al., 2025"},{"why":"Documents how pretraining data can produce model biases, motivating the need for deployment-phase bias audits.","marker":"Feng et al., 2023"},{"why":"Underpins the choice of paired t-tests as robust to normality violations and higher-power than nonparametric alternatives.","marker":"Rasch et al., 2007"},{"why":"Cited for the claim that model outputs depend heavily on word embeddings, which frames the language-and-dialect limitation.","marker":"Pennington et al., 2014"},{"why":"Cited to justify using human-written prompts because neural retrievers are biased toward LLM-generated content.","marker":"Dai et al., 2024"}],"fun_headline_variants":["LLMs assume white male default, then favor everyone else","Chatbots see white male as baseline, rate others kinder","Default human: white male; everyone else seen as nicer","LLMs' hidden default: white middle-aged male, then bias toward others","LLMs bias: default white male, but others appear more generous"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that when a demographic group's scores do not differ significantly from the control, that means the model treats that group as the default person; non-significance is not the same as proof of equivalence, so this inference carries the argument.","fun_headline_variants_meta":{"raw":{"variants":["LLMs assume white male default, then favor everyone else","Chatbots see white male as baseline, rate others kinder","Default human: white male; everyone else seen as nicer","LLMs' hidden default: white middle-aged male, then bias toward others","LLMs bias: default white male, but others appear more generous"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2824,"prompt_tokens":902,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1835}},"tokens_in":518,"tokens_out":1922,"duration_ms":15396,"temperature":1.0,"reasoning_tokens":1835,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:09:49.844015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 103 scenarios with an explicit 'typical person' control and use a pre-specified equivalence margin; if White, male, or middle-aged groups do not fall within that margin of the control, or if an explicit 'typical person' control produces a different baseline, then the claim that these groups are the implicit default is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Underpins the choice of paired t-tests as robust to normality violations and higher-power than nonparametric alternatives."}],"review_version":1}