REVIEW 3 major objections 3 minor 1 cited by
Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that LLM-driven reference selection exhibits both a male-author preference and a majority-group bias, with effects growing in larger pools.
desk verdict The subject is important and the design is plausible, but the 'majority bias' claim needs to be checked against raw pool composition before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The controlled candidate-pool experiment is the central object carrying the argument. The authors construct reference pools with pseudonymous author names whose gender is signaled only by the name, vary the gender composition of the pool, and ask the LLMs to select references. Because the names are invented, content quality and reputation are stripped away, so any systematic selection pattern can be attributed to gender cues and pool composition. The models evaluated are GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku.
What would settle it
Run the same candidate-pool experiments using real author names or real papers instead of pseudonyms, or observe LLM-generated reference lists in real review tasks against a gender-balanced baseline; if the selection patterns disappear or reverse, the measured bias is an artifact of the invented-name setup rather than a property of LLM reference selection.
Extended reading notes
Core claim
The central claim is that LLM-driven reference selection is not gender-neutral. Across four evaluated models, selection shifts with the gender composition of the candidate pool: male-authored references are favored overall, and the more numerous gender in the pool is favored regardless of which gender that is. The effect is stronger in larger candidate pools, and the abstract reports that prompt-based attempts to reduce the bias have only modest effect. Field-level differences appear, with social sciences showing the least bias, indicating that the phenomenon is not uniform across scholarly domains.
Load-bearing premise
The load-bearing premise is that the models' choices among invented pseudonymous names in a controlled candidate pool predict how they would choose among real authors in actual literature review, even though real names carry extra associations.
Editorial extensions
If this is right
- LLM-assisted literature review can inherit and amplify gender imbalance in citation and scholarly recognition.
- Larger candidate pools worsen the bias, so simply expanding search results will not make LLM reference choice fairer.
- Prompt-based mitigation is not sufficient on its own; stronger interventions are needed to counter the bias.
- The bias is not uniform across fields, with social sciences showing the least bias, implying that discipline-specific calibration may be required.
- Integrating LLMs into high-stakes academic workflows could perpetuate existing gender disparities unless mitigation strategies are developed.
Reading between the lines
- If pseudonymous names behave differently from real author names, real-world bias could be larger (because real names carry reputational signals) or smaller (because models have richer information); a direct test with real author names in actual review tasks would settle the transfer.
- The majority-group bias suggests the model may be using base-rate information about the pool, implying the effect could extend beyond gender to any salient group cue in the candidate set.
- A testable extension is to vary the task framing, such as asking for 'the most influential references' versus 'a balanced reference list,' and measure whether the bias flips or persists.
- A public benchmark built from the candidate pools could let future models be scored on gender balance, turning the bias into a measurable evaluation target.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript investigates whether LLM-driven reference selection exhibits gender bias. Using controlled experiments with pseudonymous author names and varying gender ratios in candidate reference pools, it evaluates GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku across fields. The abstract claims a persistent male-author preference, a majority-group bias favoring the more prevalent gender, amplification with larger pools, modest attenuation by prompt-based mitigation, and field differences with social sciences showing the least bias.
Significance. If the claims hold, this is an important and timely contribution to understanding AI-mediated scholarly infrastructure. The experimental design, which manipulates pool gender composition while holding the selection task constant, is a sound approach for isolating bias, and the use of multiple model families and prompt mitigation is a strength. However, because the abstract reports no per-capita normalization, effect sizes, or uncertainty measures, the significance of the majority-group bias finding is currently unverifiable and could be trivial if it reflects raw composition effects.
major comments (3)
- [Abstract] The abstract's central claim of a 'majority-group bias' is indistinguishable from a mechanical composition effect unless selection rates are normalized by the number of candidates of each gender. Please state explicitly whether the reported effect is computed as per-capita selection probability (e.g., selection rate per male and female candidate, or odds ratios), and report the relevant denominators. If the effect is based on raw counts of selected references, the finding that 'whichever gender is more prevalent' is favored follows by construction and does not support the paper's conclusion.
- [Abstract] The use of pseudonymous author names is described, but the abstract gives no indication of how names were generated or whether they were matched on length, commonness, alphabetic order, and other attributes. If the name sets differ by gender on any correlated attribute, selection differences could reflect those attributes rather than gender. The full text must report generation and matching procedures, or show that results are robust to name permutation.
- [Abstract] The abstract claims bias is 'amplified in larger candidate pools' and that social sciences show 'the least bias,' but gives no effect sizes, confidence intervals, or trial counts. Without these quantities, the reader cannot assess the precision or practical magnitude of the reported differences; please include them in the results and abstract.
minor comments (3)
- [Abstract] The abstract should name the exact model versions and prompts used, as performance varies across snapshots; a brief description in the full text would suffice.
- [Abstract] The phrase 'persistent preference for male-authored references' should be accompanied by a measure of effect size (e.g., Cohen's d or odds ratio) to be interpretable.
- [Abstract] The paper would benefit from a brief statement of how the candidate pools were constructed, including the number of candidates, the gender ratios used, and whether fields were matched across conditions.
Circularity Check
No significant circularity: the abstract reports an empirical measurement, and no claimed result reduces to its inputs by construction.
full rationale
This is an abstract-only review of an empirical study. There is no derivation chain, fitted parameter, or self-citation invoked to establish the central claim. The 'majority-group bias' is described as a pattern observed across experimentally varied candidate-pool compositions; without further text showing that the bias is defined simply as the raw majority count, it cannot be equated with the pool composition by construction. Similarly, the reported male preference and pool-size amplification are empirical contrasts, not quantities fitted from the outcome they purport to predict. A possible methodological concern about whether the reported majority effect controls for candidate-base rates is a validity question, not a circularity, and cannot be established from the abstract. The paper does not rely on an ansatz smuggled in via citation or on a uniqueness theorem from the authors' prior work. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Pseudonymous author names convey gender to LLMs similarly to real names.
- domain assumption Unbiased reference selection would select candidates in proportion to their representation in the pool.
- domain assumption The four tested LLMs are representative of LLM-based citation tools.
Cite this review
Pith. "Pith review of Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection." pith.science (2026). https://pith.science/paper/WTZ3DMYN
@misc{pith2026250802740,
author = {Pith},
title = {Pith review of: Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/WTZ3DMYN}},
note = {Machine review of arXiv:2508.02740}
}
read the original abstract
Large language models (LLMs) are rapidly being adopted as research assistants, particularly for literature review and reference recommendation, yet little is known about whether they introduce demographic bias into citation workflows. This study systematically investigates gender bias in LLM-driven reference selection using controlled experiments with pseudonymous author names. We evaluate several LLMs (GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku) by varying gender composition within candidate reference pools and analyzing selection patterns across fields. Our results reveal two forms of bias: a persistent preference for male-authored references and a majority-group bias that favors whichever gender is more prevalent in the candidate pool. These biases are amplified in larger candidate pools and only modestly attenuated by prompt-based mitigation strategies. Field-level analysis indicates that bias magnitude varies across scientific domains, with social sciences showing the least bias. Our findings indicate that LLMs can reinforce or exacerbate existing gender imbalances in scholarly recognition. Effective mitigation strategies are needed to avoid perpetuating existing gender disparities in scientific citation practices before integrating LLMs into high-stakes academic workflows.
Forward citations
Cited by 1 Pith paper
-
Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews
GUEST gives SE researchers process recommendations for planning, conducting, and reporting GenAI-supported SLRs and independent GenAI tool evaluations under mandatory human oversight.
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.