{"id":"24bc0393-1a1d-4621-8006-f60ef3499590","arxiv_id":"2501.00367","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM recommendations of important AI research favor recent, well-cited, team-authored papers, but do not measurably over-represent male, white, or developed-country scholars relative to a human-curated benchmark.","lead":"This paper asked three large language models to recommend 50 important papers and 50 important scholars in four AI subfields, then compared the recommendations with expert-curated lists. The models favored recent, highly cited, collaborative work, yet showed no statistically significant tendency to favor male, white, or developed-country scholars.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The demographic null is uninterpretable: comparing LLM scholar lists to a citation-curated 50-person benchmark cannot detect bias, since a biased benchmark makes a null result the expected outcome under bias replication.","rationale":"The paper's strongest and most novel claim is the demographic null. Everything else (recency, team size, citation trends) can stand on its own, but the abstract and conclusion are built on the scholar-level comparison. That comparison has an invalid control condition: the 'real important scholars' list is a product of citation-based importance, and the paper's own literature review shows citation-based importance is demographically biased. Therefore the experiment cannot identify 'disproportionate' recommendation. This is not a disagreement with consensus; it is an internal design problem. The paper could have compared LLM lists to the field's eligible author population, or to a benchmark chosen by criteria independent of demographic proxies. It did neither. The low Kappa race labels and n=50 further undermine the null. I agree with the reader's weakest assumption and retain the REJECT verdict; no change needed. Credit: the prompt-based multi-model setup and the paper-level analyses are reasonable, and the paper is transparent about some limitations, but these do not rescue the central demographic inference.","tokens_in":14811,"tokens_out":5185,"duration_ms":55432,"concrete_test":"Re-run the scholar-level analysis with a demographically meaningful baseline instead of the 50-person curated list: for each of ML/DL/RL/NLP, take all authors with at least one indexed publication in OpenAlex (or the top 1,000 authors by citation count) and compute the field's female share, non-white share, and non-very-high-HDI share. Then compare each LLM's 50-scholar list against these base rates with a two-proportion z-test or Fisher's exact test. If the LLM lists show significantly lower representation of women, non-white authors, or developing-country authors relative to the field population while matching the curated benchmark, the paper's central null result is an artifact of the biased benchmark.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — 'no evidence that LLMs disproportionately recommend male, white, or developed-country authors' — rests on chi-square comparisons between each LLM's 50-name scholar list and the authors' manually curated 'real important scholars' list (Methods; Appendix). That reference list is built from citation counts and domain expertise. The introduction itself documents that citation counts and recognition carry gender, race, and country biases (refs 16–21). So the benchmark is not a neutral baseline; it is an artifact of the very human biases the paper claims to contrast. If the LLM reproduces the same biased canon, the test will show no significant difference by construction. A null result therefore cannot distinguish 'no demographic preference' from 'preference aligned with existing biased distributions.' The abstract's wording overstates what is identified. The problem is compounded by two further weaknesses: (1) with n=50 and highly skewed categories, chi-square has little power to detect anything short of a large effect; (2) race labels in Table 4 have Kappa as low as 0.33, so measurement noise further attenuates any real differences. The paper-level findings about recency, team size, and citation counts are less affected by this baseline problem, but they are secondary to the demographic headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether three large language models (ChatGPT-4o, Claude, and GLM) exhibit biases when recommending important papers and scholars in machine learning, deep learning, reinforcement learning, and natural language processing. The authors prompt each model to list 50 real papers and 50 real scholars per field, compare the outputs against a manually curated 'real important' list based on citation counts and domain expertise, and test four hypotheses about citation counts, gender, race, and country of recommended authors. They report that LLM recommendations have limited accuracy, favor recent and collaborative work, are less disruptive, and show no significant demographic differences relative to the benchmark, leading them to conclude that LLMs do not disproportionately recommend male, white, or developed-country authors, in contrast to known human biases.","tokens_in":15008,"tokens_out":5934,"duration_ms":52661,"significance":"The paper addresses a timely and important question about fairness in LLM-based scholarly recommendation. Its paper-level findings—that LLMs prefer recent, multi-author, incremental papers—are descriptively useful and could inform future research on recommendation systems. The study is also transparent in listing its limitations, including the small sample size and the expert-curated benchmark. However, the headline demographic conclusion is not identified by this design: because the benchmark is constructed from citation counts and expert judgment, which the paper's own introduction notes are biased along gender, race, and country lines, a null difference between LLM outputs and the benchmark cannot be interpreted as an absence of demographic bias. The low reliability of the race prediction labels (Kappa as low as 0.33) and the small, skewed samples further weaken the demographic tests. The demographic claim should not be presented as a contrast to known human biases.","major_comments":[{"comment":"The demographic null result (Hypotheses 2–4, Table 3) is uninterpretable as evidence of non-bias because the reference list is constructed from citation counts and domain expertise, and the Introduction explicitly documents that citations and recognition carry gender, race, and country biases (refs 16–21). If an LLM reproduces the same biased canon, the chi-square comparison against this benchmark will show no significant difference by construction. The abstract's claim that 'there is no evidence that LLMs disproportionately recommend male, white, or developed-country authors' therefore overstates what the experiment can establish; the correct interpretation is that LLM recommendations are not detectably different from a biased human-curated list.","section":"Introduction; Data; Methods; Results (Scholar Level); Discussion"},{"comment":"With only 50 scholars per field per model and highly skewed demographic categories (e.g., male dominance), the chi-square tests reported in Table 3 have very low power to detect anything short of a large effect. The paper reports no effect sizes, confidence intervals, or minimum detectable effects, so 'no significant difference' is a weak basis for the conclusion of no demographic preference. A null result at this sample size should be presented as inconclusive, not as evidence of absence.","section":"Results (Scholar Level); Table 3"},{"comment":"Table 4 reports inter-method Kappa values for race prediction as low as 0.33 (GPT4o, pred_fl_reg_name), and the authors themselves note that 'more stable methods for race prediction should be explored.' With this level of measurement error, any true racial difference in LLM recommendations would be attenuated, so Hypothesis 3 cannot be meaningfully tested with the current name-based race labels. The gender prediction agreement of 1.0 is reassuring, but race remains a load-bearing measurement problem.","section":"Robustness Checks; Table 4"},{"comment":"The Discussion's first stated limitation—that the benchmark was manually curated by experts and 'inherently limits the explanatory power and broader applicability of the findings'—applies directly to the demographic conclusion, which is drawn entirely from comparisons to that benchmark. The paper should either remove the demographic claim from the abstract and conclusions or redesign the benchmark to include a neutral baseline (e.g., population or publication-base rates) rather than an elite, citation-selected list.","section":"Discussion (Limitations)"},{"comment":"There is an inconsistency in the treatment of GLM: the Results state that 'we excluded GLM's outputs from subsequent visualizations' because of its high error rate, yet Figure 7 and the accompanying text discuss GLM's gender, race, and country distributions. The paper should clarify whether GLM is included in the scholar-level demographic analyses and, if so, how its high rate of fabricated recommendations affects those comparisons.","section":"Results (Error rate; Figure 7)"}],"minor_comments":[{"comment":"Throughout the manuscript, 'essays' is used where 'papers' or 'articles' is meant (e.g., in the prompts and in the section on error rates).","section":"Throughout"},{"comment":"The Results section refers to Table 1 for error rates, but the error rates appear in Table 2; the tables are misnumbered.","section":"Results"},{"comment":"The caption of Table 4 ('Performance of different LLMs with various settings and contexts') does not describe what the Kappa values measure; it should state that these are inter-method agreement scores for race and gender prediction.","section":"Table 4"},{"comment":"The phrase 'The table below (omitted here)' in the Robustness Checks is a leftover from a draft; the table is actually presented as Table 4.","section":"Robustness Checks"},{"comment":"Figure numbering is inconsistent: the text refers to Figure 1B, 1C, and 1D for the field-specific KDE plots, but Figure 1 is the methodology flowchart; the KDE plots appear to be in Figure 2.","section":"Figure 2"},{"comment":"The 'compensation' effect mentioned in the Conclusions—that models recommended a higher proportion of scholars from developing countries than the actual distribution—is not tested statistically; the chi-square tests in Table 3 are not reported for this specific contrast.","section":"Conclusions"},{"comment":"Hypothesis 1 is stated as 'LLM tends to recommend papers with higher citation counts,' but the results show LLM-recommended papers have lower average citation counts than the benchmark; the wording of the hypothesis and the rejection interpretation should be clarified.","section":"Table 1, Hypotheses"},{"comment":"Some entries in the 'Real Important Papers' Appendix appear to be outside the stated fields (e.g., 'Meta-analysis in clinical trials' under Machine Learning); the curation criteria should be documented more explicitly.","section":"Appendix"}],"recommendation":"reject","confidential_remarks":"The paper tackles a relevant topic, but the central demographic conclusion is not identified by the current design, and the limitations are acknowledged in the Discussion. If the authors were to reframe the paper as a descriptive comparison between LLM outputs and a citation-based expert list, it might be suitable for a more specialized venue, but as it stands the headline claim is misleading. The paper's own limitations section already contains the seeds of this critique."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the paper is a legitimate extension of LLM citation-bias work into paper and scholar recommendation, and the descriptive findings about recency, team size, and disruptiveness are worth a look. But the headline claim — that LLMs show no demographic bias — does not hold up, because the 'real important' benchmark is built from citation counts and expert judgment, the same dimensions documented to carry gender, race, and country bias. Comparing LLM outputs to that benchmark can't separate 'no bias' from 'replicates existing bias.' The null is just uninterpretable.\n\nWhat's new and good: they test three models across four fields, use a recommendation prompt rather than reference generation, and report precision/recall of the models' outputs. That's a useful empirical framing. The paper-level patterns — preference for recent papers, larger teams in some fields, conservative/incremental work — are plausible and supported by the descriptive stats, even if n=50 per field limits precision.\n\nSoft spots: (1) The demographic benchmark problem I mentioned. The authors even acknowledge in the Discussion that their benchmark is manually curated and limits explanatory power, but the abstract doesn't carry that caveat. (2) With n=50 and skewed categories, chi-square has little power; a null is not evidence of absence. (3) Race labels from name-based packages are noisy, with Kappa as low as 0.33 in Table 4. (4) The abstract says LLMs recommend papers with 'greater citation counts,' but the results say the opposite — average citation counts are lower than the real list. That's a direct contradiction. (5) No code or data released, so reproducibility is limited.\n\nWho's this for? Scientometrics folks and people building AI search/recommendation tools. The paper deserves a serious referee because the task framing is new and the descriptive findings are useful, but the demographic conclusion needs either a different benchmark or a re-framing as 'no evidence of additional demographic skew beyond the benchmark.' I'd send it to review with major-revision expectations, not desk reject it.\n\nMy recommendation: engage with it, but treat the demographic null as not established.","headline":"New task framing, but the demographic null is uninterpretable next to a citation-curated benchmark.","tokens_in":15553,"tokens_out":2511,"would_cite":false,"duration_ms":23193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLM scholar recommendations show no demographic bias, while paper recommendations favor recent, collaborative, incremental work.","keywords":["large language models","literature recommendation","algorithmic bias","gender bias","racial bias","country disparity","Matthew effect","science of science"],"falsifier":"Re-run the scholar-recommendation analysis with the comparison group changed from the manually curated \"real important\" list to all active researchers in the same subfields matched by publication count and citation impact, and with demographic labels from self-identification or verified records instead of name-inference packages. If the LLM lists show a significantly higher share of male, white, or developed-country scholars than that matched population, the paper's no-demographic-bias conclusion fails; the paper's own reported race-classification Kappa values as low as 0.33 indicate the labels may be too noisy to support the null result as it stands.","tokens_in":14592,"feed_emoji":"🤖","tokens_out":5189,"duration_ms":53031,"temperature":0.7,"pith_summary":"This paper asks whether large language models, asked to recommend the 50 most important papers or scholars in machine learning, deep learning, reinforcement learning, and natural language processing, reproduce known human biases in research exposure. It finds that LLM paper recommendations are only partly accurate, with precision ranging from roughly 3% to 20%, and that they skew toward recently published work, larger author teams, and developmental papers over disruptive ones. In scholar recommendations, the paper reports no statistically significant over-representation of male, white, or developed-country authors compared with its manually curated benchmark of \"real important\" scholars, contrasting with documented human citation and recognition patterns. The paper reads this null result as evidence that current LLM recommenders, at least in these fields, do not add demographic skew beyond what the benchmark already contains, and may even slightly compensate toward developing-country authors.","feed_headline":"No gender, race, or country bias in LLM scholar picks","feed_subtitle":"But recommendations favor recent, large-team, incremental papers—and precision is low.","key_machinery":"The central machinery is a comparison between LLM-generated lists and a manually curated benchmark of 50 important papers and 50 important scholars per field, built from citation counts and domain expertise. The paper measures differences with Kolmogorov-Smirnov tests for paper-level continuous attributes and chi-square tests for demographic proportions, using OpenAlex and SciSciNet for bibliographic metadata and the World Bank's Human Development Index to classify countries. Demographic attributes are inferred from names using packages such as Surgeo, Ethnicolr, sexmachine, and a combination of nameparser, nltk, and gender-guesser, with reported robustness Kappa values as low as 0.33 for race prediction.","core_discovery":"The central claim is that LLMs' literature recommendations are demographically neutral in a specific sense: chi-square tests show no significant difference between the gender, race, and nationality distributions of recommended scholars and the distributions in a hand-built benchmark of 50 \"real important\" scholars per field, so hypotheses 2, 3, and 4 are rejected. At the paper level, LLM recommendations favor more recent documents, larger teams, and conservative, incremental research, while recommended papers are not on average more highly cited than the real-important benchmark (hypothesis 1 is also rejected). Accuracy is limited, with Claude achieving the best precision at about 20%, ChatGPT-4o about 15%, and GLM about 3%, and the paper excludes GLM from most analyses because its recommendations were largely low-quality or fabricated.","pith_inferences":["The strongest test would replace the \"real important\" benchmark with a baseline of all active researchers matched by field and productivity; a demographic skew against that baseline could still exist even if no skew appears against the hand-curated list.","Because the race classifiers used in the paper reach Kappa values as low as 0.33, the race null result is the least stable of the four demographic conclusions and could plausibly flip under self-identified or manually verified race labels.","The paper-level preferences for recent, large-team, incremental work may have indirect demographic consequences: if historically underrepresented groups publish more often in smaller teams or in older foundational work, a seemingly neutral bias could still reduce their exposure over time.","The study covers only computer science and AI, so the demographic null result may not carry over to disciplines with different publication cultures, citation densities, or author-name conventions."],"forward_implications":["If the paper is correct, LLM-based scholarly recommenders in these fields do not currently amplify gender, race, or country-of-origin disparities beyond what citation-based rankings already contain, so fairness interventions may need to target the benchmarks and training data rather than the recommender layer alone.","The systematic preference for recent, large-team, incremental papers means LLM recommendations will tend to hide older foundational work and highly disruptive single-team scholarship, shaping the literature that new researchers encounter.","Low precision and high fabrication rates, especially for GLM, imply that unscreened LLM recommendations are unreliable for serious literature searches and need verification before use.","The slight, statistically insignificant over-representation of developing-country scholars could, if it persists in larger samples, offset some existing global disparities in research visibility."],"supporting_citations":[{"why":"Petiska's finding that ChatGPT favors highly cited, older publications in environmental science motivates the Matthew Effect research question.","marker":"[12]"},{"why":"Merton's canonical statement of the Matthew Effect provides the theoretical frame for Hypothesis 1.","marker":"[14]"},{"why":"Larivière et al. document global gender disparities in science, supplying the known-human-bias baseline for Hypothesis 2.","marker":"[16]"},{"why":"Hopkins et al. report publication disparities by race and ethnicity, supplying the baseline for Hypothesis 3.","marker":"[19]"},{"why":"Gomez et al. show that leading countries receive more citations for similar research, supplying the baseline for Hypothesis 4.","marker":"[20]"},{"why":"Lin et al.'s SciSciNet provides the disruptiveness index used to compare recommended papers with real important papers.","marker":"[22]"},{"why":"Priem et al.'s OpenAlex supplies citation counts, author details, institutional affiliations, and country information for both recommended and benchmark items.","marker":"[23]"},{"why":"Zeng et al.'s GLM-130B is one of the three models under test, representing a bilinguual model developed outside the US.","marker":"[24]"},{"why":"Wu et al.'s large-teams-versus-small-teams disruptiveness result grounds the paper's comparison of disruptive and developmental recommendations.","marker":"[35]"}],"fun_headline_variants":["LLM recs unbiased by gender, race, country","LLM picks favor recent, large-team papers","No demographic bias in LLM scholar recommendations","LLM recommendations fair but low precision"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole demographic comparison assumes that the manually curated benchmark of 50 \"real important\" scholars is an unbiased reflection of who deserves recommendation; if that benchmark already reflects human biases, then matching it exactly does not prove the LLM has no bias.","fun_headline_variants_meta":{"raw":{"variants":["LLM recs unbiased by gender, race, country","LLM picks favor recent, large-team papers","No demographic bias in LLM scholar recommendations","LLM recommendations fair but low precision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1236,"prompt_tokens":774,"completion_tokens":462,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":390,"completion_tokens_details":{"reasoning_tokens":404}},"tokens_in":390,"tokens_out":462,"duration_ms":5080,"temperature":1.0,"reasoning_tokens":404,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:52:44.988384+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the scholar-recommendation analysis with the comparison group changed from the manually curated \"real important\" list to all active researchers in the same subfields matched by publication count and citation impact, and with demographic labels from self-identification or verified records instead of name-inference packages. If the LLM lists show a significantly higher share of male, white, or developed-country scholars than that matched population, the paper's no-demographic-bias conclusion fails; the paper's own reported race-classification Kappa values as low as 0.33 indicate the labels may be too noisy to support the null result as it stands.","supporting_citations":[{"cited_title":"ChatGPT cites the most-cited articles and journals, relying solely on Google Scholar's citation counts. As a result, AI may amplify the Matthew Effect in environmental science","cited_arxiv_id":"2304.06794","evidence_quote":"Petiska's finding that ChatGPT favors highly cited, older publications in environmental science motivates the Matthew Effect research question."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Merton's canonical statement of the Matthew Effect provides the theoretical frame for Hypothesis 1."},{"cited_title":"& Sugimoto, C","cited_arxiv_id":null,"evidence_quote":"Larivière et al. document global gender disparities in science, supplying the known-human-bias baseline for Hypothesis 2."},{"cited_title":"L., Jawitz, J","cited_arxiv_id":null,"evidence_quote":"Hopkins et al. report publication disparities by race and ethnicity, supplying the baseline for Hypothesis 3."},{"cited_title":"J., Herman, A","cited_arxiv_id":null,"evidence_quote":"Gomez et al. show that leading countries receive more citations for similar research, supplying the baseline for Hypothesis 4."},{"cited_title":"& Wang, D","cited_arxiv_id":null,"evidence_quote":"Lin et al.'s SciSciNet provides the disruptiveness index used to compare recommended papers with real important papers."},{"cited_title":"& Evans, J","cited_arxiv_id":null,"evidence_quote":"Wu et al.'s large-teams-versus-small-teams disruptiveness result grounds the paper's comparison of disruptive and developmental recommendations."}],"review_version":1}