Pith. sign in

REVIEW 5 major objections 8 minor 1 cited by

Who Gets Recommended? Investigating Gender, Race, and Country Disparities in Paper Recommendations from Large Language Models

T0 review · 5 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that LLM scholar recommendations show no demographic bias, while paper recommendations favor recent, collaborative, incremental work.

desk verdict New task framing, but the demographic null is uninterpretable next to a citation-curated benchmark. read the letter →

arxiv 2501.00367 v1 pith:KCNZKREN submitted 2024-12-31 cs.IR cs.CYcs.DL

classification cs.IRcs.CYcs.DL
keywords largelanguagemodelsliteraturerecommendationalgorithmicbiasgenderracialcountrydisparityMattheweffectscienceof
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models, asked to recommend the 50 most important papers or scholars in machine learning, deep learning, reinforcement learning, and natural language processing, reproduce known human biases in research exposure. It finds that LLM paper recommendations are only partly accurate, with precision ranging from roughly 3% to 20%, and that they skew toward recently published work, larger author teams, and developmental papers over disruptive ones. In scholar recommendations, the paper reports no statistically significant over-representation of male, white, or developed-country authors compared with its manually curated benchmark of "real important" scholars, contrasting with documented human citation and recognition patterns. The paper reads this null result as evidence that current LLM recommenders, at least in these fields, do not add demographic skew beyond what the benchmark already contains, and may even slightly compensate toward developing-country authors.

What carries the argument

The central machinery is a comparison between LLM-generated lists and a manually curated benchmark of 50 important papers and 50 important scholars per field, built from citation counts and domain expertise. The paper measures differences with Kolmogorov-Smirnov tests for paper-level continuous attributes and chi-square tests for demographic proportions, using OpenAlex and SciSciNet for bibliographic metadata and the World Bank's Human Development Index to classify countries. Demographic attributes are inferred from names using packages such as Surgeo, Ethnicolr, sexmachine, and a combination of nameparser, nltk, and gender-guesser, with reported robustness Kappa values as low as 0.33 for race prediction.

What would settle it

Re-run the scholar-recommendation analysis with the comparison group changed from the manually curated "real important" list to all active researchers in the same subfields matched by publication count and citation impact, and with demographic labels from self-identification or verified records instead of name-inference packages. If the LLM lists show a significantly higher share of male, white, or developed-country scholars than that matched population, the paper's no-demographic-bias conclusion fails; the paper's own reported race-classification Kappa values as low as 0.33 indicate the labels may be too noisy to support the null result as it stands.

Watch

Extended reading notes

Core claim

The central claim is that LLMs' literature recommendations are demographically neutral in a specific sense: chi-square tests show no significant difference between the gender, race, and nationality distributions of recommended scholars and the distributions in a hand-built benchmark of 50 "real important" scholars per field, so hypotheses 2, 3, and 4 are rejected. At the paper level, LLM recommendations favor more recent documents, larger teams, and conservative, incremental research, while recommended papers are not on average more highly cited than the real-important benchmark (hypothesis 1 is also rejected). Accuracy is limited, with Claude achieving the best precision at about 20%, ChatGPT-4o about 15%, and GLM about 3%, and the paper excludes GLM from most analyses because its recommendations were largely low-quality or fabricated.

Load-bearing premise

The whole demographic comparison assumes that the manually curated benchmark of 50 "real important" scholars is an unbiased reflection of who deserves recommendation; if that benchmark already reflects human biases, then matching it exactly does not prove the LLM has no bias.

Editorial extensions

If this is right

  • If the paper is correct, LLM-based scholarly recommenders in these fields do not currently amplify gender, race, or country-of-origin disparities beyond what citation-based rankings already contain, so fairness interventions may need to target the benchmarks and training data rather than the recommender layer alone.
  • The systematic preference for recent, large-team, incremental papers means LLM recommendations will tend to hide older foundational work and highly disruptive single-team scholarship, shaping the literature that new researchers encounter.
  • Low precision and high fabrication rates, especially for GLM, imply that unscreened LLM recommendations are unreliable for serious literature searches and need verification before use.
  • The slight, statistically insignificant over-representation of developing-country scholars could, if it persists in larger samples, offset some existing global disparities in research visibility.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The strongest test would replace the "real important" benchmark with a baseline of all active researchers matched by field and productivity; a demographic skew against that baseline could still exist even if no skew appears against the hand-curated list.
  • Because the race classifiers used in the paper reach Kappa values as low as 0.33, the race null result is the least stable of the four demographic conclusions and could plausibly flip under self-identified or manually verified race labels.
  • The paper-level preferences for recent, large-team, incremental work may have indirect demographic consequences: if historically underrepresented groups publish more often in smaller teams or in older foundational work, a seemingly neutral bias could still reduce their exposure over time.
  • The study covers only computer science and AI, so the demographic null result may not carry over to disciplines with different publication cultures, citation densities, or author-name conventions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 8 minor

Summary. The paper investigates whether three large language models (ChatGPT-4o, Claude, and GLM) exhibit biases when recommending important papers and scholars in machine learning, deep learning, reinforcement learning, and natural language processing. The authors prompt each model to list 50 real papers and 50 real scholars per field, compare the outputs against a manually curated 'real important' list based on citation counts and domain expertise, and test four hypotheses about citation counts, gender, race, and country of recommended authors. They report that LLM recommendations have limited accuracy, favor recent and collaborative work, are less disruptive, and show no significant demographic differences relative to the benchmark, leading them to conclude that LLMs do not disproportionately recommend male, white, or developed-country authors, in contrast to known human biases.

Significance. The paper addresses a timely and important question about fairness in LLM-based scholarly recommendation. Its paper-level findings—that LLMs prefer recent, multi-author, incremental papers—are descriptively useful and could inform future research on recommendation systems. The study is also transparent in listing its limitations, including the small sample size and the expert-curated benchmark. However, the headline demographic conclusion is not identified by this design: because the benchmark is constructed from citation counts and expert judgment, which the paper's own introduction notes are biased along gender, race, and country lines, a null difference between LLM outputs and the benchmark cannot be interpreted as an absence of demographic bias. The low reliability of the race prediction labels (Kappa as low as 0.33) and the small, skewed samples further weaken the demographic tests. The demographic claim should not be presented as a contrast to known human biases.

major comments (5)
  1. [Introduction; Data; Methods; Results (Scholar Level); Discussion] The demographic null result (Hypotheses 2–4, Table 3) is uninterpretable as evidence of non-bias because the reference list is constructed from citation counts and domain expertise, and the Introduction explicitly documents that citations and recognition carry gender, race, and country biases (refs 16–21). If an LLM reproduces the same biased canon, the chi-square comparison against this benchmark will show no significant difference by construction. The abstract's claim that 'there is no evidence that LLMs disproportionately recommend male, white, or developed-country authors' therefore overstates what the experiment can establish; the correct interpretation is that LLM recommendations are not detectably different from a biased human-curated list.
  2. [Results (Scholar Level); Table 3] With only 50 scholars per field per model and highly skewed demographic categories (e.g., male dominance), the chi-square tests reported in Table 3 have very low power to detect anything short of a large effect. The paper reports no effect sizes, confidence intervals, or minimum detectable effects, so 'no significant difference' is a weak basis for the conclusion of no demographic preference. A null result at this sample size should be presented as inconclusive, not as evidence of absence.
  3. [Robustness Checks; Table 4] Table 4 reports inter-method Kappa values for race prediction as low as 0.33 (GPT4o, pred_fl_reg_name), and the authors themselves note that 'more stable methods for race prediction should be explored.' With this level of measurement error, any true racial difference in LLM recommendations would be attenuated, so Hypothesis 3 cannot be meaningfully tested with the current name-based race labels. The gender prediction agreement of 1.0 is reassuring, but race remains a load-bearing measurement problem.
  4. [Discussion (Limitations)] The Discussion's first stated limitation—that the benchmark was manually curated by experts and 'inherently limits the explanatory power and broader applicability of the findings'—applies directly to the demographic conclusion, which is drawn entirely from comparisons to that benchmark. The paper should either remove the demographic claim from the abstract and conclusions or redesign the benchmark to include a neutral baseline (e.g., population or publication-base rates) rather than an elite, citation-selected list.
  5. [Results (Error rate; Figure 7)] There is an inconsistency in the treatment of GLM: the Results state that 'we excluded GLM's outputs from subsequent visualizations' because of its high error rate, yet Figure 7 and the accompanying text discuss GLM's gender, race, and country distributions. The paper should clarify whether GLM is included in the scholar-level demographic analyses and, if so, how its high rate of fabricated recommendations affects those comparisons.
minor comments (8)
  1. [Throughout] Throughout the manuscript, 'essays' is used where 'papers' or 'articles' is meant (e.g., in the prompts and in the section on error rates).
  2. [Results] The Results section refers to Table 1 for error rates, but the error rates appear in Table 2; the tables are misnumbered.
  3. [Table 4] The caption of Table 4 ('Performance of different LLMs with various settings and contexts') does not describe what the Kappa values measure; it should state that these are inter-method agreement scores for race and gender prediction.
  4. [Robustness Checks] The phrase 'The table below (omitted here)' in the Robustness Checks is a leftover from a draft; the table is actually presented as Table 4.
  5. [Figure 2] Figure numbering is inconsistent: the text refers to Figure 1B, 1C, and 1D for the field-specific KDE plots, but Figure 1 is the methodology flowchart; the KDE plots appear to be in Figure 2.
  6. [Conclusions] The 'compensation' effect mentioned in the Conclusions—that models recommended a higher proportion of scholars from developing countries than the actual distribution—is not tested statistically; the chi-square tests in Table 3 are not reported for this specific contrast.
  7. [Table 1, Hypotheses] Hypothesis 1 is stated as 'LLM tends to recommend papers with higher citation counts,' but the results show LLM-recommended papers have lower average citation counts than the benchmark; the wording of the hypothesis and the rejection interpretation should be clarified.
  8. [Appendix] Some entries in the 'Real Important Papers' Appendix appear to be outside the stated fields (e.g., 'Meta-analysis in clinical trials' under Machine Learning); the curation criteria should be documented more explicitly.

Circularity Check

0 steps flagged · score 2.0 of 10

No derivation-level circularity; the demographic null is an empirical benchmark comparison, though the benchmark is built from the same citation-recognition system under study.

full rationale

The paper's claims are empirical comparisons, not derivations. LLM recommendations are compared with a manually curated 'real important papers/scholars' benchmark selected via citation counts and domain expertise (Methods; Appendix). No parameter is fitted and then renamed a prediction; no equation is defined in terms of the quantity it claims to test; no uniqueness theorem or ansatz is imported from the authors' prior work. The only self-citation is ref. 21 (Tian & Bu 2024), used to support the premise that developing-country scholars occupy supporting roles; it is corroborated by ref. 20 and is not load-bearing for the central comparison. The main limitation is benchmark construct validity: because the reference list is curated from citation counts and expert judgment, and the paper itself cites evidence that citation and recognition carry gender, race, and country biases, a null chi-square against that benchmark cannot cleanly establish that LLMs lack demographic bias. However, this is a statistical and benchmark-validity weakness, not an equivalence by construction: the LLM distributions could in principle have differed from the curated list, and the robustness checks (e.g., Kappa as low as 0.33 for race labels in Table 4) acknowledge measurement noise. Thus there is no specific circular step in the derivation chain; score 2 reflects the minor benchmark-construction concern and the minor self-citation, not a forced reduction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper does not fit parameters to data in the usual sense, but it relies on hand-set thresholds and strong domain assumptions. The most consequential is that the expert-curated lists are treated as unbiased ground truth, which makes the demographic null result structurally unable to detect bias amplification if the ground truth already contains human bias.

free parameters (2)
  • citation threshold for authenticity = 100 citations
    Articles with fewer than 100 citations are automatically treated as fake recommendations; this cutoff is chosen by the authors' experience and determines which recommendations enter later distribution tests.
  • edit-distance threshold for title matching = not reported
    Recommendations whose title edit distance to the best OpenAlex match exceeds an experience-based threshold are deemed false. The threshold value is not specified, so its influence on error rates and on the remaining sample cannot be assessed.
assumptions (4)
  • domain assumption OpenAlex and SciSciNet metadata accurately reflect publication attributes, author affiliations, and countries.
    The entire analysis uses OpenAlex for citation counts, fields, and country information; errors in this metadata propagate to every comparison.
  • domain assumption Name-based packages (Surgeo, Ethnicolr, sexmachine, gender-guesser) provide valid race and gender labels.
    The reported race prediction Kappa values are as low as 0.33, so this assumption is questionable and load-bearing for the demographic null result.
  • domain assumption The manually curated lists of 50 important papers and scholars are a valid benchmark for 'real' importance.
    The lists are constructed by the authors with citation sorting and screening; they embed prior success measures and possible human biases, which undermines the demographic comparison.
  • domain assumption Absence of a statistically significant chi-square difference with n=50 licenses rejection of bias hypotheses.
    The paper interprets p>0.05 as rejecting hypotheses 2-4 without power analysis; with 50 items per list, the tests likely cannot detect small or even moderate disparities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Who Gets Recommended? Investigating Gender, Race, and Country Disparities in Paper Recommendations from Large Language Models." pith.science (2026). https://pith.science/paper/KCNZKREN

@misc{pith2026250100367,
  author       = {Pith},
  title        = {Pith review of: Who Gets Recommended? Investigating Gender, Race, and Country Disparities in Paper Recommendations from Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KCNZKREN}},
  note         = {Machine review of arXiv:2501.00367}
}
read the original abstract

This paper investigates the performance of several representative large models in the tasks of literature recommendation and explores potential biases in research exposure. The results indicate that not only LLMs' overall recommendation accuracy remains limited but also the models tend to recommend literature with greater citation counts, later publication date, and larger author teams. Yet, in scholar recommendation tasks, there is no evidence that LLMs disproportionately recommend male, white, or developed-country authors, contrasting with patterns of known human biases.

Figures

Figures reproduced from arXiv: 2501.00367 by the authors.

Figure 1
Figure 1. Flowchart of the methodology for the experimental process [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Kernel Density Estimation (KDE) of LLM recommendation results versus [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. The topic map of the domain of machine learning. [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The topic map of the domain of deep learning. [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: The topic map of the natural language processing. [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: The topic map of the domain of reinforcement learning. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Stacked bar chart of gender, race, and nationality distribution of [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whose Name Comes Up? Auditing LLM-Based Scholar Recommendations

    cs.CY 2025-05 conditional novelty 6.0 of 10

    An audit of six open-weight LLMs shows that AI-generated scholar recommendations favor senior, highly cited, White and male scientists and often fail multi-constraint queries.

Reference graph

Works this paper leans on

41 extracted references · 28 canonical work pages · cited by 1 Pith paper

  1. [1]

    & Almusharraf, N

    Imran, M. & Almusharraf, N. Analyzing the role of ChatGPT as a writing assistant at higher education level: A systematic review of the literature. CONT ED TECHNOLOGY 15, ep464 (2023)

  2. [2]

    Meyer, J. G. et al. ChatGPT and large language models in academia: opportunities and challenges. BioData Mining 16, 20 (2023)

  3. [3]

    ChatGPT for complex text evaluation tasks

    Thelwall, M. ChatGPT for complex text evaluation tasks. Journal of the Association for Information Science and Technology (2024),

  4. [4]

    F., Mohtasim, M

    Ali, N. F., Mohtasim, M. M., Mosharrof, S. & Krishna, T. G. Automated Literature Review Using NLP Techniques and LLM-Based Retrieval-Augmented Generation. Preprint at https://doi.org/10.48550/arXiv.2411.18583 (2024)

  5. [5]

    Tan, Z. et al. Large Language Models for Data Annotation and Synthesis: A Survey. Preprint at https://doi.org/10.48550/arXiv.2402.13446 (2024)

  6. [6]

    & Tan, J

    Jin, H., Zhang, Y ., Meng, D., Wang, J. & Tan, J. A Comprehensive Survey on Process-Oriented Automatic Text Summarization with Exploration of LLM-Based Methods. Preprint at https://doi.org/10.48550/arXiv.2403.02901 (2024)

  7. [7]

    Liang, W . et al. Can Large Language Models Provide Useful Feedback on Research Papers? A Large-Scale Empirical Analysis. NEJM AI 1, AIoa2400196 (2024)

  8. [8]

    & Schiebinger, L

    Zou, J. & Schiebinger, L. AI can be sexist and racist — it’s time to make it fair. Nature 559, 324–326 (2018)

Show all 41 references
  1. [9]

    & Kanchan, T

    Guleria, A., Krishan, K., Sharma, V . & Kanchan, T. ChatGPT: ethical concerns and challenges in academics and research. The Journal of Infection in Developing Countries WHO GETS RECOMMENDED? 28 17, 1292–1299 (2023)

  2. [10]

    & Zou, J

    Garg, N., Schiebinger, L., Jurafsky, D. & Zou, J. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences 115, E3635–E3644 (2018)

  3. [11]

    ChatGPT listed as author on research papers: many scientists disapprove

    Stokel-Walker, C. ChatGPT listed as author on research papers: many scientists disapprove. Nature 613, 620–621 (2023)

  4. [12]

    ChatGPT cites the most-cited articles and journals, relying solely on Google Scholar’s citation counts

    Petiska, E. ChatGPT cites the most-cited articles and journals, relying solely on Google Scholar’s citation counts. As a result, AI may amplify the Matthew Effect in environmental science. Preprint at https://doi.org/10.48550/arXiv.2304.06794 (2023)

  5. [13]

    Gallegos, I. O. et al. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 50, 1097–1179 (2024)

  6. [14]

    Merton, R. K. The Matthew Effect in Science. Science 159, 56–63 (1968)

  7. [15]

    & Gingras, Y

    Larivière, V . & Gingras, Y . The impact factor’s Matthew Effect: A natural experiment in bibliometrics. Journal of the American Society for Information Science and Technology 61, 424–427 (2010)

  8. [16]

    & Sugimoto, C

    Larivière, V ., Ni, C., Gingras, Y ., Cronin, B. & Sugimoto, C. R. Bibliometrics: Global gender disparities in science. Nature 504, 211–213 (2013)

  9. [17]

    Teich, E. G. et al. Citation inequity and gendered citation practices in contemporary physics. Nat. Phys. 18, 1161–1170 (2022)

  10. [18]

    & Pujara, J

    Lerman, K., Yu, Y ., Morstatter, F. & Pujara, J. Gendered citation patterns among the scientific elite. Proc. Natl. Acad. Sci. U.S.A. 119, e2206070119 (2022)

  11. [19]

    L., Jawitz, J

    Hopkins, A. L., Jawitz, J. W ., McCarty, C., Goldman, A. & Basu, N. B. Disparities WHO GETS RECOMMENDED? 29 in publication patterns by gender, race and ethnicity based on a survey of a random sample of authors. Scientometrics 96, 515–534 (2013)

  12. [20]

    J., Herman, A

    Gomez, C. J., Herman, A. C. & Parigi, P. Leading countries in global science increasingly receive more citations than other countries doing similar research. Nat Hum Behav 6, 919–929 (2022)

  13. [21]

    Tian, Y . & Bu, Y . Developed Countries Dominate Leading Roles in International Scientific Collaborations: Evidence from Scholars’ Self-Reported Contribution in Publications. Proceedings of the Association for Information Science and Technology 61, 1104–1106 (2024)

  14. [22]

    & Wang, D

    Lin, Z., Yin, Y ., Liu, L. & Wang, D. SciSciNet: A large-scale open data lake for the science of science research. Sci Data 10, 315 (2023)

  15. [23]

    & Orr, R

    Priem, J., Piwowar, H. & Orr, R. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. Preprint at https://doi.org/10.48550/arXiv.2205.01833 (2022)

  16. [24]

    Zeng, A. et al. GLM-130B: An Open Bilingual Pre-trained Model. Preprint at https://doi.org/10.48550/arXiv.2210.02414 (2023)

  17. [25]

    M., Mao, R., Cambria, E

    Amin, M. M., Mao, R., Cambria, E. & Schuller, B. W. A Wide Evaluation of ChatGPT on Affective Computing Tasks. IEEE Transactions on Affective Computing 15, 2204–2212 (2024)

  18. [26]

    Peng, K. et al. Towards Making the Most of ChatGPT for Machine Translation. Preprint at https://doi.org/10.48550/arXiv.2303.13780 (2023)

  19. [27]

    Brown, T. B. et al. Language Models are Few-Shot Learners. Preprint at https://doi.org/10.48550/arXiv.2005.14165 (2020). WHO GETS RECOMMENDED? 30

  20. [28]

    & Zou, J

    Chen, L., Zaharia, M. & Zou, J. How is ChatGPT’s behavior changing over time? Preprint at https://doi.org/10.48550/arXiv.2307.09009 (2023)

  21. [29]

    Sinha, A. et al. An Overview of Microsoft Academic Service (MAS) and Applications. in Proceedings of the 24th International Conference on World Wide Web 243–246 (Association for Computing Machinery, New York, NY , USA, 2015). doi:10.1145/2740908.2742839

  22. [30]

    & Cohan, A

    Beltagy, I., Lo, K. & Cohan, A. SciBERT: A Pretrained Language Model for Scientific Text. Preprint at https://doi.org/10.48550/arXiv.1903.10676 (2019)

  23. [31]

    & Wang, K

    Dong, Y ., Ma, H., Shen, Z. & Wang, K. A Century of Science: Globalization of Scientific Collaborations, Citations, and Innovations. in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 1437– 1446 (ACM, Halifax NS Canada, 2017)....

  24. [32]

    Kinney, R. et al. The Semantic Scholar Open Data Platform. Preprint at https://doi.org/10.48550/arXiv.2301.10140 (2023)

  25. [33]

    & Huang, W

    Bu, Y ., Li, M., Gu, W. & Huang, W. Topic diversity: A discipline scheme-free diversity measurement for journals. Journal of the Association for Information Science and Technology 72, 523–539 (2021)

  26. [34]

    Leydesdorff, L., Wagner, C. S. & Bornmann, L. Interdisciplinarity as diversity in citation patterns among journals: Rao-Stirling diversity, relative variety, and the Gini coefficient. Journal of Informetrics 13, 255–269 (2019)

  27. [35]

    & Evans, J

    Wu, L., Wang, D. & Evans, J. A. Large teams develop and small teams disrupt science and technology. Nature 566, 378–382 (2019)

  28. [36]

    & Loper, E

    Bird, S., Klein, E. & Loper, E. Natural Language Processing with Python: WHO GETS RECOMMENDED? 31 Analyzing Text with the Natural Language Toolkit. (O’Reilly Media, Inc., 2009)

  29. [37]

    Walters, W . H. & Wilder, E. I. Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep 13, 14045 (2023)

  30. [38]

    the answer

    Qureshi, R. et al. Are ChatGPT and large language models “the answer” to bringing us closer to systematic review automation? Syst Rev 12, 72 (2023)

  31. [39]

    & Kurt, Z

    Thelwall, M. & Kurt, Z. Research evaluation with ChatGPT: Is it age, country, length, or field biased? Preprint at https://doi.org/10.48550/arXiv.2411.09768 (2024)

  32. [40]

    Algaba, A. et al. Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias. Preprint at https://doi.org/10.48550/arXiv.2405.15739 (2024)

  33. [41]

    M., Stańczak, K., Guidotti, R

    Manerba, M. M., Stańczak, K., Guidotti, R. & Augenstein, I. Social Bias Probing: Fairness Benchmarking for Language Models. Preprint at https://doi.org/10.48550/arXiv.2311.09090 (2024). APPENDIX ‘Real’ Important Papers List ML DL RL NLP Deep Residual Learning for Image Recogni...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.