REVIEW 2 major objections 38 references
Full citation counts still rank papers and authors much like a sparse “backbone” of only the true intellectual sources.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 16:02 UTC pith:YBHHGUAM
load-bearing objection Solid multi-field empirics showing LLM-pruned citation backbones preserve paper and author rankings; the ranking-stability result is real, but the claim that the retained edges are true intellectual sources rests on thin validation. the 2 major comments →
The backbone of science: analysis of citation networks between papers and their sources
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When only the references that an LLM judges to be the true intellectual sources of each paper are retained, the resulting sparse “backbone” networks preserve the relative citation rankings of both papers and authors that are observed in the full citation network, even though the two prompts used to extract sources agree only modestly at the level of individual papers.
What carries the argument
LLM-extracted backbone networks: for every paper the model is asked (via two related prompts) to return only the “significant” or “inspirational” references; the resulting sparse directed graphs are then compared with the original citation network and with random and most-cited baselines.
Load-bearing premise
That a single large language model, guided by two hand-written prompts and checked only on a few dozen papers by hand, correctly recovers the true intellectual sources of a paper from anonymized full text.
What would settle it
A larger human-annotated gold standard of true sources for several hundred papers; if the LLM selections systematically miss the sources humans mark or if the backbone rankings diverge sharply from full-network rankings once those human sources are used, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper constructs citation 'backbones' for five fields by using DeepSeek-R1-Distill-Llama-70B with two prompts ('significant' and 'inspirational' references) on anonymized full texts from unarXive/OpenAlex. The resulting sparse networks H and I retain ~15–25% of edges. The authors compare them to the full network G and to random/most-cited baselines via in-degree distributions and power-law exponents (Table 3, Fig. 2), KS tests (Table 4), Pearson/Spearman correlations (Tables 5–6), clustering C(k), nearest-neighbor degree, robustness modularity R (Table 8), temporal growth of ⟨k⟩t and R (Fig. 5), excess degree Si (Eq. 2, Fig. 6), section-level placement (Fig. 8), and author-level in-strength/PageRank (Tables 11–12). They conclude that full-network rankings of papers and authors remain reliable proxies for impact even after discarding most references as non-sources.
Significance. If the LLM labels are approximately correct, the work supplies a scalable, multi-metric demonstration that citation rankings are robust to aggressive, context-aware sparsification. The multi-field design, dual-prompt consistency, random and most-cited baselines, excess-degree analysis, section-level diagnostics, and author-level extension are thorough and go well beyond typical citation-network papers. The finding that full networks already give a reliable relative ranking would be useful for scientometrics and research evaluation. The main limitation is that the central claim inherits the unvalidated accuracy of the LLM source extraction.
major comments (2)
- §3 (and Abstract): The load-bearing claim that full-network rankings remain reliable proxies for 'actual' source-based impact rests on the premise that the two LLM prompts recover true intellectual sources. The only external validation reported is 'initial manual tests … on few dozens of papers' with no inter-annotator agreement, expert gold standard, multi-model ablation, or public release of the extracted lists. Tables 5–6 and 11–12 show high rank correlations, but similar stability can arise from any non-random sparse filter that still prefers hubs (as the most-cited baselines already illustrate). Without stronger validation that retained edges are the true sources, the observed stability is consistent both with the authors’ interpretation and with the weaker claim that any context-aware sparse subgraph preserves rankings.
- §2.2–2.3 and Tables 3–4: The random baselines are often statistically indistinguishable from H/I by KS test, while most-cited baselines are distinguishable and more extreme. This weakens the assertion that the LLM selection is clearly non-random and non-popularity-driven; the structural similarity of H and I could partly reflect residual hub preference rather than prompt-specific recovery of sources. A clearer quantitative separation (or additional null models) is needed to support the stronger interpretation.
Circularity Check
No load-bearing circularity: purely empirical LLM-filtered network comparison; excess degree and rank correlations are descriptive, not fitted predictions.
full rationale
The paper constructs full citation networks G and two LLM-extracted backbone subgraphs H and I (via two prompts on anonymized full text), then compares in-degree distributions, clustering, nearest-neighbor degrees, robustness modularity, temporal growth, excess degree (Eq. 2), section placement, and author-level in-strength/PageRank. There is no claimed first-principles derivation or prediction that reduces to its inputs by construction. Excess degree Si(x)=ki(x)-(E(x)/E(G))ki(G) is an explicit null-model residual against random edge retention; it is not fitted to force ranking stability. Rank correlations (Tables 5–6, 11–12) and top-k overlaps are post-hoc empirical observations. Self-citations (e.g., robustness modularity [26], earlier citation-inflation work) supply analysis tools only and are not invoked as uniqueness theorems or load-bearing premises for the central claim that full-network rankings remain reliable. The sole validation of the LLM source labels is the authors’ own limited manual checks; that is a correctness/validation weakness, not circularity. Score 1 reflects only the minor, non-load-bearing self-citations of methodological tools.
Axiom & Free-Parameter Ledger
free parameters (2)
- LLM temperature and top-p =
0.1
- Jaccard restriction threshold =
3
axioms (3)
- domain assumption An LLM can identify the true intellectual sources of a paper from anonymized full text and numeric citation markers alone.
- standard math Power-law tails of in-degree distributions can be compared via maximum-likelihood exponents and two-sample KS tests.
- domain assumption Robustness modularity R quantifies community strength under random edge rewiring.
invented entities (2)
-
backbone of science (networks H and I)
no independent evidence
-
excess degree S_i
no independent evidence
read the original abstract
The bibliography of scientific papers lists items with variable degree of relevance for the contents of the paper itself. If we could identify the sources, i.e., the works that actually inspired the paper, their citations can help us uncover the genesis of scientific projects and would be more representative of the actual importance of papers and authors than the standard citation counts, when all references are considered. Here we present an analysis of the \textit{backbone of science}, i.e., the network of citations between papers and their sources. The latter are extracted from the full body of papers via Large Language Models (LLMs), which are currently very capable of correctly identifying the context in which a paper is cited. Using two different but related prompts, we find that the LLMs select only a small set of references, not taken at random, and that the resulting backbone networks are quite similar to each other with respect to their in-degree distributions, modularity, transitivity, and degree correlations. Backbone networks have higher heterogeneity in their in-degree distributions, compared to the full network, but the most cited papers are usually the same, with some important exceptions. Citation rankings among authors are also remarkably stable. We conclude that the full citation network, despite its redundancy with respect to the backbones, presents a reliable picture of the relative citation impact of papers and authors.
Figures
Reference graph
Works this paper leans on
-
[1]
Andres Algaba, Vincent Holst, Floriano Tori, Melika Mobini, Brecht Verbeken, Sylvia Wenmackers, and Vin- cent Ginis. How deep do large language models internalize scientific literature and citation practices?arXiv preprint arXiv:2504.02767, April 2025
Pith/arXiv arXiv 2025
-
[2]
Large language models reflect human citation patterns with a heightened citation bias
Andres Algaba, Carmen Mazijn, Vincent Holst, Floriano Tori, Sylvia Wenmackers, and Vincent Ginis. Large language models reflect human citation patterns with a heightened citation bias. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Findings of the Association for Computational Linguistics: NAACL 2025, pages 6844–6879. Association for Computational Li...
2025
-
[3]
Emergence of scaling in random networks.Science, 286(5439):509– 512, 1999
Albert-L´ aszl´ o Barab´ asi and R´ eka Albert. Emergence of scaling in random networks.Science, 286(5439):509– 512, 1999
1999
-
[4]
V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre. Fast unfolding of communities in large networks. J. Stat. Mech., P10008, 2008
2008
-
[5]
The anatomy of a large-scale hypertextual web search engine.Computer Networks and ISDN Systems, 30(1):107–117, 1998
Sergey Brin and Lawrence Page. The anatomy of a large-scale hypertextual web search engine.Computer Networks and ISDN Systems, 30(1):107–117, 1998
1998
-
[6]
Aaron Clauset, Cosma Rohilla Shalizi, and M. E. J. Newman. Power-law distributions in empirical data.SIAM Review, 51(4):661–703, 2009
2009
-
[7]
Crossref REST API.https://www.crossref.org, 2025
Crossref. Crossref REST API.https://www.crossref.org, 2025. Accessed: 2025
2025
-
[8]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
2025
-
[9]
S. N. Dorogovtsev, A. V. Goltsev, and J. F. F. Mendes. Pseudofractal scale-free web.Phys. Rev. E, 65:066122, Jun 2002
2002
-
[10]
J. L. Hodges. The significance probability of the smirnov two-sample test.Arkiv f¨ or Matematik, 3(5):469 – 486, 1958
1958
-
[11]
The distribution of the flora in the alpine zone.New Phytol., 11(2):37–50, 1912
Paul Jaccard. The distribution of the flora in the alpine zone.New Phytol., 11(2):37–50, 1912
1912
-
[12]
AnyStyle: Parser for bibliographic references.https://anystyle.io, 2025
Sylvester Keil. AnyStyle: Parser for bibliographic references.https://anystyle.io, 2025
2025
-
[13]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedatten- tion. InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023
2023
-
[14]
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors,Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013
2013
-
[15]
Evaluation of large language model performance and reliability for citations and references in scholarly writing: Cross-disciplinary study.J
Joseph Mugaanyi, Liuying Cai, Sumei Cheng, Caide Lu, and Jing Huang. Evaluation of large language model performance and reliability for citations and references in scholarly writing: Cross-disciplinary study.J. Med. Internet Res., 26:e52935, Apr 2024
2024
-
[16]
Optimal large language models to screen citations for systematic reviews.Res
Takehiko Oami, Yohei Okada, and Taka-aki Nakada. Optimal large language models to screen citations for systematic reviews.Res. Synth. Methods, 16(6):859–875, 2025
2025
-
[17]
Pan, Alexander M
Raj K. Pan, Alexander M. Petersen, Fabio Pammolli, and Santo Fortunato. The memory of science: Inflation, myopia, and the knowledge network.J. Informetr., 12(3):656–678, 2018
2018
-
[18]
Dynamical and correlation properties of the internet.Phys
Romualdo Pastor-Satorras, Alexei V´ azquez, and Alessandro Vespignani. Dynamical and correlation properties of the internet.Phys. Rev. Lett., 87:258701, Nov 2001
2001
-
[19]
Karl Pearson. Vii. note on regression and inheritance in the case of two parents.Proc. R. Soc. Lond., 58(347- 352):240–242, 12 1895. 18
-
[20]
Petersen, Raj K
Alexander M. Petersen, Raj K. Pan, Fabio Pammolli, and Santo Fortunato. Methods to account for citation inflation in research evaluation.Res. Policy, 48(7):1855–1865, 2019
2019
-
[21]
Citeme: Can language models accurately cite scientific claims? In A
Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. Citeme: Can language models accurately cite scientific claims? In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Processing Systems, volume 37, pages 7847–7877. Curran Associates, Inc., 2024
2024
-
[22]
A general theory of bibliometric and other cumulative advantage processes.J
Derek De Solla Price. A general theory of bibliometric and other cumulative advantage processes.J. Am. Soc. Inf. Sci., 27(5):292–306, 1976
1976
-
[23]
Universality of citation distributions: Toward an objective measure of scientific impact.Proc
Filippo Radicchi, Santo Fortunato, and Claudio Castellano. Universality of citation distributions: Toward an objective measure of scientific impact.Proc. Natl. Acad. Sci. U. S. A., 105(45):17268–17272, 2008
2008
-
[24]
Hierarchical organization in complex networks.Phys
Erzs´ ebet Ravasz and Albert-L´ aszl´ o Barab´ asi. Hierarchical organization in complex networks.Phys. Rev. E, 67:026112, Feb 2003
2003
-
[25]
unarXive: A Large Scholarly Data Set with Publications’ Full-Text, Annotated In-Text Citations, and Links to Metadata.Scientometrics, 125(3):3085–3108, December 2020
Tarek Saier and Michael F¨ arber. unarXive: A Large Scholarly Data Set with Publications’ Full-Text, Annotated In-Text Citations, and Links to Metadata.Scientometrics, 125(3):3085–3108, December 2020
2020
-
[26]
Silva, Aiiad Albeshri, Vijey Thayananthan, Wadee Alhalabi, and Santo Fortunato
Filipi N. Silva, Aiiad Albeshri, Vijey Thayananthan, Wadee Alhalabi, and Santo Fortunato. Robustness modularity in complex networks.Phys. Rev. E, 105:054308, May 2022
2022
-
[27]
The proof and measurement of association between two things.The American Journal of Psychology, 15(1):72–101, 1904
Charles Spearman. The proof and measurement of association between two things.The American Journal of Psychology, 15(1):72–101, 1904
1904
-
[28]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
2017
-
[29]
Collective dynamics of ‘small-world’networks.Nature, 393(6684):440– 442, 1998
Duncan J Watts and Steven H Strogatz. Collective dynamics of ‘small-world’networks.Nature, 393(6684):440– 442, 1998
1998
-
[30]
When large language models meet citation: A survey.arXiv preprint arXiv:2309.09727, 2023
Yang Zhang, Yufei Wang, Kai Wang, Quan Z Sheng, Lina Yao, Adnan Mahmood, Wei Emma Zhang, and Rongying Zhao. When large language models meet citation: A survey.arXiv preprint arXiv:2309.09727, 2023
Pith/arXiv arXiv 2023
-
[31]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Zican Dong, Yupeng Hou, Beichen Zhang, Yingqian Min, Junjie Zhang, Peiyu Liu, Xiaolei Wang, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Yiwen Hu, Jian-Yun Nie, and Ji-Rong Wen. A survey of large language models.Frontiers of Computer Science, 2...
-
[32]
Given article
**Focus Area Definition:** Thoroughly analyze the "Given article". Your analysis must concentrate *exclusively* on aspects relevant to the category "{category_name}". The definition for "{category_name}" is: "{category_description}"
-
[33]
1", "3",
**Selected Reference Identifiers (Refs):** After detailing your CoT, clearly list the reference numbers (e.g., "1", "3", "12") you have selected as most critical and relevant for "{category_name}". These should be only the reference numbers/identifiers
-
[34]
{category_name}
**Detailed Reasons for Each Selected Reference (Reasons):** For *each* reference number you listed in the step above, provide a detailed and specific explanation (e.g., 1-2 complete sentences per reference) stating *precisely why* that particular reference is critically important or highly relevant to "{category_name}" according to your analysis and the c...
-
[35]
A list of selected reference identifiers
-
[36]
refs", and
A list of detailed reasons explaining the relevance of each selected reference to that category . Your task is to parse this input text and accurately convert it into a single JSON object with the keys "refs", and "reasons". ### Input Text (contains selected references and detailed reasons): {output_text_of_first_trial} ### Instructions for JSON Generation:
-
[37]
refs" List:** Identify the list of reference numbers under a header similar to
**Extract "refs" List:** Identify the list of reference numbers under a header similar to "SELECTED_REFERENCES_FOR_...". Parse these numbers and format them as a list of strings in the JSON "refs" field. For example, if the text lists "1", "15", "42", the JSON field should be ‘["1", "15", "42"]‘. Ensure only the reference numbers/identifiers are included
-
[38]
reasons" List:** Identify the detailed reasons provided under a header similar to
**Extract "reasons" List:** Identify the detailed reasons provided under a header similar to "DETAILED_REASONS_FOR_...". Each reason corresponds to a reference in the "refs" list. Parse these reasons and format them as a list of strings for the JSON "reasons" field. The order and number of reasons must strictly match the order and number of items in the "...
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.