Pith. sign in

REVIEW 2 major objections 38 references

Full citation counts still rank papers and authors much like a sparse “backbone” of only the true intellectual sources.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 16:02 UTC pith:YBHHGUAM

load-bearing objection Solid multi-field empirics showing LLM-pruned citation backbones preserve paper and author rankings; the ranking-stability result is real, but the claim that the retained edges are true intellectual sources rests on thin validation. the 2 major comments →

arxiv 2607.09771 v1 pith:YBHHGUAM submitted 2026-07-07 cs.DL cs.SIphysics.soc-ph

The backbone of science: analysis of citation networks between papers and their sources

classification cs.DL cs.SIphysics.soc-ph
keywords citation networksbackbone of sciencelarge language modelssource extractioncitation impactscientometricsdegree distributions
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Scientific papers cite many works, but only a few of them truly inspired or enabled the research. Using a large language model on the full text of papers from five fields, the authors extract those core “sources” and build sparse citation networks they call the backbone of science. Two different prompts produce largely different source lists for each paper, yet the resulting backbones look structurally similar to each other: they are more heterogeneous than the full networks, but the highest-cited papers and authors remain largely the same. The full citation network therefore contains a great deal of redundancy, yet it still gives a reliable picture of relative impact. A reader who cares about how credit is assigned in science will find that ordinary citation rankings are surprisingly robust to the removal of most references.

Core claim

When only the references that an LLM judges to be the true intellectual sources of each paper are retained, the resulting sparse “backbone” networks preserve the relative citation rankings of both papers and authors that are observed in the full citation network, even though the two prompts used to extract sources agree only modestly at the level of individual papers.

What carries the argument

LLM-extracted backbone networks: for every paper the model is asked (via two related prompts) to return only the “significant” or “inspirational” references; the resulting sparse directed graphs are then compared with the original citation network and with random and most-cited baselines.

Load-bearing premise

That a single large language model, guided by two hand-written prompts and checked only on a few dozen papers by hand, correctly recovers the true intellectual sources of a paper from anonymized full text.

What would settle it

A larger human-annotated gold standard of true sources for several hundred papers; if the LLM selections systematically miss the sources humans mark or if the backbone rankings diverge sharply from full-network rankings once those human sources are used, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper constructs citation 'backbones' for five fields by using DeepSeek-R1-Distill-Llama-70B with two prompts ('significant' and 'inspirational' references) on anonymized full texts from unarXive/OpenAlex. The resulting sparse networks H and I retain ~15–25% of edges. The authors compare them to the full network G and to random/most-cited baselines via in-degree distributions and power-law exponents (Table 3, Fig. 2), KS tests (Table 4), Pearson/Spearman correlations (Tables 5–6), clustering C(k), nearest-neighbor degree, robustness modularity R (Table 8), temporal growth of ⟨k⟩t and R (Fig. 5), excess degree Si (Eq. 2, Fig. 6), section-level placement (Fig. 8), and author-level in-strength/PageRank (Tables 11–12). They conclude that full-network rankings of papers and authors remain reliable proxies for impact even after discarding most references as non-sources.

Significance. If the LLM labels are approximately correct, the work supplies a scalable, multi-metric demonstration that citation rankings are robust to aggressive, context-aware sparsification. The multi-field design, dual-prompt consistency, random and most-cited baselines, excess-degree analysis, section-level diagnostics, and author-level extension are thorough and go well beyond typical citation-network papers. The finding that full networks already give a reliable relative ranking would be useful for scientometrics and research evaluation. The main limitation is that the central claim inherits the unvalidated accuracy of the LLM source extraction.

major comments (2)
  1. §3 (and Abstract): The load-bearing claim that full-network rankings remain reliable proxies for 'actual' source-based impact rests on the premise that the two LLM prompts recover true intellectual sources. The only external validation reported is 'initial manual tests … on few dozens of papers' with no inter-annotator agreement, expert gold standard, multi-model ablation, or public release of the extracted lists. Tables 5–6 and 11–12 show high rank correlations, but similar stability can arise from any non-random sparse filter that still prefers hubs (as the most-cited baselines already illustrate). Without stronger validation that retained edges are the true sources, the observed stability is consistent both with the authors’ interpretation and with the weaker claim that any context-aware sparse subgraph preserves rankings.
  2. §2.2–2.3 and Tables 3–4: The random baselines are often statistically indistinguishable from H/I by KS test, while most-cited baselines are distinguishable and more extreme. This weakens the assertion that the LLM selection is clearly non-random and non-popularity-driven; the structural similarity of H and I could partly reflect residual hub preference rather than prompt-specific recovery of sources. A clearer quantitative separation (or additional null models) is needed to support the stronger interpretation.

Circularity Check

0 steps flagged

No load-bearing circularity: purely empirical LLM-filtered network comparison; excess degree and rank correlations are descriptive, not fitted predictions.

full rationale

The paper constructs full citation networks G and two LLM-extracted backbone subgraphs H and I (via two prompts on anonymized full text), then compares in-degree distributions, clustering, nearest-neighbor degrees, robustness modularity, temporal growth, excess degree (Eq. 2), section placement, and author-level in-strength/PageRank. There is no claimed first-principles derivation or prediction that reduces to its inputs by construction. Excess degree Si(x)=ki(x)-(E(x)/E(G))ki(G) is an explicit null-model residual against random edge retention; it is not fitted to force ranking stability. Rank correlations (Tables 5–6, 11–12) and top-k overlaps are post-hoc empirical observations. Self-citations (e.g., robustness modularity [26], earlier citation-inflation work) supply analysis tools only and are not invoked as uniqueness theorems or load-bearing premises for the central claim that full-network rankings remain reliable. The sole validation of the LLM source labels is the authors’ own limited manual checks; that is a correctness/validation weakness, not circularity. Score 1 reflects only the minor, non-load-bearing self-citations of methodological tools.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 2 invented entities

The central ranking-stability claim rests on the empirical behavior of one LLM under two prompts, on standard network null models, and on the operational definition of “source.” No free parameters are fitted to force the ranking result; the main unproved premises are the LLM’s contextual accuracy and the field-selection criteria.

free parameters (2)
  • LLM temperature and top-p = 0.1
    Both set to 0.1 by hand to reduce stochasticity; different values could change selected reference sets.
  • Jaccard restriction threshold = 3
    Overlap statistics computed only for papers where at least one prompt returned ≥3 references; the cutoff is chosen by the authors.
axioms (3)
  • domain assumption An LLM can identify the true intellectual sources of a paper from anonymized full text and numeric citation markers alone.
    Stated as the enabling premise in §1 and §4.2; only light manual checks are reported.
  • standard math Power-law tails of in-degree distributions can be compared via maximum-likelihood exponents and two-sample KS tests.
    Standard network-science practice used throughout §2.2.
  • domain assumption Robustness modularity R quantifies community strength under random edge rewiring.
    Adopted from Silva et al. (2022) and used as a structural metric in §2.3–2.4.
invented entities (2)
  • backbone of science (networks H and I) no independent evidence
    purpose: Citation subgraph retaining only LLM-selected source references.
    Operational definition introduced in the abstract and §1; no independent existence outside the extraction procedure.
  • excess degree S_i no independent evidence
    purpose: Measures over- or under-selection of a paper relative to random edge removal (Eq. 2).
    New node-level diagnostic introduced in §2.5 to expose prompt-specific preferences.

pith-pipeline@v1.1.0-grok45 · 27720 in / 2600 out tokens · 27239 ms · 2026-07-14T16:02:47.199614+00:00 · methodology

0 comments
read the original abstract

The bibliography of scientific papers lists items with variable degree of relevance for the contents of the paper itself. If we could identify the sources, i.e., the works that actually inspired the paper, their citations can help us uncover the genesis of scientific projects and would be more representative of the actual importance of papers and authors than the standard citation counts, when all references are considered. Here we present an analysis of the \textit{backbone of science}, i.e., the network of citations between papers and their sources. The latter are extracted from the full body of papers via Large Language Models (LLMs), which are currently very capable of correctly identifying the context in which a paper is cited. Using two different but related prompts, we find that the LLMs select only a small set of references, not taken at random, and that the resulting backbone networks are quite similar to each other with respect to their in-degree distributions, modularity, transitivity, and degree correlations. Backbone networks have higher heterogeneity in their in-degree distributions, compared to the full network, but the most cited papers are usually the same, with some important exceptions. Citation rankings among authors are also remarkably stable. We conclude that the full citation network, despite its redundancy with respect to the backbones, presents a reliable picture of the relative citation impact of papers and authors.

Figures

Figures reproduced from arXiv: 2607.09771 by Dimitri Marinelli, Gaurang Singh Yadav, Santo Fortunato, Satyaki Sikdar, Wonhee Jeong.

Figure 1
Figure 1. Figure 1: Histogram of the Jaccard similarity index [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: In-degree distributions across research fields and network types. Panels (a–o) are organized in a 3 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Degree-dependent clustering coefficient. The layout, color scheme, and field organization are identical to [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Average normalized degree of nearest-neighbors as a function of degree. The layout, color scheme, and [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Temporal evolution of the average in-degree (top panels) and robustness modularity (bottom panels) [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Node-level excess degrees (S(H), S(I)) across five fields. Rows categorize the analysis type: the top row (a–e) and middle row (f–j) display histograms of excess degrees for backbone networks H and I, respectively, while the bottom row (k–o) presents scatter plots comparing S(H) against S(I). To show the full distribution of excess degree, including small values that would otherwise be obscured by extreme … view at source ↗
Figure 7
Figure 7. Figure 7: Relation between the consistency of reference selection for the two different prompts and the number of [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Relative ratio of references across paper sections. The bars indicate the relative ratio of citation fractions [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 2 linked inside Pith

  1. [1]

    How deep do large language models internalize scientific literature and citation practices?arXiv preprint arXiv:2504.02767, April 2025

    Andres Algaba, Vincent Holst, Floriano Tori, Melika Mobini, Brecht Verbeken, Sylvia Wenmackers, and Vin- cent Ginis. How deep do large language models internalize scientific literature and citation practices?arXiv preprint arXiv:2504.02767, April 2025

  2. [2]

    Large language models reflect human citation patterns with a heightened citation bias

    Andres Algaba, Carmen Mazijn, Vincent Holst, Floriano Tori, Sylvia Wenmackers, and Vincent Ginis. Large language models reflect human citation patterns with a heightened citation bias. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Findings of the Association for Computational Linguistics: NAACL 2025, pages 6844–6879. Association for Computational Li...

  3. [3]

    Emergence of scaling in random networks.Science, 286(5439):509– 512, 1999

    Albert-L´ aszl´ o Barab´ asi and R´ eka Albert. Emergence of scaling in random networks.Science, 286(5439):509– 512, 1999

  4. [4]

    V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre. Fast unfolding of communities in large networks. J. Stat. Mech., P10008, 2008

  5. [5]

    The anatomy of a large-scale hypertextual web search engine.Computer Networks and ISDN Systems, 30(1):107–117, 1998

    Sergey Brin and Lawrence Page. The anatomy of a large-scale hypertextual web search engine.Computer Networks and ISDN Systems, 30(1):107–117, 1998

  6. [6]

    Aaron Clauset, Cosma Rohilla Shalizi, and M. E. J. Newman. Power-law distributions in empirical data.SIAM Review, 51(4):661–703, 2009

  7. [7]

    Crossref REST API.https://www.crossref.org, 2025

    Crossref. Crossref REST API.https://www.crossref.org, 2025. Accessed: 2025

  8. [8]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

    DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

  9. [9]

    S. N. Dorogovtsev, A. V. Goltsev, and J. F. F. Mendes. Pseudofractal scale-free web.Phys. Rev. E, 65:066122, Jun 2002

  10. [10]

    J. L. Hodges. The significance probability of the smirnov two-sample test.Arkiv f¨ or Matematik, 3(5):469 – 486, 1958

  11. [11]

    The distribution of the flora in the alpine zone.New Phytol., 11(2):37–50, 1912

    Paul Jaccard. The distribution of the flora in the alpine zone.New Phytol., 11(2):37–50, 1912

  12. [12]

    AnyStyle: Parser for bibliographic references.https://anystyle.io, 2025

    Sylvester Keil. AnyStyle: Parser for bibliographic references.https://anystyle.io, 2025

  13. [13]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedatten- tion. InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023

  14. [14]

    Distributed representations of words and phrases and their compositionality

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors,Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013

  15. [15]

    Evaluation of large language model performance and reliability for citations and references in scholarly writing: Cross-disciplinary study.J

    Joseph Mugaanyi, Liuying Cai, Sumei Cheng, Caide Lu, and Jing Huang. Evaluation of large language model performance and reliability for citations and references in scholarly writing: Cross-disciplinary study.J. Med. Internet Res., 26:e52935, Apr 2024

  16. [16]

    Optimal large language models to screen citations for systematic reviews.Res

    Takehiko Oami, Yohei Okada, and Taka-aki Nakada. Optimal large language models to screen citations for systematic reviews.Res. Synth. Methods, 16(6):859–875, 2025

  17. [17]

    Pan, Alexander M

    Raj K. Pan, Alexander M. Petersen, Fabio Pammolli, and Santo Fortunato. The memory of science: Inflation, myopia, and the knowledge network.J. Informetr., 12(3):656–678, 2018

  18. [18]

    Dynamical and correlation properties of the internet.Phys

    Romualdo Pastor-Satorras, Alexei V´ azquez, and Alessandro Vespignani. Dynamical and correlation properties of the internet.Phys. Rev. Lett., 87:258701, Nov 2001

  19. [19]

    Karl Pearson. Vii. note on regression and inheritance in the case of two parents.Proc. R. Soc. Lond., 58(347- 352):240–242, 12 1895. 18

  20. [20]

    Petersen, Raj K

    Alexander M. Petersen, Raj K. Pan, Fabio Pammolli, and Santo Fortunato. Methods to account for citation inflation in research evaluation.Res. Policy, 48(7):1855–1865, 2019

  21. [21]

    Citeme: Can language models accurately cite scientific claims? In A

    Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. Citeme: Can language models accurately cite scientific claims? In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Processing Systems, volume 37, pages 7847–7877. Curran Associates, Inc., 2024

  22. [22]

    A general theory of bibliometric and other cumulative advantage processes.J

    Derek De Solla Price. A general theory of bibliometric and other cumulative advantage processes.J. Am. Soc. Inf. Sci., 27(5):292–306, 1976

  23. [23]

    Universality of citation distributions: Toward an objective measure of scientific impact.Proc

    Filippo Radicchi, Santo Fortunato, and Claudio Castellano. Universality of citation distributions: Toward an objective measure of scientific impact.Proc. Natl. Acad. Sci. U. S. A., 105(45):17268–17272, 2008

  24. [24]

    Hierarchical organization in complex networks.Phys

    Erzs´ ebet Ravasz and Albert-L´ aszl´ o Barab´ asi. Hierarchical organization in complex networks.Phys. Rev. E, 67:026112, Feb 2003

  25. [25]

    unarXive: A Large Scholarly Data Set with Publications’ Full-Text, Annotated In-Text Citations, and Links to Metadata.Scientometrics, 125(3):3085–3108, December 2020

    Tarek Saier and Michael F¨ arber. unarXive: A Large Scholarly Data Set with Publications’ Full-Text, Annotated In-Text Citations, and Links to Metadata.Scientometrics, 125(3):3085–3108, December 2020

  26. [26]

    Silva, Aiiad Albeshri, Vijey Thayananthan, Wadee Alhalabi, and Santo Fortunato

    Filipi N. Silva, Aiiad Albeshri, Vijey Thayananthan, Wadee Alhalabi, and Santo Fortunato. Robustness modularity in complex networks.Phys. Rev. E, 105:054308, May 2022

  27. [27]

    The proof and measurement of association between two things.The American Journal of Psychology, 15(1):72–101, 1904

    Charles Spearman. The proof and measurement of association between two things.The American Journal of Psychology, 15(1):72–101, 1904

  28. [28]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017

  29. [29]

    Collective dynamics of ‘small-world’networks.Nature, 393(6684):440– 442, 1998

    Duncan J Watts and Steven H Strogatz. Collective dynamics of ‘small-world’networks.Nature, 393(6684):440– 442, 1998

  30. [30]

    When large language models meet citation: A survey.arXiv preprint arXiv:2309.09727, 2023

    Yang Zhang, Yufei Wang, Kai Wang, Quan Z Sheng, Lina Yao, Adnan Mahmood, Wei Emma Zhang, and Rongying Zhao. When large language models meet citation: A survey.arXiv preprint arXiv:2309.09727, 2023

  31. [31]

    ERDF A way of making Europe

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Zican Dong, Yupeng Hou, Beichen Zhang, Yingqian Min, Junjie Zhang, Peiyu Liu, Xiaolei Wang, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Yiwen Hu, Jian-Yun Nie, and Ji-Rong Wen. A survey of large language models.Frontiers of Computer Science, 2...

  32. [32]

    Given article

    **Focus Area Definition:** Thoroughly analyze the "Given article". Your analysis must concentrate *exclusively* on aspects relevant to the category "{category_name}". The definition for "{category_name}" is: "{category_description}"

  33. [33]

    1", "3",

    **Selected Reference Identifiers (Refs):** After detailing your CoT, clearly list the reference numbers (e.g., "1", "3", "12") you have selected as most critical and relevant for "{category_name}". These should be only the reference numbers/identifiers

  34. [34]

    {category_name}

    **Detailed Reasons for Each Selected Reference (Reasons):** For *each* reference number you listed in the step above, provide a detailed and specific explanation (e.g., 1-2 complete sentences per reference) stating *precisely why* that particular reference is critically important or highly relevant to "{category_name}" according to your analysis and the c...

  35. [35]

    A list of selected reference identifiers

  36. [36]

    refs", and

    A list of detailed reasons explaining the relevance of each selected reference to that category . Your task is to parse this input text and accurately convert it into a single JSON object with the keys "refs", and "reasons". ### Input Text (contains selected references and detailed reasons): {output_text_of_first_trial} ### Instructions for JSON Generation:

  37. [37]

    refs" List:** Identify the list of reference numbers under a header similar to

    **Extract "refs" List:** Identify the list of reference numbers under a header similar to "SELECTED_REFERENCES_FOR_...". Parse these numbers and format them as a list of strings in the JSON "refs" field. For example, if the text lists "1", "15", "42", the JSON field should be ‘["1", "15", "42"]‘. Ensure only the reference numbers/identifiers are included

  38. [38]

    reasons" List:** Identify the detailed reasons provided under a header similar to

    **Extract "reasons" List:** Identify the detailed reasons provided under a header similar to "DETAILED_REASONS_FOR_...". Each reason corresponds to a reference in the "refs" list. Parse these reasons and format them as a list of strings for the JSON "reasons" field. The order and number of reasons must strictly match the order and number of items in the "...