Pith. sign in

REVIEW 3 major objections 5 minor 22 references

A seven-category citation-purpose taxonomy, applied to 369 citations from NLP and computational social science papers, finds that only 11% of out-of-discipline references reflect deep engagement.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 10:09 UTC pith:PTZJFDZ3

load-bearing objection The taxonomy and annotation dataset are real contributions, but the headline engagement split is an artifact of the authors' own unvalidated mapping, so treat the main claim as conditional. the 3 major comments →

arxiv 2601.17020 v2 pith:PTZJFDZ3 submitted 2026-01-15 cs.DL cs.CL

How Do We Engage with Other Disciplines? A Framework to Study Meaningful Interdisciplinary Discourse in Scholarly Publications

classification cs.DL cs.CL
keywords citation purposeinterdisciplinary researchcomputational social sciencenatural language processingengagement qualityannotation studycitation analysisscientific discourse
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that measuring interdisciplinarity requires knowing how each out-of-discipline citation is used, not just how many there are. It introduces a seven-category taxonomy of citation purpose, from deep foundational use to surface-level related-work mentions, and applies it to 369 citations from NLP and computational social science papers. The central result is that high-engagement purposes account for only 11% of citations, while low-engagement purposes (Related Work and Analysis) make up 47%. The paper also shows that citation purpose correlates with the section a citation appears in, and that current language models cannot yet reliably automate this classification.

Core claim

The paper's central claim is that a citation-purpose taxonomy specifically designed for interdisciplinary contexts can capture the role out-of-discipline references play in shaping a paper's conceptual, methodological, and empirical claims. Applying the taxonomy to 369 agreed-upon citations from NLP+CSS publications, the authors find that deep engagement is rare: Substantiation+Basis and Basis, the two 'high engagement' purposes, together account for only 11% of citations, while Related Work and Analysis, labeled low engagement, account for 47%. They further find a statistically significant association between paper section and citation purpose, arguing that where a citation appears is predi

What carries the argument

The central object is a seven-category citation-purpose taxonomy (Substantiation+Basis, Basis, Substantiation, Use, Definition, Analysis, Related Work), developed through inductive annotation of interdisciplinary NLP papers. Each category is assigned an engagement level (High/Medium/Low) in a hand-set mapping, and section locations are also mapped to engagement levels. The taxonomy works by situating each citation in its surrounding paragraph and the paper's abstract claims, allowing the annotator to distinguish surface-level mentions from citations that ground the paper's methods or arguments.

Load-bearing premise

The load-bearing premise is the authors' hand-assigned mapping of each citation purpose and paper section to a fixed engagement level; if that mapping is wrong, the headline 11% and 47% figures are artifacts of the codebook rather than properties of the literature.

What would settle it

A concrete check: have independent experts rate the engagement of the same 369 citations without using the codebook, then compare their engagement scores to the taxonomy's labels. If expert-assessed engagement for Related Work citations is often high, or if the 11% high-engagement figure changes substantially under a different purpose-to-engagement mapping, the framework's central quantitative claim would be undermined.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Provides a quantitative method to assess the quality of interdisciplinary integration, moving beyond citation-count diversity metrics.
  • The finding that high engagement is rare in NLP+CSS offers empirical support for critiques of shallow interdisciplinarity in this field.
  • The section-purpose correlation offers a cheap proxy signal: citations in Method and Introduction sections are more likely to show deep engagement than those in Related Work sections.
  • The framework can be applied to compare engagement across venues, publication years, or different disciplinary pairs.
  • The low performance of automatic classifiers identifies a concrete gap for future NLP research on citation context modeling.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The hand-set engagement mapping makes the 11%/47% split a conclusion of the codebook; testing with alternative mappings or independent expert judgment could change the headline numbers.
  • A testable extension is to apply the taxonomy to other interdisciplinary pairs (e.g., biology + computer science) to see whether low engagement rates generalize beyond NLP+CSS.
  • The framework implies that institutional incentives for interdisciplinarity may be rewarding surface-level citation practices; one could test whether papers with more high-engagement citations are themselves more influential.
  • The authors' finding that semantic similarity is not predictive of purpose suggests that deeper discourse parsing is needed, pointing to a research program rather than a finished measurement tool.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a seven-category citation-purpose taxonomy to study how NLP+CSS publications engage with out-of-discipline references. The authors annotate 394 in-text citations (369 with full/partial agreement), assign each purpose and section a fixed engagement level (High/Medium/Low) in Tables 3–4, and report that low-engagement purposes (Related Work + Analysis) account for 47% of citations while high-engagement purposes (Substantiation+Basis, Basis) account for 11%. They further evaluate two automatic classification approaches, finding that neither achieves sufficient performance. The paper's stated contributions are the taxonomy, the annotated dataset, the engagement analysis, and the automated classification results.

Significance. If the engagement-level construct is validated, the framework would add a qualitative dimension to bibliometric interdisciplinarity measures, addressing a recognized gap in prior citation-purpose taxonomies. The paper's strengths include a reproducible annotated dataset with a detailed codebook, explicit inter-annotator reliability reporting (α=0.6, 64% complete agreement), and an honest assessment of the limits of current LLM-based classification (including a zero F1 for the most theoretically important category, Substantiation+Basis). The main weakness is that the central engagement-level findings are derived directly from a hand-set mapping that has not been validated externally, making the empirical claims about the field's engagement depth currently unsubstantiated.

major comments (3)
  1. [§5.1, Tables 3–4] The headline results — low engagement (Related Work + Analysis) = 47%, high engagement = 11% — are direct arithmetic consequences of the hand-set engagement mapping in Tables 3 and 4, which assign a fixed High/Medium/Low level to each purpose and section. No external validation is provided (e.g., expert engagement ratings, sensitivity analysis, or outcome-based checks). Consequently, the claim that 'the majority of citations ... indicating only surface-level engagement' restates the authors' codebook assumptions rather than an empirical property of the NLP+CSS literature. Please either reframe these quantities as definitional to the proposed framework, or validate the mapping by comparing it to independent judgments of engagement depth (and report sensitivity to plausible alternative mappings, e.g., moving 'Related Work' or 'Analysis' to Medium).
  2. [§4.3 / §5.1] Section 4.3 defines engagement as composed of purpose, section, and SPECTER-based semantic relatedness, but the relatedness scores are never incorporated into the engagement levels used in Section 5.1. Figure 4 shows only that SPECTER similarity varies little across purposes; it does not test the three-component model. The claim that 'engagement cannot be reliably inferred from semantic similarity alone' is not actually demonstrated, since the three-component model is never fit or compared against a two-component model. Either integrate relatedness into the engagement computation or clearly restrict the reported analysis to the two-component (purpose + section) model.
  3. [§4.2, Table 1] The 'Related Work' category is defined in Table 1 as a residual catch-all ('does not fit any of the other categories'). It is the largest category (129/369 = 35%) and is labeled Low engagement in Table 4, so the 47% low-engagement finding is highly sensitive to how this residual class is interpreted. The paper reports only α=0.6 for the purpose labels; it should also report how often annotators chose 'Related Work' as a fallback and whether disagreement cases (Figure 3) concentrate in this category. Without this information, the low-engagement share may be an artifact of annotator uncertainty channeled into the residual category rather than evidence of surface-level engagement.
minor comments (5)
  1. [§3.2] The sentence 'We identified 26,289 of the in-text citations as out-of-discipline, for an average of out-of-discipline citations per article' is missing the average value. Also, the preceding sentence mentions '11,158 interdisciplinary' citations; clarify whether this term is synonymous with 'out-of-discipline' or a separate notion.
  2. [§4.2] Please define 'partial agreement' (reported as 93.91%). Without a definition, readers cannot interpret the reliability of the agreed-upon dataset.
  3. [Figures 4 and 5] The captions say 'Correlation between citation purpose and context relatedness' and 'Correlation between citation section and purpose,' but Figure 4 is a distribution/box plot and Figure 5 shows chi-square residuals. The captions should be reworded to describe the actual visualization.
  4. [Appendix A.3.7] Typo: 'he citation refers to work' should be 'The citation refers to work.'
  5. [Table 8] The row for 'accuracy' is misformatted ('accuracy0.327 0.327 0.327 0'); fix the spacing and column alignment.

Circularity Check

1 steps flagged

Headline engagement findings are defined by the hand-set Tables 3–4 mapping: 'high = 11%, low = 47%' restates the authors' own purpose-to-engagement assignment.

specific steps
  1. self definitional [Sec. 4.3 (Tables 3-4) and Sec. 5.1]
    "Tables 3 and 4 contain the citation sections and purposes, with a qualitative estimate of the depth of engagement that they capture. ... Analysis and Related Work together, indicating low engagement, make up 47% of the analyzed citations. Conversely, purposes indicating high engagement make up only 11% of the citations."

    Engagement level is defined by the authors in Table 4: Related Work and Analysis are labeled Low, while Basis and Substantiation+Basis are labeled High. The Sec. 5.1 percentages are therefore not an independent empirical finding about the literature; they are the annotated purpose distribution rescaled through this hand-set mapping. The concluding claim that 'the purpose can be predictive of the level of engagement' is tautological, since the level is assigned from the purpose. No inter-annotator reliability or external validation is reported for the engagement levels themselves (α=0.6 is for purpose labels only), and the SPECTER relatedness component is computed but never incorporated into the headline engagement counts.

full rationale

The paper builds a genuinely useful annotation dataset and tests automatic classifiers, so much of the contribution is independent and not circular. However, the central quantitative result about interdisciplinary engagement—only 11% high engagement and 47% low engagement—is not a measured property of the citations but a direct consequence of the authors' a priori Tables 3-4 mapping of citation purposes and sections to engagement levels. The counts are real annotations, but the 'High/Medium/Low' labels are definitions, not discoveries. The statement that purpose is 'predictive' of engagement reduces to the codebook, because each purpose is assigned a single engagement level in Table 4. The section-level mapping in Table 3 has the same issue: labeling Introduction 'High' and Related Work 'Low' and then reporting that most citations are in Related Work is a restatement of the mapping. The paper also relies on Leto et al. (2024), prior work by two of the same authors, for the underlying data collection, but this is a normal methodological building block rather than a circular justificatory chain; the citation-purpose taxonomy itself is adapted from external work (Abu-Jbara et al., 2013) and tested with annotation agreement. The circularity is therefore partial and localized: the framework's headline engagement split is self-definitional, while the annotated dataset, reliability measure, and classifier experiments retain independent content.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 2 invented entities

The central results do not depend on fitted numeric parameters, but they depend on hand-chosen categorical assumptions: the field-of-study cutoff, the citation-context window, the purpose-to-engagement mapping, and the completeness of the taxonomy. These are logged as axioms/invented constructs because they are load-bearing for the 11%/47% findings and are not independently validated.

axioms (5)
  • domain assumption Citation purpose can be inferred from the paragraph surrounding the citation plus the paper's abstract.
    Sec. 4.1/4.2: annotations use only this context, without evidence that this context is sufficient to recover authorial intent.
  • domain assumption Engagement depth is a function of citation purpose, paper section, and SPECTER relatedness, with fixed High/Medium/Low assignments.
    Sec. 4.3 and Tables 3-4: the mapping is author-chosen and not validated against an external measure of interdisciplinary integration.
  • domain assumption Semantic Scholar field-of-study labels accurately distinguish in-discipline from out-of-discipline citations.
    Sec. 3.2: citations outside 'Computer Science' and 'Linguistics' are treated as out-of-discipline; tagging errors propagate into every out-of-discipline statistic.
  • domain assumption The seven taxonomy categories are exhaustive and mutually exclusive.
    The Related Work category is defined as a catch-all ('does not fit any of the other categories'), which requires the taxonomy to be complete; the codebook provides no formal proof of exhaustiveness.
  • domain assumption The RoBERTa track classifier trained on 3,189 abstracts identifies NLP+CSS papers sufficiently well to define the corpus.
    Sec. 3.1/A.1: macro-F1 0.87 on the test split; classification errors in the 2015-2025 corpus are not modeled and may bias the sample of articles.
invented entities (2)
  • Seven-category citation-purpose taxonomy (Substantiation+Basis, Basis, Substantiation, Use, Definition, Analysis, Related Work) no independent evidence
    purpose: Classify the argumentative role of out-of-discipline citations.
    Categories are author-defined and author-annotated; no external benchmark or independent replication confirms the categories map to real differences in interdisciplinary engagement.
  • Engagement level construct (High/Medium/Low) assigned to purposes and sections no independent evidence
    purpose: Quantify depth of interdisciplinary engagement.
    Tables 3-4 set the mapping by author judgment; the paper provides qualitative examples but no validation against an outcome or expert-judgment gold standard.

pith-pipeline@v1.3.0-alltime-deepseek · 13911 in / 11460 out tokens · 117143 ms · 2026-08-03T10:09:41.951591+00:00 · methodology

0 comments
read the original abstract

With the rising popularity of interdisciplinary work and increasing institutional incentives in this direction, there is a growing need to understand how resulting publications incorporate ideas from multiple disciplines. Existing computational approaches, such as affiliation diversity, keywords, and citation patterns, do not account for how individual citations are used to advance the citing work. Although, in line with addressing this gap, prior studies have proposed taxonomies to classify citation purpose, these frameworks are not well-suited to interdisciplinary research and do not provide quantitative measures of citation engagement quality. To address these limitations, we propose a framework for the evaluation of citation engagement in interdisciplinary Natural Language Processing (NLP) publications. Our approach introduces a citation purpose taxonomy tailored to interdisciplinary work, supported by an annotation study. We demonstrate the utility of this framework through a thorough analysis of publications at the intersection of NLP and Computational Social Science.

Figures

Figures reproduced from arXiv: 2601.17020 by Alexandria Leto, Bagyasree Sudharsan, Maria Leonor Pacheco.

Figure 1
Figure 1. Figure 1: Framework for evaluating a publication’s level [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Dataset of NLP+CSS articles across years and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Types of agreements and disagreements among annotators. Appendix A.3. We selected about 170 articles balanced across publication year and venue. We selected 394 in￾text citations from this set. 231 of the 394 exam￾ples were doubly-annotated, and 163 were triply￾annotated. We had complete agreement on 64.21% of the citations and partial agreement on 93.91%. We obtained a Krippendorf’s α of 0.6. The result￾i… view at source ↗
Figure 5
Figure 5. Figure 5: Correlation between citation section and pur [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 4
Figure 4. Figure 4: Correlation between citation purpose and con [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Labeled dataset for training track classifier [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Distribution of citation purpose by article year. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Distribution of citation purpose by article [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 11
Figure 11. Figure 11: System prompt used to generate synthetic [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 9
Figure 9. Figure 9: System prompt used to prompt ChatGPT 5.2 [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Decision tree used to prompt ChatGPT 5.2 [PITH_FULL_IMAGE:figures/full_fig_p015_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 5 canonical work pages

  1. [1]

    Amjad Abu-Jbara, Jefferson Ezra, and Dragomir Radev. 2013. https://aclanthology.org/N13-1067/ Purpose and polarity of citation: Towards NLP -based bibliometrics . In Proceedings of the 2013 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 596--606, Atlanta, Georgia

  2. [2]

    Christian Baden, Christian Pipal, Martijn Schoonvelde, and Mariken A. C. G van der Velden. 2022. https://doi.org/10.1080/19312458.2021.2015574 Three Gaps in Computational Text Analysis Methods for Social Sciences : A Research Agenda . Communication Methods and Measures, 16(1):1--18

  3. [3]

    Lorenzo Cassi, Raphaël Champeimont, Wilfriedo Mescheba, and Élisabeth de Turckheim. 2017. https://doi.org/doi:10.1371/journal.pone.0170296 Analysing institutions interdisciplinarity by extensive use of rao-stirling diversity index . PloS

  4. [4]

    Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019. https://doi.org/10.18653/v1/N19-1361 Structural scaffolds for citation intent classification in scientific publications . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...

  5. [5]

    Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. SPECTER: Document-level Representation Learning using Citation-informed Transformers . In ACL

  6. [6]

    Justin Grimmer and Brandon M. Stewart. 2013. https://doi.org/10.1093/pan/mps028 Text as data: The promise and pitfalls of automatic content analysis methods for political texts . Political Analysis, 21(3):267–297

  7. [7]

    Suchin Gururangan, Ana Marasovi \'c , Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.acl-main.740 Don ' t stop pretraining: Adapt language models to domains and tasks . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342--8360, Online

  8. [8]

    David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018. https://doi.org/10.1162/tacl_a_00028 Measuring the evolution of a scientific field through citation frames . Transactions of the Association for Computational Linguistics, 6:391--406

  9. [9]

    Graham, F.Q

    Rodney Michael Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Gorney, David W. Graham, F.Q. Hu, Regan Huff, Daniel King, Sebastian Kohlmeier, Ba...

  10. [10]

    Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, and Kyle Lo. 2022. https://doi.org/10.18653/v1/2022.naacl-main.137 M ulti C ite: Modeling realistic citations requires moving beyond the single-sentence single-label setting . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Compu...

  11. [11]

    Alexandria Leto, Shamik Roy, Alexander Hoyle, Daniel Acuna, and Maria Leonor Pacheco. 2024. https://doi.org/10.18653/v1/2024.nlpcss-1.11 A first step towards measuring interdisciplinary engagement in scientific publications: A case study on NLP + CSS research . In Proceedings of the Sixth Workshop on Natural Language Processing and Computational Social Sc...

  12. [12]

    McCarthy and Giovanna Maria Dora Dore

    Arya D. McCarthy and Giovanna Maria Dora Dore. 2023. https://doi.org/10.18653/v1/2023.acl-short.136 Theory-grounded computational text analysis . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1586--1594, Toronto, Canada

  13. [13]

    National Academy of Sciences , National Academy of Engineering , and Institute of Medicine . 2005. https://doi.org/10.17226/11153 Facilitating Interdisciplinary Research . The National Academies Press, Washington, DC

  14. [14]

    OpenAI . 2025. https://chat.openai.com/chat ChatGPT 5.2 [large language model]

  15. [15]

    Porter and Ismael Rafols

    Alan L. Porter and Ismael Rafols. 2009. https://doi.org/10.1007/s11192-008-2197-2 Is science becoming more interdisciplinary? Measuring and mapping six research fields over time . Scientometrics, 81(3):719--745

  16. [16]

    David Pride and Petr Knoth. 2020. https://doi.org/10.1145/3383583.3398617 An authoritative approach to citation classification . In Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, JCDL '20, page 337–340, New York, NY, USA. Association for Computing Machinery

  17. [17]

    Zhao Qun and Yang Menghui. 2023. https://doi.org/10.1016/j.joi.2023.101425 An efficient entropy of sum approach for measuring diversity and interdisciplinarity . Journal of Informetrics, 17(3):101425

  18. [18]

    Simone Teufel, Advaith Siddharthan, and Dan Tidhar. 2006. https://aclanthology.org/W06-1613/ Automatic classification of citation function . In Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, pages 103--110, Sydney, Australia

  19. [19]

    Richard Van Noorden. 2015. https://doi.org/10.1038/525306a Interdisciplinary research by the numbers . Nature, 525:306--7

  20. [20]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. https://arxiv.org/abs/19...

  21. [21]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  22. [22]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...