{"id":"7c0d2fa0-47f7-4533-841c-9343a2234cef","arxiv_id":"2601.17020","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new citation-purpose taxonomy applied to NLP+CSS papers finds that most out-of-discipline citations are shallow and that automated classification of citation purpose is not yet reliable.","lead":"This paper proposes a seven-category taxonomy for labeling why interdisciplinary papers cite work from other fields, and applies it to 369 hand-annotated citations in NLP+Computational Social Science papers. It finds that only 11% of citations reflect deep engagement, and that current language models cannot yet automate this labeling reliably.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Engagement-level findings depend on an unvalidated hand-set mapping of citation purpose and section to High/Medium/Low; a sensitivity analysis is needed.","rationale":"The reader's weakest assumption correctly identifies the engagement mapping in Tables 3–4 as the most load-bearing step in the argument. The entire quantitative contribution — the 11% high vs. 47% low split — depends on that hand-set mapping, which is presented without any validation or independent basis. My reading confirms this is the central soft spot: the taxonomy itself has some inter-annotator reliability (α=0.6), but the engagement levels are not independently verified, and the paper's own three-component proposal (purpose + section + relatedness) is only partially operationalized. The missing dataset link is also a serious reproducibility concern, but it does not affect the internal correctness of the mapping as much as the mapping's validity does. A sensitivity analysis would directly test whether the headline conclusions are robust to reasonable changes in the mapping; without that, the findings could be an artifact of the authors' choices rather than a property of NLP+CSS literature. No additional concerns warrant changing the reader's CONDITIONAL verdict, so the recommendation stands unchanged.","tokens_in":14193,"tokens_out":3783,"duration_ms":40553,"concrete_test":"Perform a sensitivity analysis on the 369 annotated citations: re-assign engagement levels under plausible alternative mappings (e.g., promote 'Substantiation' to High, demote 'Definition' to Low, treat Introduction and Method sections as Medium rather than High, or move 'Analysis' to Medium), and recompute the low/high percentages. If the headline 11%/47% split changes materially (e.g., high engagement rises above 25% or low engagement falls below 30%), the finding is not robust to the mapping. Additionally, have an independent expert rate a random sample of 50 citations for engagement depth on a 1–3 scale without seeing Tables 3–4, and compare to the mapping's predictions; a Cohen's κ below 0.4 would validate the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central finding in Sec. 5.1 — that only 11% of citations show high engagement while 47% show low — is not a direct property of the annotated data but a consequence of the authors' a priori engagement mapping in Tables 3–4. That mapping fixes each citation purpose and section to a single engagement level with no external validation. For instance, 'Related Work' is defined in Table 1 as a catch-all for citations not fitting other categories, and Table 4 labels it Low; 'Analysis' is also Low, while 'Substantiation' is Medium even though it supports a claim. If a 'Related Work' citation signals conceptual influence or an 'Analysis' citation signals critical engagement, the 47% figure shifts. The paper reports inter-annotator reliability only for the purpose labels (α=0.6), not for the engagement levels themselves. The SPECTER relatedness component proposed in Sec. 4.3 is computed but never incorporated into the headline engagement counts, so the three-component model remains untested. The 11%/47% split is therefore an artifact of the mapping until shown otherwise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a seven-category citation-purpose taxonomy to study how NLP+CSS publications engage with out-of-discipline references. The authors annotate 394 in-text citations (369 with full/partial agreement), assign each purpose and section a fixed engagement level (High/Medium/Low) in Tables 3–4, and report that low-engagement purposes (Related Work + Analysis) account for 47% of citations while high-engagement purposes (Substantiation+Basis, Basis) account for 11%. They further evaluate two automatic classification approaches, finding that neither achieves sufficient performance. The paper's stated contributions are the taxonomy, the annotated dataset, the engagement analysis, and the automated classification results.","tokens_in":14494,"tokens_out":5465,"duration_ms":52522,"significance":"If the engagement-level construct is validated, the framework would add a qualitative dimension to bibliometric interdisciplinarity measures, addressing a recognized gap in prior citation-purpose taxonomies. The paper's strengths include a reproducible annotated dataset with a detailed codebook, explicit inter-annotator reliability reporting (α=0.6, 64% complete agreement), and an honest assessment of the limits of current LLM-based classification (including a zero F1 for the most theoretically important category, Substantiation+Basis). The main weakness is that the central engagement-level findings are derived directly from a hand-set mapping that has not been validated externally, making the empirical claims about the field's engagement depth currently unsubstantiated.","major_comments":[{"comment":"The headline results — low engagement (Related Work + Analysis) = 47%, high engagement = 11% — are direct arithmetic consequences of the hand-set engagement mapping in Tables 3 and 4, which assign a fixed High/Medium/Low level to each purpose and section. No external validation is provided (e.g., expert engagement ratings, sensitivity analysis, or outcome-based checks). Consequently, the claim that 'the majority of citations ... indicating only surface-level engagement' restates the authors' codebook assumptions rather than an empirical property of the NLP+CSS literature. Please either reframe these quantities as definitional to the proposed framework, or validate the mapping by comparing it to independent judgments of engagement depth (and report sensitivity to plausible alternative mappings, e.g., moving 'Related Work' or 'Analysis' to Medium).","section":"§5.1, Tables 3–4"},{"comment":"Section 4.3 defines engagement as composed of purpose, section, and SPECTER-based semantic relatedness, but the relatedness scores are never incorporated into the engagement levels used in Section 5.1. Figure 4 shows only that SPECTER similarity varies little across purposes; it does not test the three-component model. The claim that 'engagement cannot be reliably inferred from semantic similarity alone' is not actually demonstrated, since the three-component model is never fit or compared against a two-component model. Either integrate relatedness into the engagement computation or clearly restrict the reported analysis to the two-component (purpose + section) model.","section":"§4.3 / §5.1"},{"comment":"The 'Related Work' category is defined in Table 1 as a residual catch-all ('does not fit any of the other categories'). It is the largest category (129/369 = 35%) and is labeled Low engagement in Table 4, so the 47% low-engagement finding is highly sensitive to how this residual class is interpreted. The paper reports only α=0.6 for the purpose labels; it should also report how often annotators chose 'Related Work' as a fallback and whether disagreement cases (Figure 3) concentrate in this category. Without this information, the low-engagement share may be an artifact of annotator uncertainty channeled into the residual category rather than evidence of surface-level engagement.","section":"§4.2, Table 1"}],"minor_comments":[{"comment":"The sentence 'We identified 26,289 of the in-text citations as out-of-discipline, for an average of out-of-discipline citations per article' is missing the average value. Also, the preceding sentence mentions '11,158 interdisciplinary' citations; clarify whether this term is synonymous with 'out-of-discipline' or a separate notion.","section":"§3.2"},{"comment":"Please define 'partial agreement' (reported as 93.91%). Without a definition, readers cannot interpret the reliability of the agreed-upon dataset.","section":"§4.2"},{"comment":"The captions say 'Correlation between citation purpose and context relatedness' and 'Correlation between citation section and purpose,' but Figure 4 is a distribution/box plot and Figure 5 shows chi-square residuals. The captions should be reworded to describe the actual visualization.","section":"Figures 4 and 5"},{"comment":"Typo: 'he citation refers to work' should be 'The citation refers to work.'","section":"Appendix A.3.7"},{"comment":"The row for 'accuracy' is misformatted ('accuracy0.327 0.327 0.327 0'); fix the spacing and column alignment.","section":"Table 8"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is the annotation framework and dataset, which are reproducible and useful. However, the central empirical claim (11% high / 47% low engagement) is essentially a restatement of the authors' hand-set mapping. This is fixable within the manuscript's scope: either reframe the finding as a framework-relative illustration or add a validation/sensitivity study. Given the importance of the issue to the paper's stated contribution, I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's real contribution is the seven-category citation-purpose taxonomy for interdisciplinary NLP+CSS work, plus a genuinely annotated dataset of 369 citations. That part is solid: multiple annotators, a codebook in the appendix, and honest reliability reporting (α=0.6, 64% complete agreement). The authors also deserve credit for testing automated classification and reporting that it doesn't work well — not everyone is that candid.\n\nThe soft spot is the engagement construct. The headline 'only 11% high engagement, 47% low' is not measured from data; it's a direct consequence of Tables 3–4, where the authors assign each purpose and section a fixed High/Medium/Low label. For example, Related Work is defined as a catch-all and then labeled Low; Analysis is also Low, despite the fact that critical comparison is arguably a form of engagement. If those mappings were different, the split would change. The authors never validate the mapping against any external measure of interdisciplinary integration, and the SPECTER relatedness scores they compute in Sec. 4.3 are never plugged into the headline counts. So the central finding is partially circular.\n\nMinor issues: the dataset link is missing (the text says 'here' with no URL), and the corpus is small (369 citations from 170 articles), so the distributional claims are estimates. The qualitative examples in Table 6 are useful, but they only illustrate the authors' own mapping.\n\nOverall, the taxonomy and dataset are worth having, likely for bibliometrics and future annotation studies. The engagement-level claims need external validation or at least a sensitivity analysis before they should be taken as findings about the NLP+CSS literature. This is not a fatal flaw, but it does mean the paper's main advertised result is not yet supported.\n\nWho is this for? Researchers studying citation purposes, interdisciplinarity metrics, or NLP+CSS research quality. It deserves a serious referee — I'd send it out rather than desk-reject, and ask for validation of the mapping plus release of the dataset. My guess is it can be revised into a solid contribution.","headline":"The taxonomy and annotation dataset are real contributions, but the headline engagement split is an artifact of the authors' own unvalidated mapping, so treat the main claim as conditional.","tokens_in":14974,"tokens_out":1909,"would_cite":true,"duration_ms":20341,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A seven-category citation-purpose taxonomy, applied to 369 citations from NLP and computational social science papers, finds that only 11% of out-of-discipline references reflect deep engagement.","keywords":["citation purpose","interdisciplinary research","computational social science","natural language processing","engagement quality","annotation study","citation analysis","scientific discourse"],"falsifier":"A concrete check: have independent experts rate the engagement of the same 369 citations without using the codebook, then compare their engagement scores to the taxonomy's labels. If expert-assessed engagement for Related Work citations is often high, or if the 11% high-engagement figure changes substantially under a different purpose-to-engagement mapping, the framework's central quantitative claim would be undermined.","tokens_in":14100,"feed_emoji":"🔗","tokens_out":2836,"duration_ms":28915,"temperature":0.7,"pith_summary":"This paper argues that measuring interdisciplinarity requires knowing how each out-of-discipline citation is used, not just how many there are. It introduces a seven-category taxonomy of citation purpose, from deep foundational use to surface-level related-work mentions, and applies it to 369 citations from NLP and computational social science papers. The central result is that high-engagement purposes account for only 11% of citations, while low-engagement purposes (Related Work and Analysis) make up 47%. The paper also shows that citation purpose correlates with the section a citation appears in, and that current language models cannot yet reliably automate this classification.","feed_headline":"Only 11% of cross-disciplinary citations show deep engagement","feed_subtitle":"A new taxonomy of citation purpose reveals that most out-of-discipline references in NLP+CSS papers are surface-level.","key_machinery":"The central object is a seven-category citation-purpose taxonomy (Substantiation+Basis, Basis, Substantiation, Use, Definition, Analysis, Related Work), developed through inductive annotation of interdisciplinary NLP papers. Each category is assigned an engagement level (High/Medium/Low) in a hand-set mapping, and section locations are also mapped to engagement levels. The taxonomy works by situating each citation in its surrounding paragraph and the paper's abstract claims, allowing the annotator to distinguish surface-level mentions from citations that ground the paper's methods or arguments.","core_discovery":"The paper's central claim is that a citation-purpose taxonomy specifically designed for interdisciplinary contexts can capture the role out-of-discipline references play in shaping a paper's conceptual, methodological, and empirical claims. Applying the taxonomy to 369 agreed-upon citations from NLP+CSS publications, the authors find that deep engagement is rare: Substantiation+Basis and Basis, the two 'high engagement' purposes, together account for only 11% of citations, while Related Work and Analysis, labeled low engagement, account for 47%. They further find a statistically significant association between paper section and citation purpose, arguing that where a citation appears is predi","pith_inferences":["The hand-set engagement mapping makes the 11%/47% split a conclusion of the codebook; testing with alternative mappings or independent expert judgment could change the headline numbers.","A testable extension is to apply the taxonomy to other interdisciplinary pairs (e.g., biology + computer science) to see whether low engagement rates generalize beyond NLP+CSS.","The framework implies that institutional incentives for interdisciplinarity may be rewarding surface-level citation practices; one could test whether papers with more high-engagement citations are themselves more influential.","The authors' finding that semantic similarity is not predictive of purpose suggests that deeper discourse parsing is needed, pointing to a research program rather than a finished measurement tool."],"forward_implications":["Provides a quantitative method to assess the quality of interdisciplinary integration, moving beyond citation-count diversity metrics.","The finding that high engagement is rare in NLP+CSS offers empirical support for critiques of shallow interdisciplinarity in this field.","The section-purpose correlation offers a cheap proxy signal: citations in Method and Introduction sections are more likely to show deep engagement than those in Related Work sections.","The framework can be applied to compare engagement across venues, publication years, or different disciplinary pairs.","The low performance of automatic classifiers identifies a concrete gap for future NLP research on citation context modeling."],"fun_headline_variants":["Only 11% of cross-disciplinary citations engage deeply","New taxonomy reveals sparse deep citation engagement","Most out-of-discipline citations stay surface-level","Deep engagement rare in NLP+CSS cross-citations","Citation purpose taxonomy shows shallow cross-field links"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is the authors' hand-assigned mapping of each citation purpose and paper section to a fixed engagement level; if that mapping is wrong, the headline 11% and 47% figures are artifacts of the codebook rather than properties of the literature.","fun_headline_variants_meta":{"raw":{"variants":["Only 11% of cross-disciplinary citations engage deeply","New taxonomy reveals sparse deep citation engagement","Most out-of-discipline citations stay surface-level","Deep engagement rare in NLP+CSS cross-citations","Citation purpose taxonomy shows shallow cross-field links"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000124,"raw_usage":{"total_tokens":901,"prompt_tokens":664,"completion_tokens":237,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":167}},"tokens_in":408,"tokens_out":237,"duration_ms":3062,"temperature":1.0,"reasoning_tokens":167,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:09:41.951591+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: have independent experts rate the engagement of the same 369 citations without using the codebook, then compare their engagement scores to the taxonomy's labels. If expert-assessed engagement for Related Work citations is often high, or if the 11% high-engagement figure changes substantially under a different purpose-to-engagement mapping, the framework's central quantitative claim would be undermined.","supporting_citations":[],"review_version":1}