{"id":"055f7589-fc7d-47a6-8244-5bd41bafd15c","arxiv_id":"1908.03793","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Citation flows show Physics and Mathematics increasingly citing Computer Science, with Math-to-machine-learning citations growing rapidly after 2010.","lead":"This paper tracks how Physics, Mathematics, and Computer Science papers cite each other using over 1.2 million arXiv papers from 1990 to 2017. It finds that Physics and Mathematics increasingly cite Computer Science, especially machine learning subfields, in recent years.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The arXiv-only title-matched reference extraction confounds the observed cross-field citation rise with arXiv coverage growth; the 'exponentially increasing' claim is not separately validated.","rationale":"The reader's weakest assumption correctly identifies the arXiv-only title-matched citation network and the five-reference filter as the most fragile part of the argument. My stress-test converges on the same point, with additional emphasis on two specific consequences: the extremely small early CS sample (86 and 104 papers in B1 and B2) and the absence of any model or normalization supporting the phrase 'exponentially increasing.' These are not internal contradictions; they are threats to external validity that can be addressed by validation against a complete citation database. The paper's descriptive trends may well be real, but the current evidence does not rule out the arXiv-coverage artifact, so the CONDITIONAL verdict is appropriate. I do not see a reason to move to a harsher verdict because the concern is concrete, testable, and potentially remediable; conversely, it is load-bearing enough that unconditional acceptance would be premature.","tokens_in":4585,"tokens_out":4384,"duration_ms":49520,"concrete_test":"Take the arXiv papers in Table 2 and retrieve their full reference lists from an external source such as Microsoft Academic Graph or Scopus, including references to non-arXiv items. Recompute the B1–B5 MA→CS and PHY→CS citation fractions and the MA→machine-learning citation counts using the same bucket definitions. If the upward trend and the exponential increase survive with complete references, the arXiv-only match is not the source of the result; if the early-period values rise or the growth curve flattens, the headline claim is an artifact of arXiv coverage. As a secondary check, fit exponential and linear models to the MA→machine-learning citation rate normalized by the number of CS/machine-learning papers in each bucket to determine whether growth exceeds what field growth alone predicts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Physics and Mathematics citations to Computer Science have drastically increased, and that Mathematics-to-machine-learning citations are exponential, rests on a citation network built only from references that match a title string to an arXiv paper (§2). This instrument systematically omits citations to non-arXiv literature. Because arXiv coverage is strongly time- and field-dependent, the observed growth can be an artifact of deposit-rate changes rather than a genuine shift in intellectual exchange. Table 2 makes the risk concrete: only 86 CS papers survive the five-reference filter in B1 (1995–1999) and 104 in B2, while Table 1 lists 141,662 CS papers overall. If early CS results appeared in journals or proceedings not deposited on arXiv, early MA→CS and PHY→CS citations are invisible by construction, inflating the apparent increase. The five-reference filter compounds this: papers that cite mostly non-arXiv literature are dropped, so later periods, with better arXiv coverage, are selectively represented. Finally, the abstract's 'exponentially increasing' phrase is not backed by a fitted model or per-paper normalization; because the number of machine-learning papers itself grows rapidly, raw citation counts can rise exponentially even when per-paper citation probability is constant. The paper reports no external validation against a complete citation database, leaving the central descriptive result vulnerable to this coverage confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes citation interactions among Physics, Mathematics, and Computer Science using a corpus of more than 1.2 million arXiv papers from 1990–2017. The authors construct a citation network by parsing .bbl files and matching referenced titles to arXiv, retaining only papers with at least five extracted references. They then use Sankey diagrams, temporal bucket signatures, and subfield-level analyses to show that citations from Physics and Mathematics to Computer Science have 'drastically increased,' and that citations from Mathematics to the machine-learning subfield of Computer Science are 'exponentially increasing.' The paper also reports shifts in the popularity of specific subfields over time.","tokens_in":4946,"tokens_out":3027,"duration_ms":33190,"significance":"If the central descriptive claims are reliable, the paper documents a measurable and policy-relevant shift in the intellectual center of gravity across three major scientific fields, with learning-related subfields of Computer Science gaining attention from Mathematics and Physics. The analysis is transparent about its filtering rules and provides the raw dataset and supplementary material, which are strengths. However, the current evidence is insufficient to establish the headline quantitative claim of exponential growth, and the arXiv-only, title-matched reference extraction raises a real risk that the observed trends are artifacts of coverage change rather than genuine shifts in citation behavior.","major_comments":[{"comment":"The arXiv-only title-matched reference extraction with the five-reference filter creates a time- and field-dependent selection that directly confounds the central trend. Only 86 Computer Science papers survive the filter in B1 (1995–1999) and 104 in B2, even though Table 1 lists 141,662 CS papers overall. If early CS research was published in venues not deposited on arXiv or cited non-arXiv literature, those papers are systematically missing, making the later increase in PHY→CS and MA→CS citations appear steeper than it really is. The manuscript does not quantify arXiv coverage per field and time period, nor does it validate against a complete citation database (e.g., MAG or Scopus). This is load-bearing because the abstract's claim of a 'drastic' increase rests on these counts.","section":"§2 (Dataset construction) and Table 2"},{"comment":"The claim that citations from Mathematics to the machine-learning subfield of Computer Science are 'exponentially increasing' is not supported by any fitted model, growth-rate estimate, confidence interval, or goodness-of-fit test. A raw citation count can grow exponentially simply because the number of machine-learning papers itself grows rapidly, even if the per-paper citation propensity is constant. The paper should either fit a statistical model (e.g., Poisson or negative-binomial regression on per-paper citation rates) or substantially temper the claim to a descriptive statement about raw counts.","section":"Abstract and §3 (subfield analysis, Figure 3)"},{"comment":"The statement that during B1 'we do not observe in-/outflow of citations to/from CS' is based on extremely small sample sizes: only 86 CS papers, 726 MA papers, and 7553 PHY papers in the filtered set. Zero observed counts are weak evidence of absence of citation flow. The paper should report uncertainty intervals or model the count data (e.g., Poisson rates) to distinguish genuine absence of interaction from small-sample sampling noise, especially because the B1 zero is presented as the starting point of the temporal narrative.","section":"§3 (Bucket B1 observations)"}],"minor_comments":[{"comment":"The threshold of 'at least five extracted references' is introduced without justification; a sensitivity analysis varying this threshold would help assess how much the main trends depend on this arbitrary cutoff.","section":"§2"},{"comment":"The number '1,41,662' uses Indian digit grouping; for an international readership, the standard comma grouping '141,662' would be clearer.","section":"Table 1"},{"comment":"The caption 'Fraction of citation going from one field to another field over the years' should be 'Fraction of citations going from one field to another field over the years' (grammatical agreement).","section":"Figure 2 caption"},{"comment":"Reference [2] is incomplete: it lists 'Organisation for economic cooperation and development (oecd). (1998)' without a title or publication details.","section":"References"},{"comment":"The field abbreviation is inconsistently rendered as 'PH' in the Conclusion and 'PHY' elsewhere; unify to one abbreviation.","section":"§1 and §3"},{"comment":"The supplementary material is linked via a tinyurl; a permanent DOI or stable repository link would be more appropriate for archival purposes.","section":"Supplementary material"}],"recommendation":"major_revision","confidential_remarks":"The paper is largely descriptive and the central quantitative claim (exponential increase) is not backed by a fitted model. The coverage confound from the arXiv-only, title-matched extraction is serious and should be addressed with external validation or a detailed coverage analysis before publication. The authors should also consider whether a data-report format with clearly stated limitations is more appropriate than the current claim-heavy framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a decent descriptive study, not a breakthrough. It gives a concrete, subfield-level picture of how Physics and Math have increasingly cited CS over two decades, using ~1.2M arXiv papers. The Sankey diagrams and temporal bucket signatures are clear and the filtering rules are transparent. The finding that Math's citations to CS machine learning are growing fast is plausible, and the subfield-level detail is genuinely useful for people studying interdisciplinary dynamics.\n\nThe soft spots are real and not minor. The citation network is built by matching referenced titles to arXiv titles only, and the paper retains only papers with at least five extracted references. Table 2 shows just 86 CS papers in the first bucket (1995–1999) out of 141,662 overall. Early CS work that appeared in journals or proceedings without arXiv deposits is invisible, so the observed rise in cross-field citations could be inflated by arXiv coverage growth rather than a genuine shift in referencing behavior. The 'exponentially increasing' claim in the abstract is not backed by any fitted model or per-paper normalization; raw citation counts can grow exponentially simply because the number of ML papers is exploding. The paper does not validate against a complete citation database or release code and data, which makes these confounds harder to rule out.\n\nI want to be fair: the qualitative trend is probably real — CS has become more central — and the paper is honest about its procedures. It is not a fatal flaw that the early CS sample is tiny; that is an artifact of the discipline's history. The problem is that the magnitude and shape of the trend, especially the 'exponential' descriptor, outrun the measurement instrument. This is fixable: normalize by field size, compare against a broader citation source like Scopus or Dimensions for a validation subsample, and either fit a growth model or soften the language.\n\nRecommendation: send it to peer review. The topic matters, the descriptive results are worth discussing, and the limitations are addressable in revision. For my own work, I would not cite the exponential claim yet, but I would consider citing the subfield-level flows once the coverage checks are done.","headline":"A visually compelling descriptive map of cross-field citation growth, but the arXiv-only title-matched data and an unmodeled 'exponential' claim make the strong version of the result fragile.","tokens_in":5335,"tokens_out":2489,"would_cite":false,"duration_ms":27004,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Citations from Physics and Mathematics to Computer Science have grown sharply since the 1990s, with Mathematics' references to machine learning increasing exponentially.","keywords":["interdisciplinarity","citation analysis","scientometrics","Physics","Mathematics","Computer Science","machine learning","temporal bucket signatures"],"falsifier":"Compute the same citation flows from a complete bibliographic index covering journals and conference proceedings, and test whether the exponential rise in Mathematics-to-machine-learning citations survives once the growth of each field's output in the sample is accounted for.","tokens_in":4397,"feed_emoji":"📈","tokens_out":7216,"duration_ms":65460,"temperature":0.7,"pith_summary":"This paper analyzes citation flows among Physics, Mathematics, and Computer Science using more than 1.2 million papers from a preprint repository. It claims that over the past two decades, citations from the core sciences (Physics and Mathematics) to Computer Science have risen sharply, with Mathematics' citations to the machine-learning subfield growing exponentially in recent years. The authors also show that subfield popularity is unstable: some subfields such as Computation and Language fade while others such as Information Theory and Learning gain ground. This matters because it indicates the intellectual center of gravity among these fields is shifting toward Computer Science, and particularly toward learning-related areas. If the measurement is correct, the interaction between these disciplines is becoming less symmetric over time.","feed_headline":"Physics and math citations to computer science surge","feed_subtitle":"Study of 1.2 million papers finds core sciences turning to CS, especially machine learning.","key_machinery":"The analysis is carried by a citation network built from the papers' reference lists, where a reference counts only if its title string matches a paper in the same preprint repository. Two visualization tools carry the argument: Sankey diagrams, which draw citation flow between fields and subfields as link widths proportional to the number of citations, and temporal bucket signatures, stacked histograms showing the relative age of cited papers within fixed five-year time buckets. These tools let the authors separate self-field from non-self-field citations and track both the volume and direction of cross-field exchange over time.","core_discovery":"The paper's central empirical discovery is a directional shift in citation patterns between 1995 and 2017. In the earliest five-year bucket, Mathematics and Physics almost exclusively cite each other, with no citation flow to or from Computer Science. By the 2010s, Computer Science had become a net recipient of citations from both Mathematics and Physics, and in the final bucket (2015–2017) the two core fields cite Computer Science papers about equally. At the subfield level, the most striking pattern is Mathematics' growing citation flow to the Computer Science subfield 'Learning,' which the authors describe as exponentially increasing and attribute to the rise of machine learning and deep learning. The paper also documents that Physics and Mathematics tend to cite older work from each other but recent papers from Computer Science, consistent with the view of computer science as a fast-moving field.","pith_inferences":["The exponential growth in Mathematics-to-machine-learning citations may reflect mathematicians moving into the foundations of learning algorithms; testing whether the citing mathematics papers cluster in subfields like probability and analysis would sharpen this picture.","Applying the same method to other triples such as biology, chemistry, and computer science could show whether the pull toward computer science is general or peculiar to physics and mathematics.","The overall rise in cross-field citations might be concentrated in a few computer science subfields; re-analysis measuring how much of the growth comes from Learning and Information Theory would tell whether the headline is 'interdisciplinarity' or 'concentration in a few topics.'","Because the dataset ends in 2017, extending the analysis to recent years would test whether the exponential growth continued past the deep-learning boom or flattened out."],"forward_implications":["If the trend persists, Mathematics and Physics are likely to keep increasing their citation share to Computer Science, especially to machine-learning subfields.","Subfield popularity is volatile across time: a subfield that is highly cited in one period (such as Computation and Language) can decline sharply, while others (Information Theory, Learning) rise.","The age profiles imply that Computer Science is a fast-moving field: it cites recent work within its own field, while Mathematics and Physics cite older work from each other.","The measured asymmetry suggests that interdisciplinarity among these fields is not a balanced exchange but a growing one-way flow toward Computer Science."],"supporting_citations":[{"why":"Provides the working definition of multidisciplinary, interdisciplinary, and transdisciplinary research that frames the study.","marker":"[1]"},{"why":"OECD definition of research interaction classes that the paper uses to position its focus on interdisciplinary research.","marker":"[2]"},{"why":"Earlier evidence that science is becoming more interdisciplinary, which the paper extends to citation-flow dynamics among these three fields.","marker":"[3]"},{"why":"Introduces temporal bucket signatures (TBS), the visualization and measurement technique used to analyze the age distribution of cited papers.","marker":"[4]"}],"fun_headline_variants":["Math and physics now cite computer science far more often","Computer science becomes citation hub for math and physics","Exponential math citations to machine learning found","Core sciences shift citations to CS over two decades","Physics, math turn to computer science in citation shift"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the citation flows measured from a sample of papers in one preprint repository represent the true citation behavior of the three fields; if the repository's coverage, title-matching, or the five-reference filter distort the sample, the observed rise could be an artifact of changing repository usage rather than real intellectual exchange.","fun_headline_variants_meta":{"raw":{"variants":["Math and physics now cite computer science far more often","Computer science becomes citation hub for math and physics","Exponential math citations to machine learning found","Core sciences shift citations to CS over two decades","Physics, math turn to computer science in citation shift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1415,"prompt_tokens":867,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":476}},"tokens_in":483,"tokens_out":548,"duration_ms":6454,"temperature":1.0,"reasoning_tokens":476,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:00:50.847378+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the same citation flows from a complete bibliographic index covering journals and conference proceedings, and test whether the exponential rise in Mathematics-to-machine-learning citations survives once the growth of each field's output in the sample is accounted for.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the working definition of multidisciplinary, interdisciplinary, and transdisciplinary research that frames the study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"OECD definition of research interaction classes that the paper uses to position its focus on interdisciplinary research."},{"cited_title":"Scientometrics 81(3), 719 (Apr 2009)","cited_arxiv_id":null,"evidence_quote":"Earlier evidence that science is becoming more interdisciplinary, which the paper extends to citation-flow dynamics among these three fields."},{"cited_title":"In: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","cited_arxiv_id":null,"evidence_quote":"Introduces temporal bucket signatures (TBS), the visualization and measurement technique used to analyze the age distribution of cited papers."}],"review_version":1}