{"id":"6185a4be-ee36-4913-b5c2-4f68363d8494","arxiv_id":"2607.16780","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Citing AI literature is tied to higher citation impact overall, yet field, career stage, and institutional AI capability redistribute that benefit, with intermediate-capability institutions gaining the most per unit of AI use.","lead":"Analyzing 12.8 million scientific papers, this study finds that citing AI research is linked to more citations, but the benefit is uneven: it depends on the field, the author's career stage, and the institution's AI strength, with mid-tier AI universities gaining the most. The paper argues that what matters is 'translational capacity' — making AI knowledge useful to non-AI audiences — not AI capability alone.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI-integration measure is not validated: OpenAlex AI subfield includes non-AI topics, so the conditional-return contrasts may be artifacts of taxonomy contamination.","rationale":"I read the paper as an observational bibliometric study whose central contribution is a conditional association: returns to AI knowledge integration vary by field, career stage, and institution. The most load-bearing assumption is that the OpenAlex AI subfield cleanly operationalizes 'AI knowledge.' The reader identified this same concern. It is load-bearing because every hypothesis test — H1a–c, H2, H3a–c, H4 — uses AIRef or AI score; if these variables contain non-AI citations, no estimated contrast can be attributed to AI. The paper's own limitation statement concedes taxonomy looseness, and no robustness check changes the classification. I do not think this invalidates the descriptive findings entirely; the correlations are real for the constructed measure. But it means the central theoretical interpretation ('translational capacity') is not established until the first-stage measure is validated. The paper is otherwise careful: it frames results as associations, reports clustered standard errors, and provides multiple robustness checks. Yet those checks do not address the foundational measurement issue. Verdict remains CONDITIONAL: the conditional-return claim is plausible but depends on a taxonomy that needs external validation.","tokens_in":21233,"tokens_out":4598,"duration_ms":47946,"concrete_test":"Compute, over the 12.8M focal papers, the distribution of the 77 AI topics among cited references. Identify the subset of 'AI-referencing' papers whose only flagged reference comes from clearly non-AI topic labels (e.g., geochemistry, seismology, solar radiation, psychiatry) or use an independent annotation of 200 randomly sampled flagged references. Then re-estimate Table 3: (a) excluding such papers from the sample, and (b) recomputing AI score over only topics with AI-specific labels. If the Q2 coefficient advantage over Q1 and Q4 shifts by more than 20%, or the Q2−Q4 Wald test becomes insignificant at p<0.05, the non-monotonic claim is not robust to taxonomy contamination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every estimate in Tables 1–5 is conditional on a first-stage classification: a paper is 'AI-referencing' if any cited work is in OpenAlex's Artificial Intelligence subfield (§3.1). Table S1 lists all 77 topics in that subfield, and several have no AI-specific content — e.g., 'Geochemistry and Geologic Mapping' (#38), 'Seismology and Earthquake Studies' (#65), 'Solar Radiation and Photovoltaics' (#68), 'Psychiatry, Mental Health, Neuroscience' (#60). Topic labels are not decisive, but the paper provides no evidence about the composition or weight of such topics among the flagged references. Because the AI dummy and AI score are simple counts over all 77 topics, a non-AI reference in any of these topics misclassifies a paper. The paper's own §5.3 acknowledges the taxonomy 'may include papers only loosely related to AI,' but no sensitivity analysis varies the taxonomy. The robustness checks (OLS in Table S4, 3-group split in Table S5) reuse the same contaminated regressor, so they cannot detect this. If a substantial share of flagged references falls in non-AI topics, the field contrasts, career-stage contrasts, and especially the institutional non-monotonicity in Table 3 (0.463/0.663/0.525/0.501) describe a mixture of AI and non-AI knowledge integration. The 'translational capacity' explanation then has no clean empirical referent. This is the most load-bearing weakness because every downstream claim inherits the classification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks how the association between AI knowledge integration and five-year citation impact varies across fields, career stages, and institutional AI capability. Using SciSciNet V2 data on 12.8 million non-computer-science papers from 360 CSRankings-listed institutions (2000–2024), the authors measure AI integration as citing at least one paper in the OpenAlex Artificial Intelligence subfield (extensive margin) and as the share of such references (intensive margin). PPML models with year, field, and document-type fixed effects and institution-clustered standard errors show a positive average association, field heterogeneity, career-stage interactions, and a non-monotonic institutional pattern in which intermediate-capability institutions (Q2/Q3) have the largest AI-score coefficients. Mechanism analyses profile AI knowledge granularity, epistemic crowding, AI-source displacement, and audience diversity. The paper interprets the pattern as evidence for 'translational capacity' — the ability to make AI knowledge meaningful and useful across scientific communities.","tokens_in":21475,"tokens_out":4798,"duration_ms":46278,"significance":"The conditional-returns question is timely and important, and the paper provides a credible observational baseline at unusual scale. The PPML specification, fixed effects, clustered standard errors, and the supplementary OLS and three-group robustness checks are appropriate and clearly described. The paper also honestly labels the analysis as observational. However, the central claim cannot be accepted as currently stated because the measure of 'AI knowledge integration' is built on a classification that is demonstrably contaminated by non-AI topics, and the headline non-monotonicity (H4) is read off point estimates without a formal statistical test. If these two issues can be addressed — with a validated or at least sensitivity-tested measure and a proper test of coefficient differences — the paper would make a valuable contribution to the science-of-science and AI-policy literatures.","major_comments":[{"comment":"The measurement of AIRef and AI score is a simple count over all 77 topics in the OpenAlex Artificial Intelligence subfield. Table S1 lists topics with no obvious AI content — e.g., 'Geochemistry and Geologic Mapping' (#38), 'Seismology and Earthquake Studies' (#65), 'Solar Radiation and Photovoltaics' (#68), 'Psychiatry, Mental Health, Neuroscience' (#60). A reference to any paper in these topics counts as an AI reference. The paper's own §5.3 concedes the taxonomy 'may include papers only loosely related to AI,' but no sensitivity analysis removes or reweights these topics. The OLS robustness check (Table S4) and the three-group split (Table S5) reuse the same contaminated regressor, so they cannot detect this problem. Because the non-monotonic pattern in Table 3 (0.463/0.663/0.525/0.501) could be driven by systematic variation across Q1–Q4 in the non-AI component of the AI score, the","section":"§3.1, Eq. (2), Table S1"},{"comment":"Hypothesis 4 is supported by comparing four separately estimated coefficients: 0.463 (Q1), 0.663 (Q2), 0.525 (Q3), 0.501 (Q4). No test is reported for whether these coefficients are statistically distinguishable from each other. The confidence intervals likely overlap substantially, and the ordering of point estimates is not by itself evidence of non-monotonicity. A formal Wald test of equality across groups — or a single interaction model with Q1 as the reference — is needed before concluding that intermediate-capability institutions obtain the largest returns. The same issue affects the three-group robustness check in Table S5. This is load-bearing because the paper's central novelty is the non-monotonic institutional pattern.","section":"§4.5, Table 3"},{"comment":"The coefficient-attenuation analysis is presented as evidence for the 'translational capacity' pathway, but the mechanism variables are measured after publication. Adding AudienceEntropy to the model can attenuate the AI-score coefficient even if it is an outcome, not a mediator; no mediation assumptions are stated. The one-sided Wald tests and 'Before/After' comparisons use different samples per variable (N ranges from 778,474 to 982,515), so the attenuation percentages (5.2%, 37.5%, 23.5%, etc.) are not comparable across rows. The paper does label the analysis descriptive and not causal, but the interpretive language — 'help explain' and 'pathway' — goes beyond what this design can support. I would ask the authors to rephrase this section as a profile of correlates and to avoid implying mediation.","section":"§4.7, Table 5"},{"comment":"The sample window ends in 2024, but the dependent variable is 'up-to-five-year' citations. For papers published after 2019, five years of citation accumulation are not fully observed. If AI-referencing propensity changes over time — as Figure 2 suggests — the censoring can confound publication-year fixed effects and, more importantly, the cross-sectional institutional contrasts if publication-year mix differs across capability quartiles. The authors should either restrict the sample to cohorts with a complete five-year window or demonstrate that the main results are robust to such a restriction.","section":"§3.1, §3.2.1"}],"minor_comments":[{"comment":"The text contains an unresolved cross-reference placeholder: '错误!未找到引用源。'. This should be fixed.","section":"§4.2"},{"comment":"Career stage is measured using the first author's academic age only. Because citation impact is a team product, a robustness check using the corresponding author, senior author, or average academic age of all authors would strengthen the career-stage claims.","section":"§3.2.3, Table 2"},{"comment":"The CSRankings selection rule ('at least eight years between 2015 and 2025') and the negative-log rank transformation are ad hoc. Please provide a rationale or a sensitivity check on the inclusion threshold and transformation.","section":"§3.2.4"},{"comment":"Several references are future-dated or in press (e.g., Bianchini et al., 2026; Cui et al., 2026; Zhao and Li, 2026; Wilinski, 2026). Verify that these citations are correct and accessible, and update any preprint identifiers if needed.","section":"References"},{"comment":"The note acknowledges that the mean age of cited AI papers can be influenced by extreme values, but the panel still uses the mean. A median or winsorized version would be more robust and consistent with the caveat.","section":"Figure 4, Panel b"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid observational design and a clearly stated research question, but the two load-bearing issues — the unvalidated AI-reference taxonomy and the absence of a formal test for the non-monotonic institutional pattern — should be resolved before publication. If the authors can provide a sensitivity analysis excluding non-AI topics and a proper coefficient-difference test, I would be willing to reconsider. The current version is not suitable for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, the genuinely new results are the career-stage asymmetry (seniors get credit for merely citing AI; juniors only for deep engagement) and the intermediate-institution peak in citation returns, with a pathway analysis linking that peak to audience diversity and AI-source displacement. Second, the paper's central measure of AI integration is built on OpenAlex's AI subfield, and the paper's own Table S1 lists topics like Geochemistry and Seismology as part of that subfield. That contamination is the soft underbelly of every estimate in Tables 1–5.\n\nCredit where it's due. The baseline H1 results reproduce prior findings, and the authors say so. The paper is careful: PPML, clustered SEs, year/field/doc-type FE, OLS and three-group splits as robustness checks, and an explicit observational framing. The descriptive mapping of who benefits is solid and policy-relevant. The career-stage and institutional contrasts are novel and internally consistent.\n\nNow the soft spots, in proportion. The taxonomy issue is real and load-bearing. A paper gets flagged as AI-referencing if it cites any OpenAlex work in the AI subfield. Topics like 'Geochemistry and Geologic Mapping,' 'Seismology and Earthquake Studies,' and 'Solar Radiation and Photovoltaics' have no obvious AI content. If those topics are a substantial share of flagged references, the field, career, and institutional contrasts are mixtures of AI and non-AI integration, and 'translational capacity' has no clean empirical referent. The paper acknowledges in §5.3 that the taxonomy 'may include papers only loosely related to AI,' but no sensitivity analysis varies the taxonomy. The robustness checks reuse the same contaminated regressor, so they can't detect this. This is fixable—curate the topic list, drop suspect topics, validate against a sample with known AI tool use—but until then the central claim is conditional on an unvalidated classification.\n\nTwo more moderate issues. The flagship non-monotonicity in Table 3 (0.463/0.663/0.525/0.501) is read off four separate regressions; no formal test compares those coefficients across groups. And the attenuation analysis in Table 5 shows that once audience entropy is added, the Wald tests become insignificant—which undercuts the 'translational capacity' mechanism. The paper labels this exploratory, so it's a limitation, not a fatal flaw. Similarly, the audience entropy variable is post-treatment and conditions on being cited; the paper acknowledges that too.\n\nFor whom is this? Science-of-science researchers and science policy people. It deserves a serious referee, not a desk reject. But the referee should push for a taxonomy robustness analysis and formal cross-group coefficient tests before the conditional-return claim can be accepted at face value.","headline":"Well-executed bibliometric study with a plausible conditional-return story, but the AI-integration measure rests on a taxonomy that visibly includes non-AI topics, and the flagship non-monotonicity lacks a formal cross-group test.","tokens_in":22076,"tokens_out":2224,"would_cite":false,"duration_ms":23612,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The citation payoff to AI knowledge is unevenly distributed and peaks at intermediate institutional AI capability, with field context and career stage shifting who benefits.","keywords":["AI knowledge integration","citation impact","translational capacity","field heterogeneity","career stage","institutional capability","non-monotonic returns","science of science"],"falsifier":"Recompute the AI-score returns after dropping the non-AI topics from the AI subfield definition (e.g., geochemistry, seismology, psychiatry); if the intermediate-capability peak disappears, the central non-monotonic claim is an artifact of the taxonomy rather than a property of AI knowledge.","tokens_in":20985,"feed_emoji":"📈","tokens_out":7509,"duration_ms":67404,"temperature":0.7,"pith_summary":"This paper tries to establish that the scientific value of AI knowledge is not evenly distributed and cannot be predicted from technical AI strength alone. Working from millions of papers and their references, the authors show that papers citing AI literature are on average cited more, but the size of that advantage shifts with the field, the first author's career stage, and the AI capability of the home institution. The most striking pattern is non-monotonic: the proportional citation return to the share of AI references is largest at institutions with intermediate AI capability and declines for the very highest-capability institutions. The authors attribute this to 'translational capacity' — the ability to make AI knowledge meaningful to communities outside AI itself — as distinct from absorptive capacity. If correct, this reframes AI-enabled science as a problem of knowledge translation, not only of technical access, which matters for how universities and funders allocate AI support.","feed_headline":"AI's citation payoff peaks at mid-tier institutions, not the top","feed_subtitle":"Proportional citation gains peak at moderate AI strength; translating AI across fields matters more than raw capability.","key_machinery":"The central machinery is the two-margin measurement of AI knowledge integration — an indicator for any AI-related reference (extensive margin) and the share of AI-related references among all references (intensive margin) — combined with paper-level contexts (field, first-authored career stage, and an institution's AI-capability quartile). The load-bearing comparison is institutional: when papers are split into four capability groups, the coefficients on the AI share are 0.463, 0.663, 0.525, and 0.501, generating the paper's signature non-monotonic pattern. The interpretive mechanism is 'translational capacity,' the ability to make AI knowledge meaningful, legitimate, and useful across scien","core_discovery":"On the paper's own terms, the central discovery is that the citation return to AI knowledge integration is conditional rather than universal. Papers that cite AI-related literature receive higher five-year citation counts on average, but the premium depends on three layers of context. Across fields, the return to both extensive and intensive AI referencing varies widely, and in some fields intensive referencing is even negatively associated with citations. By career stage, the extensive margin (simply having an AI reference) pays off mainly for senior scholars, while the intensive margin (the share of AI references) pays off mainly for junior scholars, who also tend to cite newer and higher-","pith_inferences":["Beyond the paper's claims: an externally validated measure of actual AI use (e.g., methods sections naming specific models, or code/dataset availability) could test whether the non-monotonic institutional pattern survives a proxy-free definition of AI integration.","The authors themselves note that the AI-subfield taxonomy includes topics such as geochemistry and seismology; I would infer that a robustness analysis excluding those non-AI topics from the reference measure would tell us whether the Q2 peak is a property of AI knowledge or of the taxonomy.","The career-stage asymmetry suggests a signaling-versus-substance mechanism; a natural test is whether the senior-scholar extensive-margin advantage is concentrated in high-status authors and the junior-scholar intensive-margin advantage grows with the recency and impact of the cited AI work.","The audience-entropy pathway is correlational; an intervention-style design—for example, tracking the citation trajectories of papers that deliberately frame methods for non-AI audiences—could provide stronger evidence for the translational-capacity mechanism."],"forward_implications":["If the pattern holds, simply adding AI infrastructure and expertise will not equalize returns; support for cross-field translation, such as interdisciplinary collaboration and methodological services, becomes a distinct policy lever.","Uniform incentives to adopt AI may backfire by encouraging symbolic AI referencing, which is exactly the margin that does not pay for early-career scholars; evaluation and training should reward substantive integration.","Top-tier and mid-tier institutions play complementary roles: frontier institutions supply methods and become citation gateways within AI, while intermediate institutions carry AI knowledge into other fields; a diversified portfolio of AI investment is preferable to a single excellence ranking.","Field-specific negative returns to intensive AI referencing imply that AI-knowledge diffusion policies must be tailored to disciplinary evaluation standards rather than applied across the board.","Because the largest proportional gains appear at intermediate capability, measures of scientific impact from AI use should be separated from measures of technical AI strength; 'translational capacity' deserves to be measured in its own right."],"fun_headline_variants":["Mid-tier AI labs get biggest citation boost, not top","AI citations pay off most at mid-tier institutions","Citation goldilocks: mid-tier AI strength wins","AI research impact peaks at mid-level, not elite","Translation beats raw AI power for citation gains"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The analysis assumes that references to papers classified in the 'Artificial Intelligence' subfield measure genuine AI knowledge integration, even though that subfield includes non-AI topics and references are only a proxy for actual AI use; if that classification or proxy fails, every reported contrast is contaminated.","fun_headline_variants_meta":{"raw":{"variants":["Mid-tier AI labs get biggest citation boost, not top","AI citations pay off most at mid-tier institutions","Citation goldilocks: mid-tier AI strength wins","AI research impact peaks at mid-level, not elite","Translation beats raw AI power for citation gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000105,"raw_usage":{"total_tokens":860,"prompt_tokens":722,"completion_tokens":138,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":64}},"tokens_in":466,"tokens_out":138,"duration_ms":2431,"temperature":1.0,"reasoning_tokens":64,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T19:58:28.528327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the AI-score returns after dropping the non-AI topics from the AI subfield definition (e.g., geochemistry, seismology, psychiatry); if the intermediate-capability peak disappears, the central non-monotonic claim is an artifact of the taxonomy rather than a property of AI knowledge.","supporting_citations":[],"review_version":1}