{"id":"e16cd609-6abd-44c1-a95a-d7ba21964d8c","arxiv_id":"2607.20923","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"After 2022, scientists became more interdisciplinary and exploratory, collaborated across more fields, and reported more differentiated team roles, with the largest shifts among high-AI-writing, established, and non-English LMIC authors.","lead":"This study tracks 775,323 scientists through 2011 to 2025 and finds that after 2022, researchers published in more distant fields, entered more new fields, collaborated with more diverse colleagues, and reported narrower, less overlapping contribution roles. The changes were strongest among established scientists, authors from non-English-speaking lower- and middle-income countries, and scientists with stronger AI-writing signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI-writing proxy is the load-bearing assumption; a holdout-period reanalysis is needed before 'selection-plus-reinforcement' is accepted.","rationale":"The paper is careful and largely self-consistent. The temporal post-2022 patterns are supported by many consistent measures, author fixed-effect event studies, alternative field classifications, and restricted-sample robustness checks; I do not see an internal error in the decomposition or in Equation 5. The most distinctive interpretive layer, however, is the claim that authors with stronger AI-writing signals were already more exploratory before LLM adoption and that the gap widened afterward. That claim is only as strong as the AI-writing proxy. The proxy is a bag-of-words mixture trained on bioRxiv and applied to PMC, with no demonstrated measurement invariance across fields and time. Because the proxy is computed from the same post-2022 window used to define the contrast, and because the outcome variables (Rao-Stirling, pivot size, collaborator diversity) are plausibly correlated with the vocabulary features the model detects, the pre-2023 'selection' gap could be a proxy artifact rather than a behavioral selection effect. This is the most load-bearing concern because it affects the paper's distinctive contribution over and above the more general observation that science changed after 2022. The reader's CONDITIONAL verdict already captures this uncertainty, so I do not recommend moving the verdict; it should remain conditional pending the holdout-proxy reanalysis. The absence of released code is also worth noting, but the proxy validity is the scientific crux.","tokens_in":26360,"tokens_out":11003,"duration_ms":111548,"concrete_test":"Recompute the high- versus low-AI-writing comparisons using a proxy built only from pre-LLM text: for each author, average the AI-writing fraction over their 2021-2022 PMC papers only, assign groups with the same 0.15/0.05 thresholds (or with within-field quantiles if the holdout yields too few cases), and re-estimate the 2011-2022 trajectories underlying Figures 2d-i and 3e-g. If the pre-2023 high-low gaps in Rao-Stirling, pivot size, and collaborator diversity shrink to null, the selection-plus-reinforcement interpretation is an artifact of measuring the proxy in the outcome window. A stronger variant is to calibrate the mixture model separately within each of the 26 OpenAlex fields on labeled 2015-2019 human-written papers and use field-specific z-scores to define high/low groups before rerunning the matched comparisons.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the author-level AI-writing rate, estimated by the corpus-level word-frequency mixture of Supplementary Note 1, is a valid proxy for LLM exposure that is comparable across fields and over time. Figures 2 and 3 and the Discussion's selection-plus-reinforcement interpretation all depend on this proxy. The model is a bag-of-words mixture with field-invariant human and AI word distributions trained on bioRxiv and applied to PMC; the high/low contrast uses a 0.15/0.05 threshold on the average fraction in the author's 2023-2025 papers. The same post-2022 window supplies both the group definition and the outcome measurements in which the high-low gap widens, so any stylistic feature of interdisciplinary or exploratory papers that resembles LLM text can masquerade as 'AI-writing signal'. Matching controls for primary field, country context, career stage, and productivity, but not for subfield, method vocabulary, or writing style, and the validation in Supplementary Fig. S13 (correlation with commercial detectors on abstracts) does not establish measurement invariance across the 26 fields or across time. If the proxy's error is correlated with Rao-Stirling index, pivot size, or collaborator diversity, the pre-2023 differences and post-2023 widening in Figures 2d-i and 3e-g are partly an artifact. The descriptive temporal changes survive, but the distinctive AI-intensity claims do not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses a linked PubMed Central/OpenAlex sample of 775,323 authors and 137,120 multi-author CRediT papers to describe changes in scientific exploration, collaboration, and the division of labor after the 2022 public diffusion of large language models. The authors report that scientists' field portfolios became broader and more exploratory after 2022, that authors with stronger AI-writing signals were already more interdisciplinary and exploratory before that date and the gap widened afterward, that collaboration networks became more interdisciplinary, and that reported CRediT role sets became narrower, less overlapping, and less rigidly bundled. The paper is explicitly descriptive and repeatedly cautions that the AI-writing measure is a textual proxy rather than direct evidence of LLM use.","tokens_in":26640,"tokens_out":4764,"duration_ms":47043,"significance":"If the descriptive findings hold, this is an important and timely contribution to the science-of-science literature, documenting a broad reorganization of scientific work in the LLM era. The study's strengths include its very large linked dataset, the use of multiple complementary measures (field counts, Shannon entropy, HHI, Rao-Stirling index, pivot size, reference-field diversity), author-level fixed-effects event studies, CEM matching, and a candid limitations section. The temporal trends are supported by several robustness checks. However, the AI-intensity comparisons and the selection-plus-reinforcement interpretation are more fragile because the AI-writing proxy's measurement invariance is not established; the quantitative counterfactual claims also lack uncertainty quantification. These issues are fixable with additional analyses, so the paper is promising but not yet ready in its current form.","major_comments":[{"comment":"The AI-writing fraction is estimated from a word-frequency mixture model trained on bioRxiv text and applied to PMC papers, but no evidence is provided for measurement invariance across the 26 fields or over time. Because high- and low-AI-writing groups are defined from the same 2023-2025 window in which the outcome measures are computed, any field- or time-specific stylistic feature of interdisciplinary or exploratory papers that resembles LLM text could masquerade as 'AI-writing signal'. The validation in Supplementary Fig. S13 uses only abstracts and commercial detectors, not full-text, per-field, or per-year calibration. This is load-bearing because the 'selection-plus-reinforcement' interpretation in the Discussion and the widening-gap claims in Fig. 2 and Fig. 3 depend on the proxy. I request a holdout-style validation: report field-specific and year-specific false-positive rates against labeled human/AI text, or demonstrate that the association between the proxy and Rao-Stirling index, pivot size, and collaborator diversity survives within-field and within-subfield calibration.","section":"Supplementary Note 1 and Fig. 2(c)-(i)"},{"comment":"The quantitative claims of the form '15.6% higher than expected', '13.0% higher', and '23.2% higher' are based on linear extrapolation of 2011-2022 trends, but no uncertainty is attached to the dashed counterfactual extensions. With only two to three post-2022 observation points and noticeable pre-trend noise (e.g., the pivot measure declining before rebounding), these percentage excesses may be sensitive to the choice of pre-period window and functional form. I ask for forecast intervals or bootstrap confidence bands on the excess values, and for sensitivity analyses using alternative counterfactuals (e.g., a shorter pre-COVID window, a damped trend, or a placebo break year).","section":"Figs. 1-3, Results sections on research portfolios and collaboration"},{"comment":"The claim that the association between paper-based and collaborator-based interdisciplinarity is 'consistently weaker' for high-AI-writing authors and weakens further after 2022 is central to the individual-level knowledge integration interpretation, but the manuscript does not report the estimating equation, coefficient estimates, or confidence intervals for this interaction. The event-study models in Supplementary Note 2 are described for the high-low comparisons in Supplementary Fig. S9, not for the Fig. 3h association. Please specify the model, present the interaction coefficients with standard errors, and ideally show an event-study plot with the high-low difference in the paper-collaborator Rao-Stirling slope over time.","section":"Fig. 3h and Supplementary Note 2"}],"minor_comments":[{"comment":"Career-stage bins are inconsistent: the Methods section defines early career as %u22645 years, mid-career as 6-15 years, and senior as %u226516 years, while Fig. 1's caption refers to Junior (<6), Early career (6-15), and Advanced (>15 yr). Please align the definitions.","section":"Methods and Fig. 1 caption"},{"comment":"The paper alternates between 'after 2022', 'post-2022', and dashed vertical lines at 2023. Please standardize the epoch boundary description (e.g., 'the 2023-2025 period') to avoid ambiguity about whether 2022 is included in the pre- or post-period.","section":"Throughout"},{"comment":"The pivot measure is described as 'cosine distance ... using Eq. 1', but Eq. (1) defines field distance from reference vectors of fields. Please state explicitly that the pivot measure applies the same cosine-distance formula to focal and prior reference vectors, and clarify the vector construction in both cases.","section":"Methods, Eq. (1) and Research Pivot Measures"},{"comment":"In Eq. (S8), the term n_it appears without a coefficient, making it look like a regressor without a parameter. Please write the model with an explicit coefficient for publication counts or state that it is included as a control.","section":"Supplementary Note 2, Eq. (S8)"},{"comment":"The sentence 'This is most likely driven by Chinese authors' would be stronger with a formal statistical test of country-by-year interactions rather than a visual inspection of Supplementary Fig. S7.","section":"Results, Country heterogeneity paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for this venue and the descriptive trends are likely to be of broad interest. The main concern is methodological rather than conceptual: the AI-writing proxy needs a much stronger validation story before the AI-intensity claims can be accepted, and the counterfactual extrapolations need uncertainty quantification. I would also encourage the editor to ask for a code/data availability statement, since the current Declarations section only lists data sources."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth engaging. It brings big, linked data to a question a lot of us care about: did the LLM era change not just prose but the actual organization of scientific work? The descriptive patterns are striking and mostly convincing: after 2022, authors publish in more fields, enter new fields more often, collaborate with more interdisciplinary partners, and report narrower, less overlapping CRediT role sets. The genuinely new pieces are the matched high- versus low-AI-writing author comparisons, the weakening link between own and collaborators' interdisciplinarity among high-AI authors, and the role-bundling decomposition against a randomized baseline. I don't see those specific results in the cited literature.\n\nThe paper does several things well. The sample is enormous, the measurement approach is multi-pronged, and the fixed-effect event studies with author and year controls address the obvious composition and productivity confounds. The CEM matching is careful, and the balance diagnostics are reported. To its credit, the paper is explicit that the design is descriptive and that the AI-writing measure is a proxy, not a direct observation of LLM use. The central temporal claim holds up: too many independent measures move together after 2022 for it to be an artifact of any single coding decision.\n\nThe soft spots are real but proportionate. The AI-writing intensity analysis is only as good as the proxy. The word-frequency mixture was trained on bioRxiv and applied to PMC, and the validation against commercial detectors covers abstracts only. That does not establish measurement invariance across 26 fields or over time. The high- versus low-AI-writing groups are defined using the same 2023-2025 window in which the outcome gaps widen, so a stylistic confound in interdisciplinary or exploratory papers could masquerade as AI-writing signal. This is the load-bearing assumption for the selection-plus-reinforcement interpretation, and the stress-test note is right that a holdout-style reanalysis, or at least a field-specific calibration, is needed before that interpretation is taken at face value. The descriptive temporal changes, by contrast, do not depend on this proxy.\n\nThe expected-value comparisons are linear extrapolations from 2011-2022 without uncertainty on the counterfactual. The paper mostly treats them as illustrative, but the figures could mislead a casual reader. The 0.15/0.05 thresholds are reasonable but hand-chosen. The biggest practical gap is that no code or processing scripts are released, which makes it hard to check the pipeline from raw PMC/OpenAlex to the author-level measures. The biomedical-heavy sample is acknowledged honestly, and the citation pattern covers the relevant prior work; the self-citations are to earlier measurement papers and are appropriate.\n\nThis paper deserves a serious referee. The issues are addressable rather than structural. I would send it to review with a request: validate or bound the AI-writing proxy, attach uncertainty to the counterfactual extrapolations, and release code and derived data. For science-of-science and bibliometric readers, this is a useful and in places agenda-setting contribution.","headline":"Large-scale descriptive evidence that post-2022 science became more exploratory and more modularly staffed; the AI-writing-intensity comparisons are suggestive but rest on a proxy that needs further validation.","tokens_in":27132,"tokens_out":2111,"would_cite":true,"duration_ms":22557,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that, after large language models became widely available in late 2022, scientists in a large biomedical sample began publishing in more intellectually distant fields, entering new fields more often, collaborating with…","keywords":["large language models","scientific collaboration","interdisciplinarity","research exploration","labor division","CRediT","AI-assisted writing","science of science"],"falsifier":"Take the same author population and measure LLM exposure directly, for example through linked survey responses or keystroke and editor logs, and check whether the high-low gaps in interdisciplinarity and role differentiation persist; if the AI-writing fraction does not track actual LLM use or the gaps vanish when actual use is controlled, the central claim is falsified. A complementary check is to apply the AI-writing proxy to a placebo corpus written before LLMs existed and see whether the same high-low gaps appear in counterfactual pre-2022 years; they should not.","tokens_in":26166,"feed_emoji":"🔬","tokens_out":8204,"duration_ms":69560,"temperature":0.7,"pith_summary":"After late 2022, when large language models became widely available, scientists in this large sample changed how they choose research directions, whom they collaborate with, and how they divide tasks. The paper finds that researchers publish across more intellectually distant fields, enter fields new to them at higher rates, and collaborate with a more disciplinarily diverse set of coauthors, with the biggest shifts among established scientists and researchers from non-English-speaking low- and middle-income countries. Authors whose papers show stronger AI-writing signals were already more interdisciplinary and exploratory before 2022, and the gap with low-AI-writing authors widened afterward. Within teams, contributors report fewer and more distinct roles, with software and validation duties growing while conceptual and management duties shrink. The authors read this as a broad reorganization of scientific work in the LLM era, not as proof that LLMs caused these changes.","feed_headline":"AI writing era linked to broader science and narrower team roles","feed_subtitle":"A study of 775,000 scientists ties wider field portfolios and more distinct team roles to AI-assisted writing intensity.","key_machinery":"The argument is carried by two instruments. The first is an author-level AI-writing fraction, estimated from a word-frequency mixture model that treats each paper as a mix of human-written and AI-generated sentences and fits the mixing fraction by maximum likelihood; it is the paper's proxy for LLM exposure. The second is a family of author-year portfolio measures built from field distances computed as cosine distances between fields' reference profiles, including a field-diversity index, a Rao-Stirling interdisciplinarity index, and a pivot-size measure that compares each paper's references to the author's prior three-year reference profile. For division of labor, the paper uses CRediT role statements to compute roles per author, pairwise Jaccard similarity of coauthor role sets, and a degree-preserving within-paper shuffle that tests whether role combinations deviate from random allocation. These measures together let the paper connect textual AI-writing signals to behavioral changes at the portfolio, network, and team levels.","core_discovery":"Using publication histories for 775,323 scientists and contribution statements from 137,120 multi-author papers, the paper establishes that the post-2022 period coincides with systematic increases in three linked dimensions: portfolio breadth and exploration, interdisciplinary collaboration, and differentiation of labor. It further shows that scientists with high AI-writing rates were already more interdisciplinary and exploratory before LLM diffusion and that the gap widened after 2022, supporting a selection-plus-reinforcement interpretation rather than a simple treatment effect. On division of labor, the decline in individual role counts is driven mainly by reduced sharing of the same roles among coauthors, and role combinations move closer to a randomized within-paper baseline, indicating weaker fixed role bundling. The authors are explicit that the design is descriptive and cannot separate LLM adoption from other post-2022 changes.","pith_inferences":["My inference: the selection-plus-reinforcement pattern suggests a widening rather than a converging trajectory, with already-exploratory scientists gaining the most from LLM tools and less-exploratory scientists potentially falling further behind even as tool access becomes universal.","My inference: the CRediT role shifts may partly reflect changes in how journals ask authors to report contributions rather than in how work is actually performed; comparing the same journals' reporting instructions over time would isolate the reporting-norm component.","My inference: if AI tools are substituting for some collaborative functions, one testable extension is that team sizes in high-AI-writing venues will stagnate or decline relative to pre-2022 trends after controlling for field effects, a pattern this paper does not directly test."],"forward_implications":["If the portfolio-breadth finding holds, research evaluations that treat field concentration as a stability metric will need to view post-2022 breadth as a real behavioral shift rather than an indexing artifact.","If high-AI-writing authors were already more exploratory, providing LLM access alone will not make all scientists equally interdisciplinary; the gap between early adopters and others may widen.","If the weaker link between collaborator diversity and paper interdisciplinarity holds for high-AI-writing authors, individual-level AI tool use may be partly substituting for the knowledge-integration role that cross-field collaborators previously played.","If role differentiation continues, contribution-reporting norms and evaluation rubrics that reward broad, shared role profiles will need to accommodate narrower, more modular task assignments."],"supporting_citations":[{"why":"Supplies the word-frequency mixture model used to estimate each paper's AI-writing fraction.","marker":"[10]"},{"why":"Provides the reference text distributions from which the mixture model's per-word probabilities are taken.","marker":"[9]"},{"why":"Supplies the co-reference network approach used to detect changes in scientists' research agendas.","marker":"[37]"},{"why":"Grounds the claim that scientists who use AI for writing are more likely to use AI for other research tasks, supporting the proxy interpretation.","marker":"[38]"},{"why":"Provides the CRediT-based framework for measuring how scientific labor is divided within teams.","marker":"[40]"}],"fun_headline_variants":["LLM era widens research scope and sharpens team roles","AI-writing era: broader exploration, more distinct labor","LLM era linked to wider science and narrower roles","Post-2022 science: more exploration, less role overlap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the paper-level AI-writing fraction, a word-frequency score, is a valid author-level proxy for actual LLM exposure; if it is biased by field, time, or writing style, the high-versus-low AI-writing comparisons and the selection-plus-reinforcement conclusion could be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["LLM era widens research scope and sharpens team roles","AI-writing era: broader exploration, more distinct labor","LLM era linked to wider science and narrower roles","Post-2022 science: more exploration, less role overlap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001014,"raw_usage":{"total_tokens":4304,"prompt_tokens":989,"completion_tokens":3315,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":3248}},"tokens_in":605,"tokens_out":3315,"duration_ms":19873,"temperature":1.0,"reasoning_tokens":3248,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:30:37.641413+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same author population and measure LLM exposure directly, for example through linked survey responses or keystroke and editor logs, and check whether the high-low gaps in interdisciplinarity and role differentiation persist; if the AI-writing fraction does not track actual LLM use or the gaps vanish when actual use is controlled, the central claim is falsified. A complementary check is to apply the AI-writing proxy to a placebo corpus written before LLMs existed and see whether the same high-low gaps appear in counterfactual pre-2022 years; they should not.","supporting_citations":[{"cited_title":"Zhang, Z","cited_arxiv_id":null,"evidence_quote":"Provides the CRediT-based framework for measuring how scientific labor is divided within teams."}],"review_version":2}