{"id":"d1a6350c-a963-40d9-a9ed-9c2d9fe7686d","arxiv_id":"2607.10780","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Across 300M+ OpenAlex works, the long decline in solo authorship halted and partially reversed after ChatGPT’s release, strongest where coauthor tasks are more replaceable and among authors who previously only coauthored.","lead":"Solo-authored science papers stopped their decades-long decline and partially rebounded after ChatGPT’s late-2022 release, especially in computational fields. The pattern is a measurable probe of where generative AI may be substituting for human coauthors rather than only enlarging teams.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The substitution reading of the 2022 solo-tail break remains correlational; concurrent OpenAlex/MAG and post-pandemic confounds are only partially bounded.","rationale":"The Reader correctly isolates the load-bearing step: reading the common 2022 break and field/content signatures as evidence of LLM substitution rather than concurrent confounds. The paper’s multi-filter, history-conditioned, balanced-venue, and content DiD checks are real and carefully reported; they make a pure artifact story less likely and justify keeping CONDITIONAL rather than REJECT. They do not, however, deliver clean identification of ChatGPT as the cause, and SI L’s attenuation plus the MAG/disambiguation timing leave residual risk on the interpretive claim that solo papers are an empirical probe of AI substituting for coauthor labor. No stronger internal inconsistency appears: the mean-vs-tail distinction, never-solo hazards, and computational tilt without exploration are coherent with substitution if the timestamp is informative. Thus the Reader’s verdict and weakest_assumption stand; confidence remains moderate pending independent re-runs that freeze metadata and further strip pandemic composition.","tokens_in":21935,"tokens_out":702,"duration_ms":12093,"concrete_test":"Re-estimate the main peer-reviewed and core monthly/annual Δβ and the composition-adjusted author-level Pt(k) series on the SI L balanced venue panel after (i) freezing OpenAlex author IDs and primary-field assignments to a single pre-July-2023 snapshot for all years and (ii) excluding 2020–2021 COVID-heavy biomedical subfields; if the 2022 break falls below significance or loses the computational DiD tilt of Fig. 3b while the full-sample break remains, the substitution reading is not robust to the documented concurrent confounds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s central claim is not merely that the solo share flattened after late 2022, but that this break—plus field ordering by substitutability, history-conditioned author switches, and content DiD toward computational work near authors’ prior collaborative centroids—probes LLM substitution for coauthor labor (Abstract; Results “The mechanism is consistent with LLM substitution”; Discussion). That interpretation rests on treating the November 2022 ChatGPT release as a common temporal anchor while ruling out concurrent shocks. The paper itself labels the design correlational and a dating convention (Results; Limitations). SI D argues against pure pandemic reversion via field ordering, and SI L shows a positive but attenuated break on balanced 2018–2024 venues (peer-reviewed Δβ +0.75 vs +1.72 full sample; Spearman 0.84), yet OpenAlex’s MAG transition (end-2021) and July 2023 author-disambiguation revision remain inside the post window and can alter authorship metadata within continuously indexed sources. SI M donuts and event studies support a 2022–23 trend change, but do not isolate LLM adoption from other 2022–23 shocks. Content DiD and never-solo author hazards strengthen the substitution story but still condition on the same timestamp. If residual database or multi-shock composition drives a large share of the left-tail movement, the probe-of-substitution claim weakens even though the descriptive break may stand.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper uses the full OpenAlex corpus (1990–2025; ~300M works, 26 fields) to study the left tail of the author-count distribution rather than mean team size. It reports that the long decline in the share of solo-authored papers flattens or partially reverses around ChatGPT’s public release (Nov 2022), with a companion flattening of mean author counts in many fields. The break is heterogeneous: stronger in fields where coauthor tasks are more substitutable by writing/coding/analysis tools, and weak or absent in lab/instrument-heavy fields. Author-level, composition-standardized probabilities show the rebound among previously coauthor-only and never-solo authors, including seniors. Content analyses (SPECTER2 DiD netting field-wide drift; within-author breadth/exploration) find post-2022 solo papers tilt toward computational work, stay near authors’ prior collaborative content, and narrow in scope. The authors interpret solo authorship as an observable probe of AI substitution for parts of coauthor labor, while stating the design is correlational and anchored on a dating convention.","tokens_in":22340,"tokens_out":1946,"duration_ms":40006,"significance":"If the descriptive break and its author/content signatures hold, this is a substantial contribution to the science of science and to empirical work on generative AI’s effect on knowledge production. Focusing on the solo tail rather than the mean is a clear conceptual advance, and the multi-filter design (peer-reviewed, core, preprint, clean all), composition standardization, history-conditioned hazards P_t(k), balanced venue panels, event-study/donut timing checks, and embedding DiD are serious empirical strengths. The paper also carefully separates acceleration vs substitution predictions and assigns them to different parts of the author-count distribution. The result is falsifiable in principle (field ordering, never-solo hazards, content geometry) and would matter for authorship norms, credit, and training pathways even if the causal attribution remains incomplete.","major_comments":[{"comment":"The central interpretive claim—that the 2022 left-tail break is an empirical probe of LLM substitution for coauthor labor—rests on a common temporal anchor plus field ordering and content signatures (Abstract; Results “The mechanism is consistent with LLM substitution”; Discussion). The manuscript correctly labels the design correlational and a dating convention (Results; Limitations). Even so, Abstract/title framing and the mechanism section still invite a causal reading that SI D, L, and M only partially bound. Concurrent shocks (post-pandemic publishing, other 2022–23 AI tools, journal policy shifts) remain inside the post window. Please either (i) demote the substitution language consistently to “timing-consistent with” and lead with the descriptive break, or (ii) add a sharper multi-shock falsification (e.g., placebos at other 2020–2024 AI/tool releases; field-by-field comparison of","section":"Abstract; Results (mechanism); Limitations; SI D, L, M"},{"comment":"OpenAlex infrastructure changes are a load-bearing confound risk for authorship counts. SI L shows the peer-reviewed break attenuates from +1.72 to +0.75 pp/yr on balanced 2018–2024 venues (Spearman 0.84 across fields), which bounds venue entry/exit but explicitly cannot rule out within-venue metadata or author-disambiguation changes. MAG discontinuation (end-2021) and the July 2023 author-disambiguation revision fall inside or adjacent to the post window (SI L; Limitations). Because solo status is defined from author IDs, disambiguation revisions can mechanically create or destroy solo papers. A main-text robustness that freezes author IDs / re-runs on a pre-revision snapshot, or at least reports sensitivity of Δβ to excluding mid-2023, is needed before the within-author switching claim can be treated as database-robust.","section":"Limitations; SI Section L; Materials and Methods (author IDs)"},{"comment":"Field-level and author-level breaks do not coincide in Engineering, the largest share-level rebound (+2.5 pp/yr peer-reviewed), where author-level solo probability barely moves and the balanced-venue break falls to +0.65 (SI H, L; Results). The paper notes this, but main-text claims that “the break appears among authors who had written only with others” and that composition “does not explain the result” (Abstract; Results) over-generalize. Engineering should be treated as an explicit exception in the Abstract/Results, and the substitution narrative should be restricted to fields where within-author Δβ is positive (e.g., CS, Psychology, Economics in SI Fig. S8).","section":"Abstract; Results; SI Sections H and L"},{"comment":"SI Section K (external Liang et al. corpus) finds multi-author papers carry at least as much estimated LLM-modified writing as solo papers in every arXiv field. That is a useful check against a pure surface-editing account of the solo rebound, but it is currently buried and not integrated with the main mechanism claim. If writing assistance is pervasive on teams, the substitution story must emphasize non-writing execution (coding, analysis, drafting structure) or the decision margin to publish without coauthors—not text polish. Please move a concise version of this result into the main Results/Discussion and state what it does and does not identify.","section":"Results (mechanism); Discussion; SI Section K"},{"comment":"Quality and paper-mill inflation remain open for the left-tail recovery (Limitations). The rebound is largest for preprints, where lag is shortest but low-cost solo output is also easiest. Peer-reviewed and core filters still show positive breaks with CIs excluding zero, which is important, yet venue-based filters cannot verify refereeing quality. Without any quality/impact/retraction/duplicate screen on recovered solo papers, the claim of a “reconfiguration of cognitive labor” risks conflating genuine substitution with an influx of low-value solo output. At minimum, report citation or journal-tier distributions for pre- vs post-2022 solo papers under the peer-reviewed filter, or flag quality as outside scope more prominently in the Abstract.","section":"Limitations; Results (preprint vs peer-reviewed); Abstract"}],"minor_comments":[{"comment":"Figure 1 reports Δβ as change in annual slope of the solo share; clarify in the caption whether monthly fits are rescaled to annual units and how seasonal adjustment (used for mean authors in SI C) is handled for the solo share.","section":"Figure 1; SI Section C"},{"comment":"Eq. (1)–(2) for P_t(k) and d_it are clear, but the main text sometimes uses d_active without restating that it is predetermined and resets after every solo year; a one-sentence reminder would help non-specialist readers.","section":"Materials and Methods; SI Section F"},{"comment":"Figure 3a UMAP atlas is descriptive only (as stated), but the two language-defined clusters (Turkish, Indonesian) are easy to misread as topical; consider a footnote or inset noting they are excluded from the English-only axis analysis.","section":"Figure 3; SI Section I"},{"comment":"The Acknowledgments disclose extensive Claude use for code and manuscript revision. Given the paper’s topic, a brief Methods note on which analyses were AI-assisted vs human-verified would strengthen reproducibility norms the paper itself discusses.","section":"Acknowledgments; Materials and Methods"},{"comment":"SI Fig. S3 heatmap significance uses bootstrap conditional on the observed source universe; state that limitation once in the main text when citing field-level significance counts (23/26 fields).","section":"Results; SI Figure S3"},{"comment":"Minor wording: title truncates “generative A” in the provided header; ensure the published title is complete (“generative AI”).","section":"Title"}],"recommendation":"major_revision","confidential_remarks":"The descriptive left-tail break is the paper’s real contribution and is unusually well documented for this literature; I would not reject on identification grounds alone. The revision ask is mainly to stop the Abstract/mechanism from outrunning the correlational design and to close the OpenAlex disambiguation risk enough that within-author claims are credible. If the authors refuse to elevate the Engineering composition caveat and the SI K multi≥solo writing result, the substitution framing becomes harder to defend at this journal’s standard. Scope fit is good for a computational social science / science-of-science venue; less so if the journal expects causal designs for AI-labor claims."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new result is a broad halt and partial reversal of the long decline in solo-authored papers, timed to late 2022, visible in the left tail rather than only in mean team size. Matsui runs this on the full OpenAlex snapshot with four venue filters, monthly and annual windows, donut and event-study timing checks, composition-adjusted author probabilities, history conditioning on never-recent-solo authors, balanced venue panels, and a SPECTER2 content DiD that nets out field-wide drift. That package is careful and transparent.\n\nWhat works: treating the solo tail as a probe of substitution is a clean design idea. The break is not just new entrants or field mix; it shows up among authors who previously only coauthored, including never-solo authors, and among seniors as well as juniors. Solo papers stay near authors’ prior collaborative content, narrow in breadth, and tilt toward computational topics—exactly the signature you’d want if LLMs are absorbing execution work. Field ordering (stronger in CS/math/psych/engineering-adjacent work, flat in chemistry/physics/lab-heavy fields) lines up with substitutability and with independent LLM-uptake measures. The paper is explicit that this is correlational and a dating convention, not a clean causal design. Mean author count also flattens, so the left-tail story is not an artifact of the metric alone.\n\nSoft spots, in proportion: the load-bearing step is still the substitution interpretation of the ChatGPT timestamp. SI D pushes against pure pandemic reversion via field ordering; SI L keeps a positive but attenuated break on continuously observed venues (peer-reviewed Δβ about +0.75 vs +1.72 full sample); SI M donuts and event studies support a 2022–23 trend change. Those help, but OpenAlex’s MAG transition and 2023 author-disambiguation revision sit inside the window and can move authorship metadata within venues. Engineering’s large share break is compositional rather than within-author, which the paper flags. No quality assessment of the recovered solos, and preprint excess could mix genuine substitution with low-cost solo output. None of that erases the descriptive break; it caps how hard you can lean on “AI substituted for coauthors.”\n\nThis is for science-of-science and AI-labor readers who care about authorship, credit, and division of cognitive labor. Methods and SI are thorough enough that a serious referee can engage. I would send it to peer review; the measurement contribution is real even if the causal reading stays provisional.","headline":"Solid large-scale descriptive break in the solo-authorship tail around late 2022, carefully multi-checked; the LLM-substitution reading is correlational and only partly bounded against database and multi-shock confounds.","tokens_in":22966,"tokens_out":618,"would_cite":true,"duration_ms":10209,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The long decline in solo-authored science papers halted and partially reversed after ChatGPT's late-2022 release.","keywords":["Science of science","Scientific collaboration","Authorship","Large language models","Artificial intelligence","Division of labor","Team science","Solo authorship"],"falsifier":"If independent measures of LLM adoption by field and author, or a later window after journal policies and indexing stabilize, show no corresponding solo-share break once venue composition, author disambiguation, and preprint volume are held fixed, the substitution reading would fail.","tokens_in":22753,"feed_emoji":"🤖","tokens_out":575,"duration_ms":12728,"temperature":0.7,"pith_summary":"Science has long shifted from solo work toward larger teams that divide cognitive labor. This paper asks whether generative AI extends that trend or instead lets researchers complete some tasks alone that once required coauthors. Using hundreds of millions of works across 26 fields, it shows the decades-long fall in solo authorship stopped and partially rebounded right after ChatGPT became public. The rebound is strongest where coauthor tasks are easier to automate, appears among authors who previously published only with others, and produces solo papers that stay near those authors' prior collaborative topics while narrowing and tilting computational. Solo papers without credited human coauthors thus become a visible probe of which research labor AI can replace, pointing to a reconfiguration inside papers rather than simply bigger or smaller teams.","feed_headline":"Solo science papers rebounded after ChatGPT arrived","feed_subtitle":"The decades-long drop halted where coauthor tasks are easiest for AI to take over","key_machinery":"The solo-authored left tail of the author-count distribution, treated as an observable probe of labor substitution: a solo paper is work completed without credited human coauthors, so changes in its share, who produces it, and what it contains mark the boundary of tasks generative AI can take over.","core_discovery":"The decades-long decline in the share of solo-authored papers halted and partially reversed around ChatGPT's public release in late 2022. The break is broad across most fields and publication filters, is not explained by new entrants or field-mix changes, and concentrates among authors who had recently or never published alone. Recovered solo papers stay close to the authors' own earlier coauthored content, contract in breadth, and shift toward computational topics, consistent with generative AI substituting for execution labor that coauthors once supplied.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Solo science authorship rebounds after ChatGPT release","Decades-long solo-author decline halts with generative AI","ChatGPT reverses drop in solo papers where tasks are replaceable","Solo papers recover among prior collaborators after late 2022","Generative AI substitutes for coauthors in solo research rebound"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the shared timing of the break at ChatGPT's public release, together with the field ordering by task substitutability, can be read as evidence of LLM substitution rather than concurrent confounds such as post-pandemic publishing shifts or database coverage changes.","fun_headline_variants_meta":{"raw":{"variants":["Solo science authorship rebounds after ChatGPT release","Decades-long solo-author decline halts with generative AI","ChatGPT reverses drop in solo papers where tasks are replaceable","Solo papers recover among prior collaborators after late 2022","Generative AI substitutes for coauthors in solo research rebound"]},"model":"grok-4.5","effort":"low","cost_usd":0.003836,"raw_usage":{"total_tokens":1236,"prompt_tokens":847,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":38360000,"prompt_tokens_details":{"text_tokens":847,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":305,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":847,"tokens_out":84,"duration_ms":4764,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T09:18:18.619949+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If independent measures of LLM adoption by field and author, or a later window after journal policies and indexing stabilize, show no corresponding solo-share break once venue composition, author disambiguation, and preprint volume are held fixed, the substitution reading would fail.","supporting_citations":[],"review_version":1}