{"id":"82d4d48a-ab11-4408-967b-4604b28989f6","arxiv_id":"2606.05118","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"AI publications show 5.5-10.2 percentage point higher probability of top-decile creativity than non-AI ones, with tool-oriented AI linked to recombinant novelty gains and adaptation-oriented AI linked to object novelty gains.","lead":"This paper analyzes over one million publications to compare creativity metrics between AI-using and non-AI scientific work. A smart generalist should read it to understand how different ways of adopting AI may shape research evaluation and funding priorities.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Validity of OpenAlex-based classifications for AI publications and creativity metrics is the load-bearing assumption","rationale":"The reader's weakest_assumption correctly isolates the measurement step that must hold for any downstream association to be interpretable. Because the full methods for labeling and metric construction are not visible in the supplied abstract, the observational claim remains unassessable; no stronger internal inconsistency is detectable from the given material.","tokens_in":1754,"tokens_out":337,"duration_ms":28786,"concrete_test":"Re-run the main specification after replacing the paper's AI classifier with an independent rule (publications whose abstract contains at least one of: 'deep learning', 'neural network', 'transformer', 'large language model') and recompute the top-decile probability gap; if the gap falls below 3 pp or changes sign, the original classification drives the result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline association (5.5–10.2 pp higher top-decile creativity probability for AI papers) requires that (a) AI vs. non-AI labeling and (b) recombinant/object novelty plus citation-impact deciles are measured without substantial error or field-specific bias. The abstract gives no detail on the exact OpenAlex query, keyword/ML classifier, or how recombinant novelty is operationalized (e.g., new reference combinations, concept co-occurrence). If either step correlates with unobserved factors such as subfield, team size, or data availability, the reported difference can be produced by measurement artifacts rather than genuine creativity differences. The further split into tool-oriented vs. adaptation-oriented modes compounds the classification risk.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript analyzes over one million OpenAlex publications to assess whether AI adoption advances scientific creativity. It reports that AI papers have a 5.5 to 10.2 percentage point higher probability of ranking in the top creativity decile across measures of recombinant novelty, object novelty, and citation impact. The analysis also identifies heterogeneity by AI research mode, with tool-oriented approaches linked to recombinant creativity gains and adaptation-oriented approaches to object novelty gains.","tokens_in":1887,"tokens_out":510,"duration_ms":33095,"significance":"Should the measurement of AI adoption and creativity dimensions prove robust, the results would suggest that AI contributes to science via multiple distinct pathways rather than a uniform mechanism. This has potential implications for research assessment frameworks and science policy, emphasizing the need to differentiate between forms of creativity. The large-scale data approach is a positive if accompanied by transparent methods.","major_comments":[{"comment":"Methods section: The classification of publications as AI or non-AI is central to the main claim but the abstract provides no information on the OpenAlex query, keyword set, or machine learning classifier employed; this must be detailed to allow assessment of potential selection bias or field-specific effects.","section":"Methods"},{"comment":"Methods section: Recombinant novelty and object novelty metrics are not described (e.g., whether recombinant novelty uses new reference combinations or concept co-occurrences); without this, it is impossible to evaluate whether the 5.5-10.2 pp difference reflects genuine creativity or measurement artifacts.","section":"Methods"},{"comment":"Results section: The split into tool-oriented vs. adaptation-oriented AI research requires explicit, operational definitions and robustness checks; the differential associations with creativity types are load-bearing for the heterogeneity conclusion but vulnerable to classification error.","section":"Results"},{"comment":"Discussion section: Potential confounders such as subfield, team size, or data availability are not mentioned as controlled; if these correlate with AI adoption, they could explain the observed differences without implying a causal advance in creativity.","section":"Discussion"}],"minor_comments":[{"comment":"Abstract: The abstract mentions 'over one million publications' but does not specify the time period or exact sample construction criteria.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an observational study in the science of science; ensure it fits the journal's scope if the journal focuses on core computer science rather than metascience."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for these constructive comments, which help improve the transparency and robustness of our analysis. We address each major point below and have revised the manuscript to incorporate additional methodological detail and controls.","responses":[{"response":"We agree that full details on the AI/non-AI classification are essential. The Methods section of the original manuscript contained a high-level description, but we have now expanded it substantially to report the precise OpenAlex query, the keyword list used for initial filtering, the machine-learning classifier architecture and training procedure, and out-of-sample performance metrics. These additions allow direct evaluation of selection bias and field coverage.","revision_made":"yes","referee_comment":"[Methods] Methods section: The classification of publications as AI or non-AI is central to the main claim but the abstract provides no information on the OpenAlex query, keyword set, or machine learning classifier employed; this must be detailed to allow assessment of potential selection bias or field-specific effects."},{"response":"We accept that the operational definitions were insufficiently explicit. We have added a dedicated subsection in Methods that specifies recombinant novelty as the share of previously unobserved reference-pair combinations (following the standard combinatorial approach) and object novelty as the introduction of previously unseen scientific concepts or entities. We also include the exact formulas, data sources, and validation checks against alternative concept-based measures to demonstrate that the reported differences are not artifacts of a single operationalization.","revision_made":"yes","referee_comment":"[Methods] Methods section: Recombinant novelty and object novelty metrics are not described (e.g., whether recombinant novelty uses new reference combinations or concept co-occurrences); without this, it is impossible to evaluate whether the 5.5-10.2 pp difference reflects genuine creativity or measurement artifacts."},{"response":"We have inserted explicit, replicable definitions in the revised Results section: tool-oriented papers are those that apply off-the-shelf AI models without modification, while adaptation-oriented papers are those that fine-tune, extend, or re-architect AI components for the domain problem. We further report two alternative classification schemes (keyword-based and embedding-based) together with robustness tables showing that the differential associations with recombinant versus object novelty remain stable across these specifications.","revision_made":"yes","referee_comment":"[Results] Results section: The split into tool-oriented vs. adaptation-oriented AI research requires explicit, operational definitions and robustness checks; the differential associations with creativity types are load-bearing for the heterogeneity conclusion but vulnerable to classification error."},{"response":"Subfield fixed effects were already included in all main specifications; we have now made this explicit in both the Methods and Discussion sections. In the revision we additionally control for team size (log number of authors) and include a proxy for data availability (whether the paper mentions a public dataset). After these controls the AI coefficient remains positive and significant, though we acknowledge that observational data cannot fully rule out residual confounding and have updated the Discussion to reflect this limitation.","revision_made":"partial","referee_comment":"[Discussion] Discussion section: Potential confounders such as subfield, team size, or data availability are not mentioned as controlled; if these correlate with AI adoption, they could explain the observed differences without implying a causal advance in creativity."}],"tokens_in":1457,"tokens_out":706,"duration_ms":28264,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is the reported 5.5–10.2 percentage point lift in top-decile creativity for AI papers, split by whether the work applies existing AI tools or adapts models to new domains. Tool-oriented work ties more to recombinant novelty while adaptation-oriented work ties more to object novelty. That split is the clearest addition beyond earlier AI-in-science counts.\n\nThe scale helps: over a million OpenAlex records, separate short-run and long-run citation measures, and an attempt to separate two creativity dimensions. Those choices give the heterogeneity some traction.\n\nThe soft spot is exactly the one the stress test flags. The abstract and available text give no concrete description of the AI-publication classifier, the keyword or ML rules used, or how recombinant and object novelty are computed from references or concepts. If those steps pick up field effects, team size, or data availability instead of genuine differences, the percentage-point gaps can appear without any real creativity shift. The observational design leaves the same opening for confounding.\n\nThis is for readers who track scientometrics or science policy and want empirical handles on AI modes. It is worth sending to referees because the question is live and the dataset size is real, provided the methods section survives close review on classification validity and robustness.","headline":"AI papers show higher top-decile creativity in OpenAlex data with mode-specific novelty patterns, but the result hinges on untested classification steps.","tokens_in":2365,"tokens_out":330,"would_cite":false,"duration_ms":24274,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"AI publications are 5.5 to 10.2 percentage points more likely than non-AI ones to rank in the top creativity decile.","keywords":["artificial intelligence","scientific creativity","recombinant novelty","object novelty","research modes","citation impact","OpenAlex"],"falsifier":"Repeating the decile-ranking analysis after replacing the OpenAlex AI labels or the novelty/impact formulas with independent alternative definitions and finding that the 5.5–10.2 point gap disappears.","tokens_in":2641,"feed_emoji":"","tokens_out":671,"duration_ms":22028,"temperature":0.7,"pith_summary":"The paper asks whether adopting artificial intelligence increases scientific creativity and finds that it does, but not uniformly. Using more than one million OpenAlex publications, the authors measure creativity through recombinant novelty, object novelty, and citation impact, then compare AI and non-AI papers. They report that AI papers reach the top creativity decile at substantially higher rates. The gains differ by research mode: applying existing AI tools to new domains produces the strongest recombinant-novelty boost, while adapting AI models to domain problems produces the strongest object-novelty boost. The results imply that AI advances science through separate creative routes rather than a single mechanism.","feed_headline":"AI papers reach top creativity decile 5–10 points more often","feed_subtitle":"Million-publication study shows tool use and model adaptation each raise a different kind of novelty.","key_machinery":"The distinction between tool-oriented AI research (applying existing models) and adaptation-oriented AI research (modifying models for domain problems) and their separate links to recombinant novelty versus object novelty.","core_discovery":"AI publications are significantly more likely to achieve top-decile creativity relative to non-AI publications, with 5.5 to 10.2 percentage point higher likelihood to rank in the top creativity decile. Tool-oriented AI research is associated with the largest gains in recombinant-based creativity, while adaptation-oriented AI research is associated with relatively higher object-based creativity. These findings show that AI does not advance science through a single mechanism but through structurally distinct creative pathways that depend on how AI is incorporated into the research process.","pith_inferences":["Policy that funds only one mode of AI adoption may miss gains available from the other mode.","Citation-based impact alone may understate the creativity contribution of adaptation-oriented AI work.","Future studies could test whether the same pattern holds when novelty is measured by expert panels instead of text-based indicators."],"forward_implications":["Tool-oriented AI use produces the largest increase in recombinant novelty.","Adaptation-oriented AI use produces the largest increase in object novelty.","Different ways of folding AI into research generate different types of scientific contribution.","Assessment frameworks should separate recombinant from conceptual forms of creativity when evaluating AI-assisted work."],"fun_headline_variants":["AI papers more often in top creativity decile by 5-10 points","Tool-oriented AI shows largest recombinant novelty gains","Adaptation-oriented AI shows higher object novelty gains","AI creativity advances through distinct pathways by mode"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The rules used to label publications as AI-related and to score recombinant novelty, object novelty, and citation impact in OpenAlex correctly capture real AI use and genuine creativity without large measurement error or bias.","fun_headline_variants_meta":{"raw":{"variants":["AI papers more often in top creativity decile by 5-10 points","Tool-oriented AI shows largest recombinant novelty gains","Adaptation-oriented AI shows higher object novelty gains","AI creativity advances through distinct pathways by mode"]},"model":"grok-4.3","cost_usd":0.007919,"raw_usage":{"total_tokens":3550,"prompt_tokens":711,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":79190500,"prompt_tokens_details":{"text_tokens":711,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2778,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":711,"tokens_out":61,"duration_ms":39738,"temperature":1.0,"reasoning_tokens":2778,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T03:27:01.547238+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the decile-ranking analysis after replacing the OpenAlex AI labels or the novelty/impact formulas with independent alternative definitions and finding that the 5.5–10.2 point gap disappears.","supporting_citations":[],"review_version":1}