{"id":"00954f0e-679a-4064-9abb-d1b0b04b1f91","arxiv_id":"2505.12666","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Researchers outside core AI fields became markedly more application-oriented, transdisciplinary, and socially accountable in their LLM-era papers, while AI insiders mainly responded by diversifying collaborations.","lead":"Using OpenAlex publication records, this study compares how computer science insiders and domain outsiders changed their research after publishing LLM-related work. It finds outsiders shifted toward applied, interdisciplinary, socially accountable science, while insiders mainly broadened their collaboration networks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The GPT-3.5 few-shot evaluator is trained on examples that already embody the predicted Mode 1-to-Mode 2 shift, and Section 5.4 concedes no human validation, so the observed outsider/insider differences may be classifier artifacts rather than changes in knowledge production.","rationale":"The reader's causal-inference concern about the matched-pair design is real, but the more load-bearing issue is the outcome measure itself. Even a perfectly controlled before/after design would not rescue the central claim if the dependent variables are produced by a classifier whose only few-shot examples already encode the predicted shift and whose validity is disclosed as unexamined in Section 5.4. The bibliometric validations do not remove this concern because they use the same LLM-vs-pre-LLM contrast and keyword/concept proxies that are mechanically correlated with LLM-ness. The evaluation-criteria inconsistency in Section 4.5 is a visible symptom of this unreliability. The proposed human-annotation check would settle whether the observed patterns are real; until it is run, the CONDITIONAL verdict remains appropriate rather than moving to REJECT, because the descriptive patterns could survive better measurement.","tokens_in":18821,"tokens_out":6383,"duration_ms":69704,"concrete_test":"Have two human annotators blind to period independently code a stratified random sample of 300 abstracts (150 pre-LLM, 150 LLM-era, balanced insider/outsider) using the paper's five-dimension scheme; compute agreement (e.g., Cohen's kappa) with the GPT-3.5 labels and compare human-annotated Mode 2 rates between pre-LLM and LLM abstracts. If GPT-3.5 agrees with humans at chance level, or if humans do not show the same outsider/insider differential shifts, the reported patterns are measurement artifacts rather than knowledge-production changes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3.3 constructs the GPT-3.5 prompt with exactly two few-shot pairs: a pre-LLM abstract labeled Mode 1 on all five dimensions and an LLM abstract labeled Mode 2 on four dimensions, with rationales that assert the shift. Because the only labeled examples instantiate the paper's hypothesis, the zero-temperature classifier is primed to label LLM-era abstracts as Mode 2 and pre-LLM abstracts as Mode 1; the vocabulary of LLM abstracts (evaluation, benchmark, bias, fairness, accuracy) lines up with the Mode 2 definitions in Table 1. The bibliometric 'validation' in Section 3.3.4 is exposed to the same confound: LLM papers are more likely to be tagged with multiple OpenAlex disciplines and to contain evaluation/accountability keywords for reasons unrelated to Mode 2 knowledge production. Section 5.4 explicitly concedes that the model is not prompted to justify its classifications and that human evaluation has not been performed. The direct contradiction in Section 4.5 (insiders rise from 0.15 to 0.27 on evaluation criteria while outsiders rise only from 0.12 to 0.16, yet the text claims the opposite) further shows the fragility of these measurements. The central claim that outsiders shift toward Mode 2 and insiders do not is therefore not grounded until the classifier's construct validity is established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how researchers' positions relative to LLM development shape changes in their scientific knowledge production after they begin publishing LLM-related work. Using OpenAlex, the authors identify 7,106 researchers (2,998 insiders in computer science/AI/NLP and 4,108 outsiders from other fields), match each researcher's LLM-era paper to their most similar pre-LLM paper via sentence embeddings, and classify abstracts along five Mode 1/Mode 2 dimensions using GPT-3.5 few-shot prompting. The paper reports that outsiders shift toward application-focused, transdisciplinary, socially accountable, and evaluation-oriented research, while insiders mainly increase collaboration heterogeneity. These findings are positioned as evidence that LLMs catalyze both innovation and reorganization in scientific communities, with implications for CSCW research and design.","tokens_in":19148,"tokens_out":5314,"duration_ms":59263,"significance":"If the measurement approach is valid, the paper makes a useful empirical contribution: it scales the insider/outsider and Mode 1/Mode 2 frameworks to a large bibliometric corpus, uses a within-researcher matched-pair design, and attempts to triangulate LLM classifications with bibliometric indicators. The specific, falsifiable predictions about differential shifts across five dimensions are well suited to CSCW debates about AI-mediated knowledge work. However, the central quantitative claims currently rest on an unvalidated LLM classifier and a before/after design without a control group, so the paper's significance will depend on whether these validity concerns can be addressed with additional analysis.","major_comments":[{"comment":"The primary outcome measure is GPT-3.5 abstract classification, but the paper concedes in Section 5.4 that the model is not prompted to justify its classifications and that no human evaluation has been performed. Because the two few-shot examples in Section 3.3.3 were hand-picked from the dataset to illustrate a Mode 1-to-Mode 2 shift, the zero-temperature classifier may be primed toward the paper's hypothesis. The bibliometric validation in Section 3.3.4 does not resolve this for dimensions whose keyword lists overlap with ordinary LLM vocabulary (e.g., 'evaluation', 'benchmark', 'accuracy'), and the paper does not state whether '0 = Not enough information' labels are included in the denominators of the reported proportions. The authors should add a human-annotated gold-standard set with inter-annotator agreement, report per-dimension classifier accuracy, and compare against a neutral or reversed-label few-shot prompt.","section":"Section 3.3.3 and Section 5.4"},{"comment":"The text states that the increase in novel quality control is more substantial among outsiders, but the reported model probabilities show the opposite: outsiders move from 0.12 to 0.16 (+0.04) while insiders move from 0.15 to 0.27 (+0.12). This direct numerical contradiction undermines the Section 5.1 summary claim that outsiders are pushing toward 'novel approaches to quality control' more strongly than insiders, and it must be corrected before the evaluation-criteria dimension can be interpreted.","section":"Section 4.5"},{"comment":"All reported shifts are point estimates without confidence intervals, standard errors, or significance tests, despite the Conclusion's use of the phrase 'significant shifts.' Because the design is a within-researcher matched comparison, paired tests (e.g., McNemar tests or bootstrap paired differences) are feasible and should be reported for each of the five dimensions and for the insider/outsider contrasts. Without such uncertainty quantification, the reader cannot assess whether differences of 0.04 to 0.12 are meaningful or simply artifacts of the large sample.","section":"Section 4 (all subsections)"},{"comment":"The abstract and Section 5.1 attribute the observed changes to LLM adoption ('LLMs catalyze both innovation and reorganization'), but the matched-pair design in Section 3.3.2 compares each researcher's LLM-era paper with their own most similar pre-LLM paper and includes no control group of researchers who did not publish LLM-related work. Secular trends toward applied, interdisciplinary, and accountability-oriented science could produce the same before/after shifts. The causal wording should be softened to associational language, or the analysis should be supplemented with a difference-in-differences design using non-LLM papers as controls.","section":"Abstract and Section 3.3.2"}],"minor_comments":[{"comment":"The phrase 'scientistic practice' appears to be a typo for 'scientific practice.'","section":"Section 2.2.1"},{"comment":"The word 'heterogenous' should be 'heterogeneous.'","section":"Section 5.1"},{"comment":"The ACM Reference Format block still contains placeholder text ('2018', 'Conference acronym ’XX'), and several references include access dates or incomplete metadata; these should be cleaned before submission.","section":"Reference formatting"},{"comment":"The text says outsiders 'become notably more application-driven compared to insiders,' but only outsider values are reported in the prose for Figure 3a; please report the insider values as well or refer explicitly to the figure for both groups.","section":"Section 4.1"},{"comment":"The capitalization of the modes is inconsistent ('Mode 1' vs 'mode 1', 'Mode 2' vs 'mode 2'); please standardize, especially in Table 1 and the prompt in Section 3.3.3.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a timely empirical study with a plausible framework, but the central measurement is currently unvalidated and one reported result (Section 4.5) directly contradicts its own numbers. I would ask for human validation of the classifier, statistical inference, and a corrected evaluation-criteria analysis before considering acceptance. The lack of a control group is also a serious limitation for the causal language in the abstract, though it could be addressed by rewording or by adding a difference-in-differences analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about arXiv:2505.12666. First, it does something genuinely new at scale: it pairs 7,106 researchers' LLM-era papers with their own matched pre-LLM papers and scores both on five Mode 1/Mode 2 dimensions, using OpenAlex and a GPT-3.5 classifier with bibliometric checks. Second, the central measurement is not yet solid enough to support the claims. The few-shot examples used to prompt the classifier are hand-picked from the dataset to already show a Mode 1 to Mode 2 shift, and Section 5.4 concedes there is no human validation. The abstract's \"catalyze\" language is not backed by any control group or adjustment for secular trends.\n\nCredit where due: the insider/outsider and Mode 1/Mode 2 frameworks are borrowed, but the joint operationalization is new. The matching design, the attempt to validate with keyword and discipline counts, and the discussion of CSCW implications (mediator roles, domain-specific tools) are thoughtful. The workflow is described clearly enough to reproduce. For a descriptive study, the patterns reported for application focus, transdisciplinarity, and accountability are plausible.\n\nThe soft spots are real and central. (1) The classifier is primed: the only labeled examples are one pre-LLM abstract called Mode 1 everywhere and one LLM abstract called Mode 2 on four dimensions. That does not establish construct validity. The keyword validations share the same confound, since LLM papers naturally mention benchmarks, evaluation, bias, and so on. (2) There are no confidence intervals or significance tests anywhere in Section 4; the reported differences are small (e.g., 0.15 to 0.32 vs 0.19 to 0.29) and may be noise. (3) Section 4.5 contains a direct internal contradiction: the text says outsiders show a more substantial increase in broader evaluation criteria, but the numbers given show insiders rising from 0.15 to 0.27 while outsiders rise from 0.12 to 0.16. That is the opposite of the stated claim. (4) The matched-pair before/after design has no control group, so the observed shifts could reflect general trends toward applied, interdisciplinary science rather than LLM adoption. The authors disclose some limitations in 5.4, which is honest, but disclosure does not fix the mismatch between evidence and the causal framing.\n\nWho is this for? Researchers in CSCW and science-of-science who want a large-scale descriptive snapshot and a cautionary example of LLM-as-evaluator pitfalls. It does not yet deserve to be cited as evidence for differential Mode 2 shifts. It does deserve a serious referee: the question is important, the data work is non-trivial, and the flaws are addressable in revision (human-annotated validation, control group, corrected Section 4.5, softened causal language).\n\nMy recommendation: send to peer review with expectations of major revision, not desk reject.","headline":"A large-scale descriptive comparison of insider/outsider LLM adoption that reads well but rests on an unvalidated classifier and a self-contradictory evaluation section.","tokens_in":19629,"tokens_out":3358,"would_cite":false,"duration_ms":32121,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models are shifting outsiders toward applied, transdisciplinary, accountable research while pushing insiders toward broader collaborations.","keywords":["large language models","knowledge production","insider-outsider","Mode 2 science","research collaboration","bibliometrics","few-shot classification","CSCW"],"falsifier":"Re-running the same before/after matched-pair measurement on researchers whose paired papers do not involve LLMs (for example, a comparison built around another recent topic shift) would settle the attribution: if the same shifts toward application, transdisciplinarity, and accountability appear, the pattern is a secular trend rather than an LLM effect.","tokens_in":18598,"feed_emoji":"🤖","tokens_out":10661,"duration_ms":99218,"temperature":0.7,"pith_summary":"The paper sets out to show that large language models are not just productivity tools but catalysts that change the direction of scientific knowledge production, and that the change depends on a researcher's position as insider or outsider. On a matched sample of 7,106 researchers who published LLM-related work between 2023 and 2025, it compares each LLM-era paper with the author's most similar pre-LLM paper. The paper finds outsiders shifting toward application-focused, transdisciplinary, socially accountable research with broader evaluation language, while insiders respond mainly by diversifying their collaboration networks. If these patterns hold, they give the computer-supported cooperative work field concrete targets: designing domain-specific tools for outsider-led AI research and studying the mediator roles that transdisciplinary AI collaborations will need.","feed_headline":"LLMs push outsiders to applied science, insiders to new collaborations","feed_subtitle":"Of 7,106 researchers, outsiders shift toward applied science; insiders widen collaborations.","key_machinery":"The carrying mechanism is an evaluation workflow that combines the insider-outsider lens of radical innovation with the Mode 1/Mode 2 knowledge-production framework. Mode 1 knowledge is academic, disciplinary, homogeneous, autonomous, and peer-reviewed; Mode 2 is application-oriented, transdisciplinary, heterogeneous, socially accountable, and judged by broader stakeholders. The workflow classifies each of 7,106 researchers as insider or outsider by their dominant pre-2023 publication discipline, pairs each researcher's LLM-era paper with their most semantically similar pre-LLM paper via sentence embeddings, then uses few-shot prompting of a commercial language model to score both abstracts on the five dimensions (with a \"not enough information\" category). Bibliometric proxies for each dimension validate the model output. This design gives the paper its before/after contrast on similar topics and is what turns the qualitative insider-outsider and Mode 1/Mode 2 theories into measurable shifts.","core_discovery":"The paper's central claim is that LLM adoption reshapes knowledge production asymmetrically: researchers outside AI/NLP (outsiders) use LLMs as an entry point to Mode 2 knowledge production, while researchers inside AI/NLP (insiders) respond mainly by restructuring collaboration. On the five Mode 1/Mode 2 dimensions, the few-shot classification shows outsiders moving from 0.75 to 0.81 on application orientation, from 0.15 to 0.32 on transdisciplinarity, from 0.59 to 0.68 on social accountability, and from 0.12 to 0.16 on novel quality-control language, while insiders' collaboration-heterogeneity score rises from 0.26 to 0.31. Bibliometric proxies—application keywords, mean number of disciplines, institution-type diversity, and accountability keywords—point the same way, with the evaluation dimension the clear exception, where keyword and model results diverge. The paper reads the pattern as outsiders pushing toward more applied, transdisciplinary, accountable science with new quality-control approaches, and insiders selectively adapting through more institutionally diverse partnerships to hold their position in a field now shared with LLM-enabled outsiders.","pith_inferences":["A natural extension the paper does not run: applying the same matched-pair design to earlier tool shifts, such as cloud computing or pre-trained embeddings, would show whether the outsider push toward Mode 2 is specific to LLMs or generic to any major toolkit.","The paper's own interpretation implies a testable prediction: as LLMs lower programming barriers, outsiders' need for technical co-authors should decline, so their collaboration heterogeneity should flatten while insiders' continues to rise.","If outsider LLM work is genuinely Mode 2, a further untested consequence is that outsider LLM papers should diffuse across disciplines faster than insider LLM papers, visible in citation and policy-document traces.","The paper's disclosed lack of human validation for its language-model classifications means a human-rating study on a random abstract sample is the immediate next check before the dimension-level numbers are taken at face value."],"forward_implications":["Outsider-led LLM research will keep moving toward application and social relevance, so domain-specific tools and workflows will matter more than generic research assistants.","Insiders' most visible adaptive response will be institutionally diverse collaborations, especially with healthcare, industry, and government partners, rather than deeper epistemic shifts.","Aggregate measures of LLM impact that pool insiders and outsiders will blur the two distinct response paths and should be reported separately.","LLM-enabled transdisciplinary collaboration will increasingly require active mediation across domain language, norms, and evaluation standards, a role the field is positioned to study."],"supporting_citations":[{"why":"Supplies the Mode 1/Mode 2 knowledge-production framework that defines the five compared dimensions.","marker":"[19]"},{"why":"Adapts and operationalizes the Mode 1/Mode 2 framework into the five-dimension comparison the workflow classifies.","marker":"[21]"},{"why":"Provides the insider-outsider theory of radical innovation that motivates separating researchers by their position relative to LLM development.","marker":"[22]"},{"why":"Extends the insider-outsider account by arguing outsiders are freer from field rules, grounding the prediction that outsiders drive change.","marker":"[63]"},{"why":"Supplies the large scholarly knowledge graph from which the 7,106 researchers, their disciplines, institutions, and matched papers are drawn.","marker":"[52]"},{"why":"Provides the sentence-embedding method used to pair each LLM-era paper with the researcher's most similar pre-LLM paper.","marker":"[55]"},{"why":"Establishes few-shot prompting, the technique that lets the commercial language model classify abstracts on the five dimensions.","marker":"[8]"}],"fun_headline_variants":["Outsiders turn applied, insiders rewire collaborations under LLMs","LLMs reshape science: outsiders apply, insiders partner up","Insider-outsider split: LLMs drive applied shift and network changes","LLM era: outsiders go applied and transdisciplinary, insiders diversify","LLMs shift outsiders to applied science, insiders to diverse ties"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the measured shift from a researcher's most similar pre-LLM paper to their LLM-era paper is caused by LLM adoption rather than by secular trends toward applied and interdisciplinary science, since there is no control group of researchers who never published LLM work.","fun_headline_variants_meta":{"raw":{"variants":["Outsiders turn applied, insiders rewire collaborations under LLMs","LLMs reshape science: outsiders apply, insiders partner up","Insider-outsider split: LLMs drive applied shift and network changes","LLM era: outsiders go applied and transdisciplinary, insiders diversify","LLMs shift outsiders to applied science, insiders to diverse ties"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001164,"raw_usage":{"total_tokens":4795,"prompt_tokens":902,"completion_tokens":3893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":3800}},"tokens_in":518,"tokens_out":3893,"duration_ms":26570,"temperature":1.0,"reasoning_tokens":3800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:29:11.801448+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the same before/after matched-pair measurement on researchers whose paired papers do not involve LLMs (for example, a comparison built around another recent topic shift) would settle the attribution: if the same shifts toward application, transdisciplinarity, and accountability appear, the pattern is a secular trend rather than an LLM effect.","supporting_citations":[{"cited_title":"1994.The new production of knowledge: The dynamics of science and research in contemporary societies","cited_arxiv_id":null,"evidence_quote":"Supplies the Mode 1/Mode 2 knowledge-production framework that defines the five compared dimensions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Adapts and operationalizes the Mode 1/Mode 2 framework into the five-dimension comparison the workflow classifies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the insider-outsider theory of radical innovation that motivates separating researchers by their position relative to LLM development."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the insider-outsider account by arguing outsiders are freer from field rules, grounding the prediction that outsiders drive change."}],"review_version":1}