Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Adapting to LLMs: How Insiders and Outsiders Reshape Scientific Knowledge Production

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Large language models are shifting outsiders toward applied, transdisciplinary, accountable research while pushing insiders toward broader collaborations.

desk verdict A large-scale descriptive comparison of insider/outsider LLM adoption that reads well but rests on an unvalidated classifier and a self-contradictory evaluation section. read the letter →

arxiv 2505.12666 v1 pith:MUTELB3B submitted 2025-05-19 cs.HC

classification cs.HC
keywords largelanguagemodelsknowledgeproductioninsider-outsiderMode2scienceresearchcollaborationbibliometricsfew-shotclassificationCSCW
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that large language models are not just productivity tools but catalysts that change the direction of scientific knowledge production, and that the change depends on a researcher's position as insider or outsider. On a matched sample of 7,106 researchers who published LLM-related work between 2023 and 2025, it compares each LLM-era paper with the author's most similar pre-LLM paper. The paper finds outsiders shifting toward application-focused, transdisciplinary, socially accountable research with broader evaluation language, while insiders respond mainly by diversifying their collaboration networks. If these patterns hold, they give the computer-supported cooperative work field concrete targets: designing domain-specific tools for outsider-led AI research and studying the mediator roles that transdisciplinary AI collaborations will need.

What carries the argument

The carrying mechanism is an evaluation workflow that combines the insider-outsider lens of radical innovation with the Mode 1/Mode 2 knowledge-production framework. Mode 1 knowledge is academic, disciplinary, homogeneous, autonomous, and peer-reviewed; Mode 2 is application-oriented, transdisciplinary, heterogeneous, socially accountable, and judged by broader stakeholders. The workflow classifies each of 7,106 researchers as insider or outsider by their dominant pre-2023 publication discipline, pairs each researcher's LLM-era paper with their most semantically similar pre-LLM paper via sentence embeddings, then uses few-shot prompting of a commercial language model to score both abstracts on the five dimensions (with a "not enough information" category). Bibliometric proxies for each dimension validate the model output. This design gives the paper its before/after contrast on similar topics and is what turns the qualitative insider-outsider and Mode 1/Mode 2 theories into measurable shifts.

What would settle it

Re-running the same before/after matched-pair measurement on researchers whose paired papers do not involve LLMs (for example, a comparison built around another recent topic shift) would settle the attribution: if the same shifts toward application, transdisciplinarity, and accountability appear, the pattern is a secular trend rather than an LLM effect.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM adoption reshapes knowledge production asymmetrically: researchers outside AI/NLP (outsiders) use LLMs as an entry point to Mode 2 knowledge production, while researchers inside AI/NLP (insiders) respond mainly by restructuring collaboration. On the five Mode 1/Mode 2 dimensions, the few-shot classification shows outsiders moving from 0.75 to 0.81 on application orientation, from 0.15 to 0.32 on transdisciplinarity, from 0.59 to 0.68 on social accountability, and from 0.12 to 0.16 on novel quality-control language, while insiders' collaboration-heterogeneity score rises from 0.26 to 0.31. Bibliometric proxies—application keywords, mean number of disciplines, institution-type diversity, and accountability keywords—point the same way, with the evaluation dimension the clear exception, where keyword and model results diverge. The paper reads the pattern as outsiders pushing toward more applied, transdisciplinary, accountable science with new quality-control approaches, and insiders selectively adapting through more institutionally diverse partnerships to hold their position in a field now shared with LLM-enabled outsiders.

Load-bearing premise

The load-bearing premise is that the measured shift from a researcher's most similar pre-LLM paper to their LLM-era paper is caused by LLM adoption rather than by secular trends toward applied and interdisciplinary science, since there is no control group of researchers who never published LLM work.

Editorial extensions

If this is right

  • Outsider-led LLM research will keep moving toward application and social relevance, so domain-specific tools and workflows will matter more than generic research assistants.
  • Insiders' most visible adaptive response will be institutionally diverse collaborations, especially with healthcare, industry, and government partners, rather than deeper epistemic shifts.
  • Aggregate measures of LLM impact that pool insiders and outsiders will blur the two distinct response paths and should be reported separately.
  • LLM-enabled transdisciplinary collaboration will increasingly require active mediation across domain language, norms, and evaluation standards, a role the field is positioned to study.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not run: applying the same matched-pair design to earlier tool shifts, such as cloud computing or pre-trained embeddings, would show whether the outsider push toward Mode 2 is specific to LLMs or generic to any major toolkit.
  • The paper's own interpretation implies a testable prediction: as LLMs lower programming barriers, outsiders' need for technical co-authors should decline, so their collaboration heterogeneity should flatten while insiders' continues to rise.
  • If outsider LLM work is genuinely Mode 2, a further untested consequence is that outsider LLM papers should diffuse across disciplines faster than insider LLM papers, visible in citation and policy-document traces.
  • The paper's disclosed lack of human validation for its language-model classifications means a human-rating study on a random abstract sample is the immediate next check before the dimension-level numbers are taken at face value.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies how researchers' positions relative to LLM development shape changes in their scientific knowledge production after they begin publishing LLM-related work. Using OpenAlex, the authors identify 7,106 researchers (2,998 insiders in computer science/AI/NLP and 4,108 outsiders from other fields), match each researcher's LLM-era paper to their most similar pre-LLM paper via sentence embeddings, and classify abstracts along five Mode 1/Mode 2 dimensions using GPT-3.5 few-shot prompting. The paper reports that outsiders shift toward application-focused, transdisciplinary, socially accountable, and evaluation-oriented research, while insiders mainly increase collaboration heterogeneity. These findings are positioned as evidence that LLMs catalyze both innovation and reorganization in scientific communities, with implications for CSCW research and design.

Significance. If the measurement approach is valid, the paper makes a useful empirical contribution: it scales the insider/outsider and Mode 1/Mode 2 frameworks to a large bibliometric corpus, uses a within-researcher matched-pair design, and attempts to triangulate LLM classifications with bibliometric indicators. The specific, falsifiable predictions about differential shifts across five dimensions are well suited to CSCW debates about AI-mediated knowledge work. However, the central quantitative claims currently rest on an unvalidated LLM classifier and a before/after design without a control group, so the paper's significance will depend on whether these validity concerns can be addressed with additional analysis.

major comments (4)
  1. [Section 3.3.3 and Section 5.4] The primary outcome measure is GPT-3.5 abstract classification, but the paper concedes in Section 5.4 that the model is not prompted to justify its classifications and that no human evaluation has been performed. Because the two few-shot examples in Section 3.3.3 were hand-picked from the dataset to illustrate a Mode 1-to-Mode 2 shift, the zero-temperature classifier may be primed toward the paper's hypothesis. The bibliometric validation in Section 3.3.4 does not resolve this for dimensions whose keyword lists overlap with ordinary LLM vocabulary (e.g., 'evaluation', 'benchmark', 'accuracy'), and the paper does not state whether '0 = Not enough information' labels are included in the denominators of the reported proportions. The authors should add a human-annotated gold-standard set with inter-annotator agreement, report per-dimension classifier accuracy, and compare against a neutral or reversed-label few-shot prompt.
  2. [Section 4.5] The text states that the increase in novel quality control is more substantial among outsiders, but the reported model probabilities show the opposite: outsiders move from 0.12 to 0.16 (+0.04) while insiders move from 0.15 to 0.27 (+0.12). This direct numerical contradiction undermines the Section 5.1 summary claim that outsiders are pushing toward 'novel approaches to quality control' more strongly than insiders, and it must be corrected before the evaluation-criteria dimension can be interpreted.
  3. [Section 4 (all subsections)] All reported shifts are point estimates without confidence intervals, standard errors, or significance tests, despite the Conclusion's use of the phrase 'significant shifts.' Because the design is a within-researcher matched comparison, paired tests (e.g., McNemar tests or bootstrap paired differences) are feasible and should be reported for each of the five dimensions and for the insider/outsider contrasts. Without such uncertainty quantification, the reader cannot assess whether differences of 0.04 to 0.12 are meaningful or simply artifacts of the large sample.
  4. [Abstract and Section 3.3.2] The abstract and Section 5.1 attribute the observed changes to LLM adoption ('LLMs catalyze both innovation and reorganization'), but the matched-pair design in Section 3.3.2 compares each researcher's LLM-era paper with their own most similar pre-LLM paper and includes no control group of researchers who did not publish LLM-related work. Secular trends toward applied, interdisciplinary, and accountability-oriented science could produce the same before/after shifts. The causal wording should be softened to associational language, or the analysis should be supplemented with a difference-in-differences design using non-LLM papers as controls.
minor comments (5)
  1. [Section 2.2.1] The phrase 'scientistic practice' appears to be a typo for 'scientific practice.'
  2. [Section 5.1] The word 'heterogenous' should be 'heterogeneous.'
  3. [Reference formatting] The ACM Reference Format block still contains placeholder text ('2018', 'Conference acronym ’XX'), and several references include access dates or incomplete metadata; these should be cleaned before submission.
  4. [Section 4.1] The text says outsiders 'become notably more application-driven compared to insiders,' but only outsider values are reported in the prose for Figure 3a; please report the insider values as well or refer explicitly to the figure for both groups.
  5. [Throughout] The capitalization of the modes is inconsistent ('Mode 1' vs 'mode 1', 'Mode 2' vs 'mode 2'); please standardize, especially in Table 1 and the prompt in Section 3.3.3.

Circularity Check

2 steps flagged · score 6.0 of 10

Outsider Mode-2 shift is partly constructed: the outsider definition plus LLM-paper selection guarantees transdisciplinarity growth, and the few-shot prompt's only labeled examples encode the expected Mode 1-to-Mode 2 shift.

  1. self definitional [Section 3.3.1 (insider/outsider definition) and Section 4.2 (transdisciplinarity result)]
    "We define insiders as researchers whose primary area of publication is in the field of Computer Science... If this primary discipline is computer science, the researcher is classified as an insider... otherwise, they are classified as an outsider. ... Regarding disciplinary orientation—specifically whether LLM-related research is becoming more interdisciplinary—our few-shot classification experiments reveal a notable increase in transdisciplinarity among outsiders’ papers. The proportion of transdisciplinary work rose from 0.15 in pre-LLM papers to 0.32 in LLM papers (Fig 4a - deep blue)."

    Outsiders are defined as researchers whose primary field is not computer science, while LLM papers are selected because their titles/abstracts contain 'large language model' or 'chatgpt'—computer-science/AI concepts. The matched pre-LLM paper is in the researcher's home field. Thus an outsider's LLM-era paper will almost necessarily carry both the home discipline and computer science in OpenAlex, mechanically inflating the reported rise in unique disciplines (2.5 to 3.1) and the model-based transdisciplinarity increase (0.15 to 0.32). The same selection logic pushes outsider LLM papers toward the applied, problem-driven language of Mode 2.

  2. fitted input called prediction [Section 3.3.3 (few-shot prompt) and Section 5.4 (limitations)]
    "To guide the model’s classification, we selected two pairs of papers from our dataset as the few-shot examples... We labeled the left example as Mode 1 knowledge production across all five dimensions... In contrast, the right example reflects four key dimensions for Mode 2 knowledge production. ... While we use GPT-3.5 classification to analyze knowledge production shifts, we currently do not prompt the model to explain its reasoning or provide evidence for its classifications."

    The only labeled examples supplied to GPT-3.5 are a pre-LLM abstract labeled all-Mode-1 and an LLM abstract labeled four-dimension Mode-2—exactly the paper's predicted before/after shift. These hand-picked examples effectively define the outcome variable for the classifier, and the resulting classifications are then presented as empirical evidence of the shift. Section 5.4 concedes that the model is not asked to justify classifications and that human evaluation has not been performed, so there is no external anchor for the labels. The bibliometric validation (Section 3.3.4) is exposed to a similar confound because LLM-era abstracts naturally contain the evaluation, accountability, and application keywords that Table 1 associates with Mode 2.

full rationale

The paper has real independent content: the collaboration-heterogeneity result (insiders rising more) is not forced by the insider/outsider definition, and the bibliometric checks for application keywords and accountability terms provide some external grounding. However, the central claim that outsiders shift toward Mode 2 is substantially circular. First, the transdisciplinarity and application-focus components are nearly built into the research design: outsiders are defined as non-CS researchers, and their LLM-era papers are selected because they concern a CS/AI topic, so cross-disciplinarity and application framing are expected by construction. Second, the few-shot evaluator is primed with a single Mode-1 pre-LLM example and a single Mode-2 LLM example, instantiating the paper's conclusion before any measurement occurs; the absence of human evaluation (admitted in Section 5.4) leaves no independent check. The Section 4.5 internal inconsistency (insiders rise from 0.15 to 0.27 while outsiders rise from 0.12 to 0.16, yet the text says outsiders show the more substantial increase) further underscores that these measurements are fragile, though that inconsistency is a correctness issue rather than circularity. Overall, the derivation chain partially reduces to its own inputs, meriting a score of 6.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central claims rest on the Mode 1/Mode 2 taxonomy, the validity of GPT-3.5 classifications without human labels, the matched-pair baseline, and the assumption that keyword presence identifies LLM adoption. The authors disclose several of these in Section 5.4, but none are independently verified.

assumptions (6)
  • domain assumption The Mode 1/Mode 2 knowledge-production framework (Gibbons et al., Hessels and van Lente) is a valid description of scientific knowledge production.
    Section 3.1 adopts this framework to define the five measured dimensions; if the framework is not valid for contemporary science, the measurement construct is questionable.
  • ad hoc to paper GPT-3.5-Turbo's abstract classifications along the five dimensions are valid outcome measures without human evaluation.
    Section 3.3.3 uses the model with temperature 0 and two hand-picked few-shot examples; Section 5.4 acknowledges that no human evaluation was performed.
  • ad hoc to paper The most similar pre-LLM paper by sentence embedding is a valid baseline for the same researcher's pre-LLM knowledge production.
    Section 3.3.2 selects one matching prior paper per researcher; this assumes the chosen paper is representative and that no other confounds explain the difference.
  • domain assumption Publishing a paper with 'large language model' or 'ChatGPT' in the title or abstract indicates LLM adoption by the authors.
    Section 3.2.1 uses this keyword filter; researchers who use LLMs without such keywords are invisible to the study.
  • ad hoc to paper Treating authors affiliated with the same institution as a single unique researcher is an acceptable disambiguation strategy.
    Section 3.2.2 makes this choice to compensate for OpenAlex disambiguation limits; it can merge distinct researchers and lose movers.
  • domain assumption Paper abstracts contain sufficient information to classify all five Mode 1/Mode 2 dimensions.
    Section 3.3.3 allows a label of 0 for insufficient information, but the paper does not report how often 0 was assigned, so the extent of missingness is unknown.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adapting to LLMs: How Insiders and Outsiders Reshape Scientific Knowledge Production." pith.science (2026). https://pith.science/paper/MUTELB3B

@misc{pith2026250512666,
  author       = {Pith},
  title        = {Pith review of: Adapting to LLMs: How Insiders and Outsiders Reshape Scientific Knowledge Production},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MUTELB3B}},
  note         = {Machine review of arXiv:2505.12666}
}
read the original abstract

CSCW has long examined how emerging technologies reshape the ways researchers collaborate and produce knowledge, with scientific knowledge production as a central area of focus. As AI becomes increasingly integrated into scientific research, understanding how researchers adapt to it reveals timely opportunities for CSCW research -- particularly in supporting new forms of collaboration, knowledge practices, and infrastructure in AI-driven science. This study quantifies LLM impacts on scientific knowledge production based on an evaluation workflow that combines an insider-outsider perspective with a knowledge production framework. Our findings reveal how LLMs catalyze both innovation and reorganization in scientific communities, offering insights into the broader transformation of knowledge production in the age of generative AI and sheds light on new research opportunities in CSCW.

Figures

Figures reproduced from arXiv: 2505.12666 by the authors.

Figure 1
Figure 1. The computational process of evaluating the impact of LLM in scientific knowledge production [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. The definition of insider and outsider 3.3.2 Matching similar pairs between LLM papers and pre-LLM papers. To assess the impact of LLMs on researchers’ knowledge production, we match each LLM-related paper with the most similar-topic paper authored by the same researcher in the five years prior to 2023. This allows us to compare the difference of knowledge production on similar topics before and after the rise of LL… view at source ↗
Figure 3
Figure 3. Comparison of application focus between pre-LLM and LLM papers for both insiders and outsiders. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comparison of disciplinary orientation between pre-LLM and LLM papers: (a) Model-based probability [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Comparison of collaboration patterns between pre-LLM and LLM papers for both insiders and outsiders. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Comparison of social accountability between pre-LLM and LLM papers for both insiders and outsiders. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Comparison of evaluation criteria between pre-LLM and LLM papers for both insiders and outsiders. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scientific exploration, collaboration and labor division in the large language model era

    cs.DL 2026-07 conditional novelty 6.0 of 10

    After 2022, scientists became more interdisciplinary and exploratory, collaborated across more fields, and reported more differentiated team roles, with the largest shifts among high-AI-writing, established, and non-E...

Reference graph

Works this paper leans on

67 extracted references · 26 canonical work pages · cited by 1 Pith paper

  1. [1]

    Mark S Ackerman, Juri Dachtera, Volkmar Pipek, and Volker Wulf. 2013. Sharing knowledge and expertise: The CSCW view of knowledge management. Computer supported cooperative work: CSCW: an international journal 22, 4-6 (21 Aug. 2013), 531–573. https://doi.org/10.1007/s10606-013-9192-8

  2. [2]

    Guilherme F C F Almeida, José Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de Araújo. 2024. Exploring the psychology of LLMs’ moral and legal reasoning. Artificial intelligence 333, 104145 (1 Aug. 2024), 104145. https: //doi.org/10.1016/j.artint.2024.104145

  3. [3]

    Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D’arcy, David Wadden, Matt Latzke, Minyang Tian, Pan Ji, Shengyan Liu, Hao Tong, Bohao Wu, Yanyu Xiong, Luke Zettlemoyer, Graham Neubig, Dan Weld, Doug Downey, Wen-Tau Yih, Pang Wei Koh, and Hannaneh Hajishirzi. 2024. OpenSch...

  4. [4]

    Christopher A Bail. 2024. Can Generative AI improve social science? Proceedings of the National Academy of Sciences of the United States of America 121, 21 (21 May 2024), e2314021121. https://doi.org/10.1073/pnas.2314021121

  5. [5]

    Jean-Louis Barsoux, Cyril Bouquet, and Michael Wade. [n. d.]. Why Outside Perspectives Are Critical for Inno- vation Breakthroughs. https://sloanreview.mit.edu/article/why-outside-perspectives-are-critical-for-innovation- breakthroughs/. Accessed: 2025-5-9

  6. [6]

    Marcel Binz, Stephan Alaniz, Adina Roskies, Balazs Aczel, Carl T Bergstrom, Colin Allen, Daniel Schad, Dirk Wulff, Jevin D West, Qiong Zhang, Richard M Shiffrin, Samuel J Gershman, Vencislav Popov, Emily M Bender, Marco Marelli, Matthew M Botvinick, Zeynep Akata, and Eric Schulz. 2025. How should the advancement of large language models affect the practic...

  7. [7]

    Pierre Bourdieu. 1975. The specificity of the scientific field and the social conditions of the progress of reason. Social sciences information. Information sur les sciences sociales 14, 6 (Dec. 1975), 19–47. https://doi.org/10.1177/ 053901847501400602

  8. [8]

    Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, ...

Show all 67 references
  1. [9]

    Joel Chan, Joseph Chee Chang, Tom Hope, Dafna Shahaf, and Aniket Kittur. 2018. SOLVENT: A mixed initiative system for finding analogies between research papers. Proceedings of the ACM on human-computer interaction 2, CSCW (1 Nov. 2018), 1–21. https://doi.org/10.1145/3274300

  2. [10]

    Inyoung Cheong, King Xia, K J Kevin Feng, Quan Ze Chen, and Amy X Zhang. 2024. (A)I am not a lawyer, but...: Engaging legal experts towards responsible LLM policies for legal advice. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, U...

  3. [11]

    Clayton M Christensen. 2015. The innovator’s dilemma: when new technologies cause great firms to fail . Harvard Business Review Press

  4. [12]

    Antonio Correia, Shoaib Jameel, Daniel Schneider, Benjamim Fonseca, and Hugo Paredes. 2019. The effect of scientific collaboration on CSCW research: A scientometric study. In 2019 IEEE 23rd International Conference on Computer Supported Cooperative Work in Design (CSCWD) . IEE...

  5. [13]

    Fernando Delgado, Solon Barocas, and Karen Levy. 2022. An Uncommon Task: Participatory Design in Legal AI. Proc. ACM Hum. -Comput. Interact. CSCW1 (8 March 2022). https://doi.org/10.1145/3512898 arXiv:2203.06246 [cs.CY]

  6. [14]

    K J Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2024. Cocoa: Co-planning and Co-execution with AI agents. arXiv [cs.HC] (14 Dec. 2024). arXiv:2412.10999 [cs.HC]

  7. [15]

    Nikolaus Franke, Marion K Poetz, and Martin Schreier. 2014. Integrating problem solvers from analogous markets in new product ideation. Management science 60, 4 (April 2014), 1063–1081. https://doi.org/10.1287/mnsc.2013.1805 , Vol. 1, No. 1, Article . Publication date: May 201...

  8. [16]

    Jie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang, Zheng Zhang, Toby Jia-Jun Li, and Simon Tangi Perrault. 2024. CollabCoder: A lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models. In Proceedings of the CHI Conference on...

  9. [17]

    Jian Gao and Dashun Wang. 2024. Quantifying the use and potential benefits of artificial intelligence in scientific research. Nature human behaviour 8, 12 (Dec. 2024), 2281–2292. https://doi.org/10.1038/s41562-024-02020-5

  10. [18]

    Alireza Ghafarollahi and Markus J Buehler. 2024. SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning. arXiv [cs.AI] (9 Sept. 2024). arXiv:2409.05556 [cs.AI]

  11. [19]

    1994.The new production of knowledge: The dynamics of science and research in contemporary societies

    Michael Gibbons, Camille Limoges, Helga Nowotny, Simon Schwartzman, Peter Scott, and Martin Trow. 1994.The new production of knowledge: The dynamics of science and research in contemporary societies . SAGE Publications, London, England. https://doi.org/10.4135/9781446221853

  12. [20]

    Thilo Hagendorff, Ishita Dasgupta, Marcel Binz, Stephanie C Y Chan, Andrew Lampinen, Jane X Wang, Zeynep Akata, and Eric Schulz. 2023. Machine Psychology. arXiv [cs.CL] (24 March 2023). arXiv:2303.13988 [cs.CL]

  13. [21]

    Laurens K Hessels and Harro van Lente. 2008. Re-thinking new knowledge production: A literature review and a research agenda. Research policy 37, 4 (May 2008), 740–760. https://doi.org/10.1016/j.respol.2008.01.008

  14. [22]

    Charles W L Hill and Frank T Rothaermel. 2003. The performance of incumbent firms in the face of radical technological innovation. Academy of management review 28, 2 (April 2003), 257–274. https://doi.org/10.5465/amr.2003.9416161

  15. [23]

    Ryan Hill, Yian Yin, Carolyn Stein, Xizhao Wang, Dashun Wang, and Benjamin F Jones. 2021. Adaptability and the pivot penalty in science and technology. arXiv [cs.DL] (13 July 2021). arXiv:2107.06476 [cs.DL]

  16. [24]

    Chenyan Jia, Michelle S Lam, Minh Chau Mai, Jeffrey T Hancock, and Michael S Bernstein. 2024. Embedding democratic values into social media AIs via societal objective functions. Proceedings of the ACM on human-computer interaction 8, CSCW1 (17 April 2024), 1–36. https://doi.or...

  17. [25]

    Marina Jirotka, Charlotte P Lee, and Gary M Olson. 2013. Supporting scientific collaboration: Methods, tools and concepts. Computer supported cooperative work: CSCW: an international journal 22, 4-6 (19 Aug. 2013), 667–715. https://doi.org/10.1007/s10606-012-9184-0

  18. [26]

    Hyeonsu Kang, Joseph Chee Chang, Yongsung Kim, and Aniket Kittur. 2022. Threddy: An interactive system for personalized thread-based exploration and organization of scientific literature. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology...

  19. [27]

    Hyeonsu B Kang, Tongshuang Wu, Joseph Chee Chang, and Aniket Kittur. 2023. Synergi: A Mixed-Initiative System for Scholarly Synthesis and Sensemaking. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23, Article 43) . Association...

  20. [28]

    KapaniaShivani, WangRuiyi, Litoby Jia-Jun, LiTianshi, and ShenHong. 2025. ’I’m categorizing LLM as a productivity tool’: Examining ethics of LLM use in HCI research practices. Proceedings of the ACM on human-computer interaction (2 May 2025). https://doi.org/10.1145/3711000

  21. [29]

    Aniket Kittur and Robert E Kraut. 2008. Harnessing the wisdom of crowds in wikipedia: quality through coordination. In Proceedings of the 2008 ACM conference on Computer supported cooperative work . ACM, New York, NY, USA. https://doi.org/10.1145/1460563.1460572

  22. [30]

    Aniket Kittur and Robert E Kraut. 2010. Beyond Wikipedia: coordination and conflict in online production groups. In Proceedings of the 2010 ACM conference on Computer supported cooperative work . ACM, New York, NY, USA. https://doi.org/10.1145/1718918.1718959

  23. [31]

    Hadas Kotek, Sarah Babinski, Rikker Dockum, and Christopher Geissler. 2020. Gender representation in linguistic example sentences. Proceedings of the Linguistic Society of America 5, 1 (23 March 2020), 514. https://doi.org/10.3765/ plsa.v5i1.4723

  24. [32]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in Large Language Models. In Proceedings of The ACM Collective Intelligence Conference . ACM, New York, NY, USA. https://doi.org/10.1145/3582269. 3615599

  25. [33]

    Thomas S Kuhn. 2012. The Structure of Scientific Revolutions. https://press.uchicago.edu/ucp/books/book/chicago/S/ bo13179781.html. Accessed: 2025-2-5

  26. [34]

    Yoonjoo Lee, Hyeonsu B Kang, Matt Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue

  27. [35]

    Hanlin Li, Brent Hecht, and Stevie Chancellor. 2022. All that’s happening behind the scenes: Putting the spotlight on volunteer moderator labor in Reddit. Proceedings of the International AAAI Conference on Web and Social Media 16 (31 May 2022), 584–595. https://doi.org/10.160...

  28. [36]

    Tianhao Li, Sandesh Shetty, Advaith Kamath, Ajay Jaiswal, Xiaoqian Jiang, Ying Ding, and Yejin Kim. 2024. CancerGPT for few shot drug pair synergy prediction using large pretrained language models. npj digital medicine 7, 1 (19 Feb. 2024), 40. https://doi.org/10.1038/s41746-02...

  29. [37]

    Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y Zou. 2024. Mapping the increasing use of LLMs in scientific papers. arXiv [cs.CL] (1 A...

  30. [38]

    Houjiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou, Daisy Pinaroc, Matthew Lease, and Min Kyung Lee. 2024. Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI. Proceedings of the ACM on human-computer interaction 8, CSCW2 (7 Nov. 2024...

  31. [39]

    Yiren Liu, Si Chen, Haocong Cheng, Mengxia Yu, Xiao Ran, Andrew Mo, Yiliu Tang, and Yun Huang. 2024. CoQuest: Exploring Research Question Co-Creation with an LLM-based Agent. In Proceedings of the CHI Conference on Human Factors in Computing Systems , Vol. 46. ACM, New York, N...

  32. [40]

    Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun, Kevin K Nejad, Felipe Yáñez, Bati Yilmaz, Kangjoo Lee, Alexandra O Cohen, Valentina Borghesani, Anton Pashkov, Daniele Marinazzo, Jonathan Nicholas, Alessandro Salatiello, Ilia Sucholutsky, Pasquale Minervini, Sepehr Razavi, Rober...

  33. [41]

    Yao Lyu, Jie Cai, and John M Carroll. 2025. A systematic literature review of infrastructure studies in SIGCHI. arXiv [cs.HC] (13 April 2025). arXiv:2504.09612 [cs.HC]

  34. [42]

    Siddharth Mishra-Sharma, Yiding Song, and Jesse Thaler. 2024. PAPERCLIP: Associating astronomical observations and natural language with multi-modal models. arXiv [astro-ph.IM] (13 March 2024). arXiv:2403.08851 [astro-ph.IM]

  35. [43]

    Miryam Naddaf. 2025. How are researchers using AI? Survey reveals pros and cons for science. Nature (4 Feb. 2025). https://doi.org/10.1038/d41586-025-00343-5

  36. [44]

    Awais Naeem, Tianhao Li, Huang-Ru Liao, Jiawei Xu, Aby M Mathew, Zehao Zhu, Zhen Tan, Ajay Kumar Jaiswal, Raffi A Salibian, Ziniu Hu, Tianlong Chen, and Ying Ding. 2024. Path-RAG: Knowledge-guided key region retrieval for Open-ended pathology visual question answering. arXiv [...

  37. [45]

    Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . ACM, New York, NY, USA, 1–13. https://doi.org/10.1145/359...

  38. [46]

    Andrew B Neang, Will Sutherland, Michael W Beach, and Charlotte P Lee. 2021. Data integration as coordination: The articulation of data work in an ocean science collaboration. Proceedings of the ACM on human-computer interaction 4, CSCW3 (5 Jan. 2021), 1–25. https://doi.org/10...

  39. [47]

    Sachita Nishal, Jasmine Sinchai, and Nicholas Diakopoulos. 2024. Understanding Practices around Computational News Discovery Tools in the Domain of Science Journalism. Proc. ACM Hum.-Comput. Interact. 8, CSCW1 (26 April 2024), 1–36. https://doi.org/10.1145/3637419

  40. [48]

    Gabrielle O’Brien. 2025. How scientists use large language models to program. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–16. https://doi.org/10.1145/3706598.3713668

  41. [49]

    Nicholas G Otis, Solène Delecourt, Katelyn Cranney, and Rembrand Koning. 2024. Global Evidence on Gender Gaps and Generative AI. Harvard Business School

  42. [50]

    Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, and Danielle Bragg. 2025. Understanding the LLM-ification of CHI: Unpacking the impact of LLMs at CHI through a systematic literature review. In Proceedings of the 2025 CHI Conferenc...

  43. [51]

    Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. AngleKindling: Supporting Journalistic Angle Ideation with Large Language Models. In Proceedings of the 2023 CHI Conference on...

  44. [52]

    Jason Priem, Heather Piwowar, and Richard Orr. 2022. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv [cs.DL] (3 May 2022). arXiv:2205.01833 [cs.DL]

  45. [53]

    Mirjana Prpa, Giovanni Maria Troiano, Matthew Wood, and Yvonne Coady. 2024. Challenges and Opportunities of LLM-Based Synthetic Personae and Data in HCI. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24, Article 461) . Associat...

  46. [54]

    Kevin Pu, K J Kevin Feng, Tovi Grossman, Tom Hope, Bhavana Dalvi Mishra, Matt Latzke, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. 2024. IdeaSynth: Iterative research idea development through evolving and composing idea facets with literature-grounded feedback. arXi...

  47. [55]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv [cs.CL] (27 Aug. 2019). arXiv:1908.10084 [cs.CL]

  48. [56]

    2012.The new production of knowledge: The dynamics of science and research in contemporary societies

    Peter Scott, Michael Gibbons, Helga Nowotny, Camille Limoges, Martin Trow, and Simon Schwartzman. 2012.The new production of knowledge: The dynamics of science and research in contemporary societies . SAGE Publications, London, England. 1–192 pages

  49. [57]

    David J Smith. 2012. Technological discontinuities, outsiders and social capital: a case study from Formula 1. European journal of innovation management 15, 3 (27 July 2012), 332–350. https://doi.org/10.1108/14601061211243666

  50. [58]

    Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving Support. In Proceedings of the 2021 CHI Conference on Human...

  51. [59]

    Chris Stokel-Walker and Richard Van Noorden. 2023. What ChatGPT and generative AI mean for science. Nature 614, 7947 (Feb. 2023), 214–216. https://doi.org/10.1038/d41586-023-00340-6

  52. [60]

    Zechang Sun, Yuan-Sen Ting, Yaobo Liang, Nan Duan, Song Huang, and Zheng Cai. 2024. Interpreting multi-band galaxy observations with large language model-based agents. arXiv [astro-ph.IM] (23 Sept. 2024), arXiv:2409.14807. https://doi.org/10.48550/arXiv.2409.14807 arXiv:2409.1...

  53. [61]

    Yu Tao and Kush R Varshney. 2021. Insiders and outsiders in research on machine learning and society. arXiv [cs.CY] (3 Feb. 2021). arXiv:2102.02279 [cs.CY]

  54. [62]

    Mary Uhl-Bien and Michael Arena. 2018. Leadership for organizational adaptability: A theoretical synthesis and integrative framework. The leadership quarterly 29, 1 (Feb. 2018), 89–104. https://doi.org/10.1016/j.leaqua.2017.12.009

  55. [63]

    Ibo Van De Poel. 2000. On the role of outsiders in technical development.Technology Analysis and Strategic Management 12, 3 (Sept. 2000), 383–397. https://doi.org/10.1080/09537320050130615

  56. [64]

    Richard Van Noorden and Jeffrey M Perkel. 2023. AI and science: what 1,600 researchers think. Nature 621, 7980 (27 Sept. 2023), 672–675. https://doi.org/10.1038/d41586-023-02980-0

  57. [65]

    Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, Anima Anandkumar, Karianne Bergen, Carla P Gomes, Shirley Ho, Pushmeet Kohli, Joan Lasenby, Jure Leskovec, Tie-Yan Liu, Arjun Manrai, Debora ...

  58. [66]

    Alyson L Young and Wayne G Lutters. 2017. Infrastructuring for cross-disciplinary synthetic science: Meta-study research in land system science. Computer supported cooperative work: CSCW: an international journal 26, 1-2 (3 April 2017), 165–203. https://doi.org/10.1007/s10606-...

  59. [2024]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems

    PaperWeaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers. In Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–19. https://doi.org/10.1145/3613904.3642196

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.