REVIEW 3 major objections 5 minor 31 references
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that SE research must proactively shape LLM adoption, since LLMs will disrupt research practice whatever their role; rigor demands human oversight, interpretability, and community guidelines.
desk verdict A genuinely useful position paper that maps LLM effects on SE research via McLuhan's Tetrad, with an honest speculative core and a self-referential blind spot worth flagging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is McLuhan's Tetrad of Media Laws, a four-question grid asking what a technology enhances, makes obsolete, retrieves, and reverses into when taken to extremes. The paper applies this grid to each stage of a generic SE research pipeline (question formulation, design, data collection, processing, analysis and interpretation, writing and dissemination, and cross-cutting impacts) to organize speculation about LLM effects into a structured analysis. The Tetrad does the argument's load-bearing work: each 'reverse' entry identifies a risk, and those risks motivate the paper's call for human oversight, reporting guidelines, benchmarks, and education.
What would settle it
Measure the distribution of research topics produced by LLM-assisted brainstorming versus human-only brainstorming across matched groups of researchers: if LLM-assisted questions are not more similar to one another than human-generated questions (as quantified by embedding-based distance), the paper's predicted 'creativity echo chamber' reversal does not hold, weakening the case that human oversight is essential for novelty.
Extended reading notes
Core claim
On its own terms, the paper establishes that the effects of LLMs on SE research are not a single trajectory but a four-sided one: the same technology that accelerates hypothesis generation and data analysis also makes manual practices obsolete, revives older habits of informal discourse and cross-domain theory, and threatens to reverse into homogenized creativity and fabricated findings when pushed to extremes. Applying the Tetrad across the research pipeline produces a structured map of these tensions, and the paper's central claim is that the outcome depends on choices the community makes now—whether to use LLMs only as tools, or to reshape research culture around them with human agency as the anchor. The authors hold that with deliberate integration, LLM use leads to more efficient, data-driven investigations and deeper intellectual engagement, while without oversight it leads to lower researcher skills and compromised rigor.
Load-bearing premise
The whole structured analysis rests on the assumption that one group's speculative application of McLuhan's Tetrad during a two-day symposium yields a reliable map of LLM effects on the research pipeline; the paper itself acknowledges that the effects are 'based on (qualified) assumptions, grounded on early and fragmented experience,' so if that lens or that group's collective speculation is unrepresentative, the paper's structured conclusions lose their foundation.
Editorial extensions
If this is right
- Researchers can redirect time saved by LLMs from tedious tasks toward deeper intellectual engagement, shifting the field's output from sheer quantity of papers toward more impactful research.
- Automation of code generation is expected to revive interest in formal specification, verification, requirements engineering, and other human-in-the-loop topics.
- Without human oversight, LLM-assisted research is at risk of creativity echo chambers, homogenized research questions, and a decline in foundational skills among early-career researchers.
- Adopting transparent reporting standards, validation metrics such as inter-rater reliability, and public benchmarks for LLM-generated artifacts would make LLM-assisted studies reproducible, comparable, and less prone to undetected bias.
Reading between the lines
- The Tetrad's reverse quadrant implies a measurable prediction: research groups that rely heavily on LLMs for ideation may converge on a narrower set of topics than those that do not; this could be tested by comparing topic diversity across labs with different LLM-use policies, using embedding-based distance on published titles and abstracts.
- If the paper's skill-decline concern is right, the effect should show up in longitudinal assessments of graduate students' manual coding and statistical analysis abilities as LLM use becomes routine in training programs.
- The call for benchmarks suggests a concrete deliverable: a publicly available dataset with ground-truth labels for SE research tasks (annotation, summarization, causal inference) that the community could use to decide where LLMs can be reliably deployed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that LLMs will fundamentally disrupt software engineering research and that the SE research community must proactively integrate LLMs into research practices while preserving human agency. The authors apply Marshall McLuhan's Tetrad of Media Laws (Enhance, Obsolesce, Retrieve, Reverse) to a generic research pipeline, producing a structured map of LLM effects across seven stages (Table 1). The analysis is based on collective speculation by ten researchers at the 2nd Copenhagen Symposium on Human-Centered AI in SE, and the paper explicitly acknowledges it rests on qualified assumptions. The paper concludes with a call to action: experiment with LLMs, develop transparent reporting guidelines, create benchmarks, and provide education. The central claim is normative and depends entirely on the credibility of the Tetrad analysis.
Significance. The paper's value lies in its timely, structured synthesis of opportunities and risks of LLMs in SE research, and in its concrete call to action. It is deliberately a position paper: it makes no empirical claims and instead offers a framework for community discussion. The four proposed actions (experimentation, reporting guidelines, benchmarks, education) are actionable and likely to be useful to the SE research community. However, as a basis for guiding research practice, the paper is limited by its admitted speculative nature and by a self-referential methodological blind spot: the table structure that anchors the analysis was itself generated by an LLM, without the transparency the paper demands of others. These issues are fixable within the scope of a position paper, and the manuscript has clear strengths in framing and proposed next steps.
major comments (3)
- [Section 2, Table 1; Section 3 action (2)] The paper discloses that "The structure of Table 1, including the research pipeline phases, was generated using a GPT associated with the Disruptive Playbook..., then refined and filled in by the authors," but provides no audit trail—no prompts, no raw GPT output, no cell-by-cell record of human refinement. This is load-bearing because Table 1 is the central evidence map for the paper's claims, and the reversal/warning cells (e.g., "creativity echo chamber," "homogenized research questions") may reproduce the LLM's priors rather than reflect independent analysis. The paper itself, in Section 3 action (2), demands exactly this transparency ("specifying the LLM model and version used, the exact prompting strategies employed..., and the mechanisms for human oversight"), so the authors should apply their own standard to this paper by adding an appendix with prompts, raw outputs, and a description of the human refinement process.
- [Section 1, Section 3] The method of collective speculation is underspecified. The authors state that "This speculation was conducted collaboratively by a team of 10 researchers during the 2nd Symposium on Human-Centered AI in SE," but do not describe the process: how the Tetrad was applied to each pipeline stage, how individual contributions were aggregated, how disagreements were resolved, or whether any reliability check (e.g., independent replication, inter-rater agreement) was attempted. Since the paper's structured conclusions rest on this exercise, the absence of methodological detail is a major gap that prevents readers from assessing the robustness of Table 1.
- [Section 2 and Table 1 (cross-cutting row)] Several table entries are asserted without supporting argument or citation. For example, the cross-cutting reversal "Lower skills of researchers" and retrieve "Impactful research" are presented as findings, yet Section 3 later acknowledges that the analysis is "based on (qualified) assumptions, grounded on early and fragmented experience." This tension between the table's assertive structure and the paper's own caveat should be resolved: the paper should clearly label speculative entries as hypotheses (e.g., in the table caption or in the text) and, where possible, cite the early evidence it mentions (e.g., [1], [3], [6]) to distinguish observed effects from conjectured ones.
minor comments (5)
- [Section 2.1] The heading "Obsolece" is a typo for "Obsolesce", and Section 2.2 contains "qantitative, qalitative" and "techiques", which should be corrected.
- [Section 2.2] The sentence "Driven by probabilistic patterns, LLMs may generate plausible but incorrect conclusions, misattribute sources, or identify false patterns from noisy data, distorting findings and compromising research validity. One example is about using LLMs as annotators." transitions awkwardly; consider revising the second sentence to flow more naturally.
- [Section 1] The paper applies McLuhan's Tetrad without justifying this choice over alternative frameworks (e.g., affordances, socio-technical systems, task-technology fit). A brief justification would strengthen the framing.
- [Table 1] Some pipeline stages in Table 1 (e.g., Data Collection, Data Processing) are not discussed in the text; the paper should either expand on them or explicitly state that only prioritized stages are elaborated due to space.
- [Section 3, action (3)] The benchmark proposal would benefit from referencing existing LLM evaluation efforts in SE (e.g., SWE-bench) to situate the proposal within ongoing community work.
Circularity Check
No circular derivation: this is a position paper with no fitted parameters, equations, or empirical predictions; self-citations supply framing, not load-bearing proof.
full rationale
The paper is a position piece that applies McLuhan's Tetrad to speculate on LLM effects across SE research stages; it contains no equations, no fitted parameters, and no empirical predictions, so the standard circularity patterns (self-definitional reductions, fitted input called prediction, uniqueness imported from authors, ansatz smuggled via citation) do not apply. The only construction-like step is the disclosure in Section 2 that 'The structure of Table 1, including the research pipeline phases, was generated using a GPT associated with the Disruptive Playbook..., then refined and filled in by the authors.' This is a transparency statement about process, not a derivation: the cell contents are authored and supported by external citations ([1], [3], [4], [6], [8], [19], etc.), and the paper does not claim the GPT output is evidence. The central normative claim—that the SE research community should proactively engage with LLMs while preserving human agency—is argued from concerns about bias, oversight, rigor, and reproducibility, not reduced to a self-citation. The paper itself flags its speculative basis ('our speculation would also change as LLMs evolve'; 'based on (qualified) assumptions, grounded on early and fragmented experience'), which is a limitation rather than circularity. The self-referential point that an LLM assisted in generating a taxonomy about LLM risks is a methodological caveat about provenance, but no quoted text shows the conclusion is equivalent to the model's output; human refinement is explicitly stated. Therefore no circular step meets the evidence bar.
Assumptions & free parameters
assumptions (3)
- domain assumption McLuhan's Tetrad of Media Laws is an appropriate framework for analyzing the impact of LLMs on SE research.
- domain assumption The research pipeline stages in Table 1 represent the major phases of SE research.
- ad hoc to paper The collective speculation of the 10 researchers at the 2nd Copenhagen Symposium is sufficient to draw conclusions about LLM effects.
Cite this review
Pith. "Pith review of Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research." pith.science (2026). https://pith.science/paper/2RSTOK5Q
@misc{pith2026250612691,
author = {Pith},
title = {Pith review of: Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/2RSTOK5Q}},
note = {Machine review of arXiv:2506.12691}
}
read the original abstract
The adoption of Large Language Models (LLMs) is not only transforming software engineering (SE) practice but is also poised to fundamentally disrupt how research is conducted in the field. While perspectives on this transformation range from viewing LLMs as mere productivity tools to considering them revolutionary forces, we argue that the SE research community must proactively engage with and shape the integration of LLMs into research practices, emphasizing human agency in this transformation. As LLMs rapidly become integral to SE research - both as tools that support investigations and as subjects of study - a human-centric perspective is essential. Ensuring human oversight and interpretability is necessary for upholding scientific rigor, fostering ethical responsibility, and driving advancements in the field. Drawing from discussions at the 2nd Copenhagen Symposium on Human-Centered AI in SE, this position paper employs McLuhan's Tetrad of Media Laws to analyze the impact of LLMs on SE research. Through this theoretical lens, we examine how LLMs enhance research capabilities through accelerated ideation and automated processes, make some traditional research practices obsolete, retrieve valuable aspects of historical research approaches, and risk reversal effects when taken to extremes. Our analysis reveals opportunities for innovation and potential pitfalls that require careful consideration. We conclude with a call to action for the SE research community to proactively harness the benefits of LLMs while developing frameworks and guidelines to mitigate their risks, to ensure continued rigor and impact of research in an AI-augmented future.
Reference graph
Works this paper leans on
-
[1]
Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel
-
[3]
Cauã Barros, Bruna Azevedo, Valdemar Neto, Mohamad Kassab, Marcos Kali- nowski, Hugo Nascimento, and Michelle Bandeira. 2025. Large Language Model for Qualitative Research - A Systematic Mapping Study. In Workshop on Method- ological Issues with Empirical Studies in Software Engineering (WSESE@ICSE’25)
work page 2025
-
[6]
Katia Romero Felizardo, Márcia Sampaio Lima, Anderson Deizepe, Tayana Uchôa Conte, and Igor Steinmacher. 2024. ChatGPT application in Systematic Literature Reviews in Software Engineering: an evaluation of its accuracy to support the selection activity. In Empirical Software Engineering and Measurement . 25–36
work page 2024
-
[2]
Muneera Bano, Rashina Hoda, Didar Zowghi, and Christoph Treude. 2024. Large language models for qualitative research in software engineering: exploring opportunities and challenges. Automated Software Engineering 31, 1 (2024), 8
work page 2024
-
[4]
Stefano De Paoli. 2024. Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provoca- tion on the limits of the approach. Soc Sci Comput Rev 42, 4 (2024), 997–1019
work page 2024
-
[5]
Markman Ellis. 2008. An introduction to the coffee-house: A discursive model. Language & Communication 28, 2 (2008), 156–164
work page 2008
-
[7]
Marco Gerosa, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. 2024. Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering 31, 1 (2024), 13
work page 2024
-
[8]
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Vikram Dhillo...
work page 2025
Show all 31 references
-
[9]
Timothy C Guetterman, Michael D Fetters, and John W Creswell. 2015. Integrat- ing quantitative and qualitative results in health science mixed methods research through joint displays. The Annals of Family Medicine 13, 6 (2015), 554–561
2015
-
[10]
Qusai Khraisha, Sophie Put, Johanna Kappenberg, Azza Warraitch, and Kristin Hadfield. 2024. Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages...
2024
-
[11]
Xuân-Lan Lam Hoai and Thierry Simonart. 2023. Comparing meta-analyses with ChatGPT in the evaluation of the effectiveness and tolerance of systemic therapies in moderate-to-severe plaque psoriasis. J Clin Med 12, 16 (2023), 5410
2023
-
[12]
Tobias Lorey, Paul Ralph, and Michael Felderer. 2022. Social science theories in software engineering research. In Int’l Conf. on Software Engineering. 1994–2005
2022
-
[13]
Walter S Mathis, Sophia Zhao, Nicholas Pratt, Jeremy Weleff, and Stefano De Paoli
-
[14]
Marshall McLuhan. 1977. Laws of the Media. ETC: A Review of General Semantics (1977), 173–179
1977
-
[15]
Marshall McLuhan. 2017. The medium is the message. In Commun Theory . Routledge, 390–402
2017
-
[16]
Bertrand Meyer. 2023. AI Does Not Help Programmers. https://cacm.acm.org/ blogcacm/ai-does-not-help-programmers/
2023
-
[17]
Paul Ralph, Rashina Hoda, and Christoph Treude. 2020. ACM SIGSOFT empirical standards. (2020)
2020
-
[18]
Daniel Russo, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil Ernst, Thomas Fritz, et al. 2024. Generative ai in software engineering must be human-centered: The copenhagen manifesto. J. Syst. Softw....
2024
-
[19]
Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew Kun, and Hagit Ben Shoshan. 2024. AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation. In CHI Conf. on Human Factors in Computing Systems . 1–17
2024
-
[20]
Igor Steinmacher, Jacob Mcauley Penney, Katia Romero Felizardo, Alessandro F Garcia, and Marco A Gerosa. 2024. Can ChatGPT emulate humans in software engineering surveys?. InProc. of the 18th ACM/IEEE Int’l. Symposium on Empirical Software Engineering and Measurement . 414–419
2024
-
[21]
Klaas-Jan Stol. 2024. Teaching Theorizing in Software Engineering Research. arXiv:2406.17174 [cs.SE] https://arxiv.org/abs/2406.17174
2024 arXiv
-
[22]
Margaret-Anne Storey, Rashina Hoda, Alessandra Maciel Paz Milani, and Maria Teresa Baldassarre. 2025. Guiding Principles for Using Mixed Methods Research in Software Engineering. http://arxiv.org/abs/2404.06011
2025 arXiv
-
[23]
Margaret-Anne Storey, Daniel Russo, Nicole Novielli, Takashi Kobayashi, and Dong Wang. 2024. A disruptive research playbook for studying disruptive innovations. ACM TOSEM 33, 8 (2024), 1–29
2024
-
[24]
Persson, Gerbrand Ceder, and Anubhav Jain
Vahe Tshitoyan, John Dagdelen, Leigh Weston, Alexander Dunn, Ziqin Rong, Olga Kononova, Kristin A. Persson, Gerbrand Ceder, and Anubhav Jain. 2019. Unsupervised word embeddings capture latent knowledge from materials science literature. Nat. 571, 7763 (2019), 95–98. doi:10.103...
2019 doi
-
[25]
Stefan Wagner, Marvin Muñoz Barón, Davide Falessi, and Sebastian Baltes
-
[26]
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, Anima Anandkumar, Karianne Bergen, Carla Gomes, Shirley Ho, Pushmeet Kohli, Joan Lasenby, Jure Leskovec, Tie-Yan Liu, Arjun Manrai, Debora Ma...
2023
-
[27]
Ting Zhang, Ivana Clairine Irsan, Ferdian Thung, and David Lo. 2024. Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models. ACM Trans. Softw. Eng. Methodol. (Sept. 2024). doi:10.1145/3697009
2024 doi
-
[28]
arXiv:2411.07668 [cs.SE] https://arxiv.org/abs/2411.07668
Towards Evaluation Guidelines for Empirical Studies involving LLMs. arXiv:2411.07668 [cs.SE] https://arxiv.org/abs/2411.07668
-
[31]
Chengbo Zheng, Yuanhao Zhang, Zeyu Huang, Chuhan Shi, Minrui Xu, and Xiaojuan Ma. 2024. DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration. In ACM Symposium on User Interface Software and Technology. 1–20
2024
-
[2024]
Inductive thematic analysis of healthcare qualitative interviews using open- source large language models: How does it compare to traditional methods? Computer Methods and Programs in Biomedicine 255 (2024), 108356
2024
-
[2025]
In 2025 IEEE/ACM 22nd Int’l
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?. In 2025 IEEE/ACM 22nd Int’l. Conf. on Mining Software Repositories (MSR)
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.