REVIEW 3 major objections 4 minor 1 cited by
GenAI Is No Silver Bullet for Qualitative Research in Software Engineering
T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Autonomous GenAI cannot replace qualitative researchers in software engineering; the current evidence supports only narrow, low-context tasks like deductive coding, summarization, and transcription.
desk verdict A useful, well-hedged synthesis arguing that GenAI is not a silver bullet for qualitative SE research, but the paper's own empirical scan is too thin to carry much weight and a few categorical claims are stronger than the evidence justifies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing device is a two-dimensional map of qualitative research: research strategies (respondent, field, lab, data) crossed with dimensions of epistemology (postpositivist vs. constructivist), coding strategy (inductive, deductive, hybrid), data granularity and type, and iteration/researcher roles. This map does the argumentative work of predicting where GenAI fits—postpositivist, deductive, low-context, fine-grained annotation tasks—and where it does not: constructivist, inductive, context-rich, reflexive interpretation. The paper also reworks standard quality criteria (reliability, validity, reflexivity, ethics) as lenses for evaluating GenAI-assisted studies.
What would settle it
Conduct a blinded, preregistered study in which a zero-shot large language model—given no codebook and no examples—analyzes interview transcripts from a novel, previously unresearched organizational context, and have experienced grounded-theory researchers, blind to the source, rate whether the model's themes capture latent interpretations and contextual nuance as well as human-generated themes; if the model consistently matches or exceeds human ratings, the paper's claim that GenAI cannot make sense of novel contexts and is unsuitable for interpretive work would be refuted.
Extended reading notes
Core claim
The paper's central claim is that current evidence supports GenAI only for low-context, deductive annotation tasks and for acceleration aids such as transcription and summarization, while genuinely interpretive work—thematic synthesis, grounded theory, ethnography, and other constructivist analyses—remains outside GenAI's demonstrated capabilities and is in epistemological tension with it. The authors ground this in a review of the spectrum of qualitative research strategies, a small empirical scan of 2025 conference papers showing minimal adoption in software engineering venues (zero of the ICSE and CHASE papers reviewed reported GenAI use in coding, while seven of 209 CSCW qualitative pape
Load-bearing premise
The load-bearing premise is that GenAI tools, being statistical neural networks, cannot make sense of novel contexts and cannot participate in the co-construction of meaning that constructivist qualitative research requires; the paper asserts this without direct empirical evidence, and if a future model demonstrably did either, the central argument would need substantial revision.
Editorial extensions
If this is right
- GenAI can be used safely as a fast second coder for well-defined deductive codebooks, as a transcription and translation tool, and as a source of candidate themes that humans must validate.
- Inductive and interpretive studies, including grounded theory, ethnography, and reflexive thematic analysis, should not treat GenAI output as analysis; at most it can surface patterns for human interpretation.
- Software engineering venues should require disclosure of any GenAI involvement in qualitative coding, including subtle tool autocomplete features, along with prompts, model versions, and parameter settings.
- Reliability metrics such as inter-coder agreement are insufficient on their own; they must be paired with reflexivity, positionality statements, and meaningful member checking when GenAI participates.
- Future benchmarking should compare human-human, human-AI, and AI-AI agreement across diverse software engineering artifact types to map where GenAI can substitute for multiple coders and where it cannot.
Reading between the lines
- Beyond the paper: if constructive, interpretive research remains resistant to GenAI, the field may split into scalable, artifact-heavy deductive analyses that increasingly use GenAI and a smaller, more human-centered interpretive tradition whose value rises precisely because it cannot be automated.
- Beyond the paper: the paper's framework yields a testable prediction—GenAI performance in coding tasks should track the clarity of the codebook and the contextual dependence of the labels; a benchmark that varies both dimensions could draw a capability frontier.
- Beyond the paper: the tentative evidence that adoption is driven more by venue timing than by discipline resistance could be tested by tracking whether ICSE and CHASE papers catch up to CSCW in GenAI use in subsequent years.
- Beyond the paper: the epistemological mismatch might narrow if the community develops standards for documenting a model's positionality, such as reporting training-data provenance and prompt choices, in much the way human positionality statements are used.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a position/reflection piece for the software-engineering community. It argues that claims that GenAI can automate qualitative analysis overgeneralize from narrow successes. The authors distinguish research strategies (respondent, field, lab, data) and dimensions of qualitative work (epistemology, coding strategy, granularity, iteration and researcher roles), then review a convenience sample of recent ICSE, CHASE, and CSCW papers and synthesize the emerging literature on LLM-assisted coding, summarization, and thematic analysis. Their central conclusion is that GenAI is useful mainly as a coding or summarization aid—particularly for deductive, low-context annotation—but is not an autonomous qualitative researcher, especially for constructivist and interpretivist studies. The paper closes with quality considerations and a five-item research agenda for benchmarking, interpretive methods, human-AI workflows, standards, and paradigm reconciliation.
Significance. If accepted, the paper provides a timely counterweight to enthusiasm about GenAI in qualitative SE research and offers useful guidance on where automated assistance is currently appropriate. It is heavily cited, carefully hedged in most of its formulations, and explicitly discloses the preliminary nature of its empirical scan, which is accompanied by a replication package—a strength. However, the paper is primarily a synthesis and argument; its only new empirical contribution is a small, convenience-sampled scan whose loading on the central conclusion is limited but still present. The broadest claim—that GenAI tools cannot make sense of novel contexts and are therefore unsuitable for constructivist work—is asserted rather than established; it is a falsifiable empirical hypothesis, but no direct test is offered. If that claim is weakened to 'current evidence does not show,' the paper remains a valuable scoping review and roadmap for the community.
major comments (3)
- [§4.2 and §2.1] The statement 'GenAI tools cannot make sense of novel contexts' is load-bearing: it underlies the conclusion that GenAI cannot be an autonomous qualitative researcher and is epistemologically incompatible with constructivist methods. Yet 'sense making' is never operationally defined, and no evidence is given for the categorical impossibility. The paper itself cites LLMs successfully summarizing/translating novel texts and producing themes that are 'less aware of latent interpretations' [30]—a graded, empirical failure mode, not an absence of sense-making. This tension weakens the central claim. Please either (a) define sense-making and specify a falsifiable test, or (b) replace the categorical 'cannot' with 'current evidence does not show,' which is the paper's own formulation in §5. Without this change, the philosophical claim overreaches the evidence.
- [§3.1, Table 1] The screening pass uses ChatGPT without reporting precision, recall, or validation against a hand-checked sample; the keyword search and single-author manual check are not reliability-assessed. The claim that 'none of the ICSE or CHASE papers' used GenAI and that CSCW prevalence is 7/209 (3.3%) rests on this unvalidated pipeline. There is also a numeric inconsistency: the text later says '7 CSCW (2.7%)' without a denominator; 7/312 is 2.2% and 7/209 is 3.3%. Since the section is explicitly preliminary, this is fixable by reporting the ChatGPT accuracy on a small gold-standard subset, clarifying the denominator, and labeling the counts as indicative. I would not reject the paper over this, but the numbers need cleaning.
- [§4.2, §5] The paper conflates an empirical claim about current LLM performance with a principled/epistemological claim about the nature of a 'statistical neural network'. §4.2 says LLMs 'lack socially embedded sense making' and are 'fundamentally at odds' with constructivist approaches, while §5 summarizes the argument as 'current evidence shows.' These have different scopes. If the stronger claim is intended, the paper should engage with the possibility that constructivist researchers might use GenAI in reflexive workflows and explain why the weak-but-real contextual sensitivity demonstrated in translation/summarization tasks fails to count. If the weaker claim is intended, the §4.2 wording should be aligned. Distinguishing the two would make the central argument more precise and defensible.
minor comments (4)
- [Table 1] Use one denominator for all percentages. The same seven papers are described as 3.3% of 209 and 2.7% of an unstated base; 7/312 is 2.2%.
- [§3.1] Specify whether the keyword search was applied to full text or only to titles/abstracts, and state the exact search string so the filtering step is independently replicable.
- [§4.1] The sentence 'This will dramatically scale up the number of artifacts such studies can analyze' reads as promotional; consider a neutral phrasing consistent with the paper's cautious stance.
- [Footnote 6] The aside that some researchers question rigor criteria is interesting but not integrated into the quality-criteria discussion; a sentence linking it to the argument would help.
Circularity Check
No circular derivation; self-citations are supporting literature, not load-bearing.
full rationale
This paper is a position/literature review, not a derivation or prediction exercise. It makes no fitted parameters, no equations, and no first-principles claims that could reduce to its own inputs by construction. The central argument—that GenAI support must be tailored to research strategy, data type, and epistemology, and that autonomous GenAI qualitative analysis is not currently supported by evidence—is assembled from external, independently published and falsifiable studies (e.g., Ahmed et al. [20], Shah et al. [35], Montes et al. [30], the mapping study [23], Wen et al. [36]) and from the authors' own descriptive scan in §3.1. The paper does cite the authors' own prior work (Storey et al. [11] for research-strategy distinctions; Ahmed et al. [20] for annotation evidence; Treude & Hata [15] for fairness), but these are used as ordinary literature support rather than as a uniqueness theorem or as the sole justification for the central claim. Removing those self-citations would not collapse the argument, because the conclusion rests mainly on external evidence and on the paper's own review of current usage. The closest concern is §4.2's assertion that 'GenAI tools cannot make sense of novel contexts' and the §2.1 claim that a statistical neural network cannot participate in co-construction of meaning; these are philosophical/empirical assertions without an operational definition of 'sense-making.' That is an unsupported-premise or correctness risk, not circularity, since the assertion is not derived from or equivalent to the paper's inputs. Overall, the paper is self-contained as an evidence-review and position statement; the minor self-citations are non-load-bearing, so the circularity score is 1.
Assumptions & free parameters
assumptions (4)
- domain assumption Storey et al.'s socio-technical framework classifies SE research strategies as respondent, field, lab, and data strategies.
- domain assumption Postpositivist and constructivist epistemologies are distinct, and GenAI fits 'most naturally with positivist epistemologies'.
- domain assumption LLMs lack 'socially embedded sense making' and therefore cannot participate in co-construction of meaning.
- domain assumption Inter-coder agreement metrics such as Cohen's κ and Krippendorff's α are meaningful measures of GenAI quality in qualitative coding.
Cite this review
Pith. "Pith review of GenAI Is No Silver Bullet for Qualitative Research in Software Engineering." pith.science (2026). https://pith.science/paper/6MXDM4YT
@misc{pith2026260308951,
author = {Pith},
title = {Pith review of: GenAI Is No Silver Bullet for Qualitative Research in Software Engineering},
year = {2026},
howpublished = {\url{https://pith.science/paper/6MXDM4YT}},
note = {Machine review of arXiv:2603.08951}
}
read the original abstract
Qualitative research gives rich insights into the quintessentially human aspects of software engineering as a socio-technical system. Qualitative research spans diverse strategies and methods, from interpretivist, in situ observational field studies, to deductive coding of data from mining studies. Advances in large language models and generative AI (GenAI) have prompted claims that artificial intelligence could automate qualitative analysis. Such claims are overgeneralizing from narrow successes. GenAI support must be carefully adapted to the data of interest, but also to the characteristics of a particular research strategy. In this Frontiers of SE paper, we discuss the emerging use of GenAI in relation to the broad spectrum of qualitative research in software engineering. We outline the dimensions of qualitative work in software engineering, review emerging empirical evidence for GenAI assistance, examine the pros and cons of GenAI-mediated qualitative research practices, and revisit qualitative research quality factors, in light of GenAI. Our goal is to inform researchers about the promises and pitfalls of GenAI-assisted qualitative research. We conclude with future plans to advance understanding of its use in software engineering.
Forward citations
Cited by 1 Pith paper
-
An Empirical Study of Model Context Protocol Applications
A large-scale study of 1,723 MCP host applications finds partial convergence on file-based configuration (85.2%) and official SDKs (81.1%), but only 37.2% gate tool execution behind a blocking approval step.
Reference graph
Works this paper leans on
-
[30]
Cristina Martinez Montes, Robert Feldt, Cristina Miguel Martos, Sofia Ouhbi, Shweta Premanandan, and Daniel Graziotin. 2025. Large language models in thematic analysis: prompt engineering, evaluation, and guidelines for qualitative software engineering research. (2025). doi:10.48550/ar xiv.2510.18456
-
[1]
Louise H Kidder and Michelle Fine. 1987. Qualitative and quantitative methods: When stories converge. InMultiple Methods in Program Evaluation. Number 35 in New Directions for Program Evaluation. Jossey-Bass. Manuscript submitted to ACM 10 Neil A. Ernst and Christoph Treude
1987
-
[2]
Margarete Sandelowski. 1993. Rigor or rigor mortis: the problem of rigor in qualitative research revisited.Advances in Nursing Science, 16, 2, 1–8
1993
-
[3]
Blake D. Poland. 1995. Transcription quality as an aspect of rigor in qualitative research.Qualitative Inquiry, 1, 3, (Sept. 1995), 290–310. doi:10.1177 /107780049500100302
1995
-
[4]
Siw Elisabeth Hove and Bente Anda. 2005. Experiences from conducting semi-structured interviews in empirical software engineering research. In11th IEEE International Software Metrics Symposium (METRICS’05). IEEE, 10–pp
2005
-
[5]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology.Qualitative Research in Psychology, 3, 2, (Jan. 2006), 77–101. doi:10.1191/1478088706qp063oa
-
[6]
2008.Selecting empirical methods for software engineering research.Guide to Advanced Empirical Software Engineering. Springer London, 285–311. isbn: 9781848000445. doi:10.1007/978-1-84800-044-5_11
-
[7]
Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering.Empirical software engineering, 14, 2, 131–164
2009
Show all 38 references
-
[8]
2011.Self-organizing agile teams: A grounded theory
Rashina Hoda. 2011.Self-organizing agile teams: A grounded theory. Ph.D. Dissertation. Open Access Te Herenga Waka-Victoria University of Wellington
2011
-
[9]
Linda Birt, Suzanne Scott, Debbie Cavers, Christine Campbell, and Fiona Walter. 2016. Member checking: a tool to enhance trustworthiness or merely a nod to validation?Qualitative Health Research, 26, 13, (July 2016), 1802–1811. doi:10.1177/1049732316654870
2016 doi
-
[10]
Kirsti Malterud, Volkert Dirk Siersma, and Ann Dorrit Guassora. 2016. Sample size in qualitative interview studies: guided by information power. Qualitative Health Research, 26, 13, (July 2016), 1753–1760. doi:10.1177/1049732315617444
2016 doi
-
[11]
Ernst, Courtney Williams, and Eirini Kalliamvakou
Margaret-Anne Storey, Neil A. Ernst, Courtney Williams, and Eirini Kalliamvakou. 2020. The who, what, how of software engineering research: a socio-technical framework.Empirical Software Engineering, 25, 5, (Aug. 2020), 4097–4129. doi:10.1007/s10664-020-09858-z
2020 doi
-
[12]
Vivek Arora, Enrique Larios Vargas, Maurício Aniche, and Arie van Deursen. 2021. Secure software engineering in the financial services: a practitioners’ perspective.arXiv preprint arXiv:2104.03476
2021 arXiv
-
[13]
Shih-Chieh Dai, Aiping Xiong, and Lun-Wei Ku. 2023. LLM-in-the-loop: leveraging large language model for thematic analysis. InFindings of the Association for Computational Linguistics: EMNLP 2023. Houda Bouamor, Juan Pino, and Kalika Bali, (Eds.) Association for Computational ...
2023 doi
-
[14]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InProceedings of the 40th International Conference on Machine Learning, 28492–28518. https://proceedings.mlr.press/v202...
2023
-
[15]
Christoph Treude and Hideaki Hata. 2023. She elicits requirements and he tests: software engineering gender bias in large language models. In 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR). IEEE, 624–629
2023
-
[16]
open codes
John Chen, Alexandros Lotsos, Sihan Cheng, Caiyi Wang, Lexie Zhao, Jessica Hullman, Bruce Sherin, Uri Wilensky, and Michael Horn. 2024. A computational method for measuring" open codes" in qualitative analysis.arXiv preprint arXiv:2411.12142
2024 arXiv
-
[17]
Robert M Davison et al. 2024. The ethics of using generative ai for qualitative data analysis.Information Systems Journal, 34, 5, 1433–1439
2024
-
[18]
Yvonne Dittrich, Helen Sharp, and Cleidson de Souza. 2024. Teaching and learning ethnography for software engineering contexts. InHandbook on Teaching Empirical Software Engineering. Springer, 593–630
2024
-
[19]
Per Lenberg, Robert Feldt, Lucas Gren, Lars Göran Wallgren Tengberg, Inga Tidefors, and Daniel Graziotin. 2024. Qualitative software engineering research: reflections and guidelines.Journal of Software: Evolution and Process, 36, 6, e2607
2024
-
[20]
Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025. Can LLMs replace manual annotation of software engineering artifacts? In2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR). IEEE, 526–538
2025
-
[21]
Lama Alqazlan, Zheng Fang, Michael Castelle, and Rob Procter. 2025. A novel, human-in-the-loop computational grounded theory framework for big social data.Big Data & Society, 12, 2, 20539517251347598
2025
-
[22]
Sebastian Baltes et al. 2025. Evaluation guidelines for empirical studies in software engineering involving llms.arXiv preprint arXiv:2508.15503
2025 arXiv
-
[23]
Cauã Ferreira Barros, Bruna Borges Azevedo, Valdemar Vicente Graciano Neto, Mohamad Kassab, Marcos Kalinowski, Hugo Alexandre D Do Nascimento, and Michelle CGSP Bandeira. 2025. Large language model for qualitative research: a systematic mapping study. In2025 IEEE/ACM Internati...
2025
-
[24]
Eva Blondeel, Patricia Everaert, and Evelien Opdecam. 2025. A practical guide to implementing chatgpt as a secondary coder in qualitative research.International Journal of Accounting Information Systems, 56, (Dec. 2025), 100754. doi:10.1016/j.accinf.2025.100754
2025
-
[25]
Breno Felix De Sousa, Ronnie de Souza Santos, and Kiev Gama. 2025. Integrating positionality statements in empirical software engineering research. In2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, (Ma...
2025
-
[26]
Zackary Okun Dunivin. 2025. Scaling hermeneutics: a guide to qualitative coding with llms for reflexive content analysis.EPJ Data Science, 14, 1, 28
2025
-
[27]
Fine Michelle
Tanisha Jowsey, Virginia Braun, Victoria Clarke, Deborah Lupton, and et al. Fine Michelle. 2025. We reject the use of generative artificial intelligence for reflexive qualitative research.Qualitative Inquiry, (Dec. 2025). doi:10.1177/10778004251401851
2025 doi
-
[28]
Daye Kang, Zhuolun Han, Jiahe Tian, Muhan Zhang, and Jeffrey M Rzeszotarski. 2025. Themeviz: understanding the effect of human-ai collaboration in theme development with an llm-enhanced interactive visual system.Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct....
2025 doi
-
[29]
Yanheng Li, Da Wang, and Yuping Wang. 2025. Understanding voter fraud misinformation videos during the 2024 taiwan election on youtube. Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct. 2025), 1–40. doi:10.1145/3757697
2025 doi
-
[31]
Duc Cuong Nguyen and Catherine Welch. 2025. Generative artificial intelligence in qualitative data analysis: analyzing—or just chatting? Organizational Research Methods, 29, 1, (Sept. 2025), 3–39. doi:10.1177/10944281251377154
2025 doi
-
[32]
Kien Nguyen-Trung. 2025. Chatgpt in thematic analysis: can ai become a research assistant in qualitative research?Quality & Quantity, 59, 6, (June 2025), 4945–4978. doi:10.1007/s11135-025-02165-z
2025 doi
-
[33]
Shruti Phadke. 2025. Exit stories: using reddit self-disclosures to understand disengagement from problematic communities.Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct. 2025), 1–27. doi:10.1145/3757592
2025 doi
-
[34]
Carolyn Seaman, Rashina Hoda, and Robert Feldt. 2025. Qualitative research methods in software engineering: past, present, and future.IEEE Transactions on Software Engineering
2025
-
[35]
Syed Tauhid Ullah Shah, Mohamad Hussein, Ann Barcomb, and Mohammad Moshirpour. 2025. From inductive to deductive: llms-based qualitative data analysis in requirements engineering.arXiv preprint arXiv:2504.19384
2025 arXiv
-
[36]
Chuanchi Wen, Paul Clough, Rachel Paton, and Rebecca Middleton. 2025. Leveraging large language models for thematic analysis: a case study in the charity sector.AI & Society, (Aug. 2025). doi:10.1007/s00146-025-02487-4
2025 doi
-
[37]
Sargam Yadav, Abhishek Kaushik, and Asifa Mehmood Qureshi. 2025. Thematic analysis of expert opinions on the use of large language models in software development.Applied AI Letters, 6, 3, e127
2025
-
[38]
Hanmo You, Zan Wang, Bin Lin, and Junjie Chen. 2025. Navigating the testing of evolving deep learning systems: an exploratory interview study. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, (Apr. 2025), 2726–2738. doi:10.1109/icse55347.2025...
2025
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.