Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

GenAI Is No Silver Bullet for Qualitative Research in Software Engineering

T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Autonomous GenAI cannot replace qualitative researchers in software engineering; the current evidence supports only narrow, low-context tasks like deductive coding, summarization, and transcription.

desk verdict A useful, well-hedged synthesis arguing that GenAI is not a silver bullet for qualitative SE research, but the paper's own empirical scan is too thin to carry much weight and a few categorical claims are stronger than the evidence justifies. read the letter →

arxiv 2603.08951 v2 pith:6MXDM4YT submitted 2026-03-09 cs.SE

classification cs.SE
keywords qualitativeresearchsoftwareengineeringgenerativeAIlargelanguagemodelsthematicanalysisgroundedtheorydeductivecodingmethods
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that claims of generative AI automating qualitative research in software engineering are overgeneralized from narrow successes. It argues that GenAI support must be tailored not only to the data but to the specific research strategy, epistemological stance, coding approach, and researcher role. Reviewing recent evidence, it finds GenAI useful mainly as a transcription, summarization, translation, or deductive-coding aid, not as an autonomous qualitative analyst. The paper matters because it gives software engineering researchers a structured way to decide where GenAI can help and where it risks undermining interpretive depth and constructivist values.

What carries the argument

The central organizing device is a two-dimensional map of qualitative research: research strategies (respondent, field, lab, data) crossed with dimensions of epistemology (postpositivist vs. constructivist), coding strategy (inductive, deductive, hybrid), data granularity and type, and iteration/researcher roles. This map does the argumentative work of predicting where GenAI fits—postpositivist, deductive, low-context, fine-grained annotation tasks—and where it does not: constructivist, inductive, context-rich, reflexive interpretation. The paper also reworks standard quality criteria (reliability, validity, reflexivity, ethics) as lenses for evaluating GenAI-assisted studies.

What would settle it

Conduct a blinded, preregistered study in which a zero-shot large language model—given no codebook and no examples—analyzes interview transcripts from a novel, previously unresearched organizational context, and have experienced grounded-theory researchers, blind to the source, rate whether the model's themes capture latent interpretations and contextual nuance as well as human-generated themes; if the model consistently matches or exceeds human ratings, the paper's claim that GenAI cannot make sense of novel contexts and is unsuitable for interpretive work would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that current evidence supports GenAI only for low-context, deductive annotation tasks and for acceleration aids such as transcription and summarization, while genuinely interpretive work—thematic synthesis, grounded theory, ethnography, and other constructivist analyses—remains outside GenAI's demonstrated capabilities and is in epistemological tension with it. The authors ground this in a review of the spectrum of qualitative research strategies, a small empirical scan of 2025 conference papers showing minimal adoption in software engineering venues (zero of the ICSE and CHASE papers reviewed reported GenAI use in coding, while seven of 209 CSCW qualitative pape

Load-bearing premise

The load-bearing premise is that GenAI tools, being statistical neural networks, cannot make sense of novel contexts and cannot participate in the co-construction of meaning that constructivist qualitative research requires; the paper asserts this without direct empirical evidence, and if a future model demonstrably did either, the central argument would need substantial revision.

Editorial extensions

If this is right

  • GenAI can be used safely as a fast second coder for well-defined deductive codebooks, as a transcription and translation tool, and as a source of candidate themes that humans must validate.
  • Inductive and interpretive studies, including grounded theory, ethnography, and reflexive thematic analysis, should not treat GenAI output as analysis; at most it can surface patterns for human interpretation.
  • Software engineering venues should require disclosure of any GenAI involvement in qualitative coding, including subtle tool autocomplete features, along with prompts, model versions, and parameter settings.
  • Reliability metrics such as inter-coder agreement are insufficient on their own; they must be paired with reflexivity, positionality statements, and meaningful member checking when GenAI participates.
  • Future benchmarking should compare human-human, human-AI, and AI-AI agreement across diverse software engineering artifact types to map where GenAI can substitute for multiple coders and where it cannot.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if constructive, interpretive research remains resistant to GenAI, the field may split into scalable, artifact-heavy deductive analyses that increasingly use GenAI and a smaller, more human-centered interpretive tradition whose value rises precisely because it cannot be automated.
  • Beyond the paper: the paper's framework yields a testable prediction—GenAI performance in coding tasks should track the clarity of the codebook and the contextual dependence of the labels; a benchmark that varies both dimensions could draw a capability frontier.
  • Beyond the paper: the tentative evidence that adoption is driven more by venue timing than by discipline resistance could be tested by tracking whether ICSE and CHASE papers catch up to CSCW in GenAI use in subsequent years.
  • Beyond the paper: the epistemological mismatch might narrow if the community develops standards for documenting a model's positionality, such as reporting training-data provenance and prompt choices, in much the way human positionality statements are used.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper is a position/reflection piece for the software-engineering community. It argues that claims that GenAI can automate qualitative analysis overgeneralize from narrow successes. The authors distinguish research strategies (respondent, field, lab, data) and dimensions of qualitative work (epistemology, coding strategy, granularity, iteration and researcher roles), then review a convenience sample of recent ICSE, CHASE, and CSCW papers and synthesize the emerging literature on LLM-assisted coding, summarization, and thematic analysis. Their central conclusion is that GenAI is useful mainly as a coding or summarization aid—particularly for deductive, low-context annotation—but is not an autonomous qualitative researcher, especially for constructivist and interpretivist studies. The paper closes with quality considerations and a five-item research agenda for benchmarking, interpretive methods, human-AI workflows, standards, and paradigm reconciliation.

Significance. If accepted, the paper provides a timely counterweight to enthusiasm about GenAI in qualitative SE research and offers useful guidance on where automated assistance is currently appropriate. It is heavily cited, carefully hedged in most of its formulations, and explicitly discloses the preliminary nature of its empirical scan, which is accompanied by a replication package—a strength. However, the paper is primarily a synthesis and argument; its only new empirical contribution is a small, convenience-sampled scan whose loading on the central conclusion is limited but still present. The broadest claim—that GenAI tools cannot make sense of novel contexts and are therefore unsuitable for constructivist work—is asserted rather than established; it is a falsifiable empirical hypothesis, but no direct test is offered. If that claim is weakened to 'current evidence does not show,' the paper remains a valuable scoping review and roadmap for the community.

major comments (3)
  1. [§4.2 and §2.1] The statement 'GenAI tools cannot make sense of novel contexts' is load-bearing: it underlies the conclusion that GenAI cannot be an autonomous qualitative researcher and is epistemologically incompatible with constructivist methods. Yet 'sense making' is never operationally defined, and no evidence is given for the categorical impossibility. The paper itself cites LLMs successfully summarizing/translating novel texts and producing themes that are 'less aware of latent interpretations' [30]—a graded, empirical failure mode, not an absence of sense-making. This tension weakens the central claim. Please either (a) define sense-making and specify a falsifiable test, or (b) replace the categorical 'cannot' with 'current evidence does not show,' which is the paper's own formulation in §5. Without this change, the philosophical claim overreaches the evidence.
  2. [§3.1, Table 1] The screening pass uses ChatGPT without reporting precision, recall, or validation against a hand-checked sample; the keyword search and single-author manual check are not reliability-assessed. The claim that 'none of the ICSE or CHASE papers' used GenAI and that CSCW prevalence is 7/209 (3.3%) rests on this unvalidated pipeline. There is also a numeric inconsistency: the text later says '7 CSCW (2.7%)' without a denominator; 7/312 is 2.2% and 7/209 is 3.3%. Since the section is explicitly preliminary, this is fixable by reporting the ChatGPT accuracy on a small gold-standard subset, clarifying the denominator, and labeling the counts as indicative. I would not reject the paper over this, but the numbers need cleaning.
  3. [§4.2, §5] The paper conflates an empirical claim about current LLM performance with a principled/epistemological claim about the nature of a 'statistical neural network'. §4.2 says LLMs 'lack socially embedded sense making' and are 'fundamentally at odds' with constructivist approaches, while §5 summarizes the argument as 'current evidence shows.' These have different scopes. If the stronger claim is intended, the paper should engage with the possibility that constructivist researchers might use GenAI in reflexive workflows and explain why the weak-but-real contextual sensitivity demonstrated in translation/summarization tasks fails to count. If the weaker claim is intended, the §4.2 wording should be aligned. Distinguishing the two would make the central argument more precise and defensible.
minor comments (4)
  1. [Table 1] Use one denominator for all percentages. The same seven papers are described as 3.3% of 209 and 2.7% of an unstated base; 7/312 is 2.2%.
  2. [§3.1] Specify whether the keyword search was applied to full text or only to titles/abstracts, and state the exact search string so the filtering step is independently replicable.
  3. [§4.1] The sentence 'This will dramatically scale up the number of artifacts such studies can analyze' reads as promotional; consider a neutral phrasing consistent with the paper's cautious stance.
  4. [Footnote 6] The aside that some researchers question rigor criteria is interesting but not integrated into the quality-criteria discussion; a sentence linking it to the argument would help.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; self-citations are supporting literature, not load-bearing.

full rationale

This paper is a position/literature review, not a derivation or prediction exercise. It makes no fitted parameters, no equations, and no first-principles claims that could reduce to its own inputs by construction. The central argument—that GenAI support must be tailored to research strategy, data type, and epistemology, and that autonomous GenAI qualitative analysis is not currently supported by evidence—is assembled from external, independently published and falsifiable studies (e.g., Ahmed et al. [20], Shah et al. [35], Montes et al. [30], the mapping study [23], Wen et al. [36]) and from the authors' own descriptive scan in §3.1. The paper does cite the authors' own prior work (Storey et al. [11] for research-strategy distinctions; Ahmed et al. [20] for annotation evidence; Treude & Hata [15] for fairness), but these are used as ordinary literature support rather than as a uniqueness theorem or as the sole justification for the central claim. Removing those self-citations would not collapse the argument, because the conclusion rests mainly on external evidence and on the paper's own review of current usage. The closest concern is §4.2's assertion that 'GenAI tools cannot make sense of novel contexts' and the §2.1 claim that a statistical neural network cannot participate in co-construction of meaning; these are philosophical/empirical assertions without an operational definition of 'sense-making.' That is an unsupported-premise or correctness risk, not circularity, since the assertion is not derived from or equivalent to the paper's inputs. Overall, the paper is self-contained as an evidence-review and position statement; the minor self-citations are non-load-bearing, so the circularity score is 1.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim is an interpretive synthesis rather than a derivation. It leans on a set of domain assumptions about research strategy taxonomies, epistemology, and LLM capabilities. No free parameters or invented entities are introduced.

assumptions (4)
  • domain assumption Storey et al.'s socio-technical framework classifies SE research strategies as respondent, field, lab, and data strategies.
    Used in §2 to structure the entire argument about where GenAI can fit; the paper adopts this framework without justifying it against alternatives.
  • domain assumption Postpositivist and constructivist epistemologies are distinct, and GenAI fits 'most naturally with positivist epistemologies'.
    Invoked in §2.1 and §4.2 to argue an 'epistemological mismatch'; relies on a contested philosophical taxonomy.
  • domain assumption LLMs lack 'socially embedded sense making' and therefore cannot participate in co-construction of meaning.
    Section 4.2; this premise drives the conclusion that GenAI cannot support constructivist methods; asserted, not empirically demonstrated.
  • domain assumption Inter-coder agreement metrics such as Cohen's κ and Krippendorff's α are meaningful measures of GenAI quality in qualitative coding.
    Sections 3.2 and 5 use these metrics as evidence; this presupposes a postpositivist evaluation model that the paper itself notes is contested [2].

how reviews work

0 comments
Cite this review

Pith. "Pith review of GenAI Is No Silver Bullet for Qualitative Research in Software Engineering." pith.science (2026). https://pith.science/paper/6MXDM4YT

@misc{pith2026260308951,
  author       = {Pith},
  title        = {Pith review of: GenAI Is No Silver Bullet for Qualitative Research in Software Engineering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6MXDM4YT}},
  note         = {Machine review of arXiv:2603.08951}
}
read the original abstract

Qualitative research gives rich insights into the quintessentially human aspects of software engineering as a socio-technical system. Qualitative research spans diverse strategies and methods, from interpretivist, in situ observational field studies, to deductive coding of data from mining studies. Advances in large language models and generative AI (GenAI) have prompted claims that artificial intelligence could automate qualitative analysis. Such claims are overgeneralizing from narrow successes. GenAI support must be carefully adapted to the data of interest, but also to the characteristics of a particular research strategy. In this Frontiers of SE paper, we discuss the emerging use of GenAI in relation to the broad spectrum of qualitative research in software engineering. We outline the dimensions of qualitative work in software engineering, review emerging empirical evidence for GenAI assistance, examine the pros and cons of GenAI-mediated qualitative research practices, and revisit qualitative research quality factors, in light of GenAI. Our goal is to inform researchers about the promises and pitfalls of GenAI-assisted qualitative research. We conclude with future plans to advance understanding of its use in software engineering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Empirical Study of Model Context Protocol Applications

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A large-scale study of 1,723 MCP host applications finds partial convergence on file-based configuration (85.2%) and official SDKs (81.1%), but only 37.2% gate tool execution behind a blocking approval step.

Reference graph

Works this paper leans on

38 extracted references · 3 canonical work pages · cited by 1 Pith paper

  1. [30]

    Cristina Martinez Montes, Robert Feldt, Cristina Miguel Martos, Sofia Ouhbi, Shweta Premanandan, and Daniel Graziotin. 2025. Large language models in thematic analysis: prompt engineering, evaluation, and guidelines for qualitative software engineering research. (2025). doi:10.48550/ar xiv.2510.18456

  2. [1]

    Louise H Kidder and Michelle Fine. 1987. Qualitative and quantitative methods: When stories converge. InMultiple Methods in Program Evaluation. Number 35 in New Directions for Program Evaluation. Jossey-Bass. Manuscript submitted to ACM 10 Neil A. Ernst and Christoph Treude

  3. [2]

    Margarete Sandelowski. 1993. Rigor or rigor mortis: the problem of rigor in qualitative research revisited.Advances in Nursing Science, 16, 2, 1–8

  4. [3]

    Blake D. Poland. 1995. Transcription quality as an aspect of rigor in qualitative research.Qualitative Inquiry, 1, 3, (Sept. 1995), 290–310. doi:10.1177 /107780049500100302

  5. [4]

    Siw Elisabeth Hove and Bente Anda. 2005. Experiences from conducting semi-structured interviews in empirical software engineering research. In11th IEEE International Software Metrics Symposium (METRICS’05). IEEE, 10–pp

  6. [5]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology.Qualitative Research in Psychology, 3, 2, (Jan. 2006), 77–101. doi:10.1191/1478088706qp063oa

  7. [6]

    Springer London, 285–311

    2008.Selecting empirical methods for software engineering research.Guide to Advanced Empirical Software Engineering. Springer London, 285–311. isbn: 9781848000445. doi:10.1007/978-1-84800-044-5_11

  8. [7]

    Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering.Empirical software engineering, 14, 2, 131–164

Show all 38 references
  1. [8]

    2011.Self-organizing agile teams: A grounded theory

    Rashina Hoda. 2011.Self-organizing agile teams: A grounded theory. Ph.D. Dissertation. Open Access Te Herenga Waka-Victoria University of Wellington

  2. [9]

    Linda Birt, Suzanne Scott, Debbie Cavers, Christine Campbell, and Fiona Walter. 2016. Member checking: a tool to enhance trustworthiness or merely a nod to validation?Qualitative Health Research, 26, 13, (July 2016), 1802–1811. doi:10.1177/1049732316654870

  3. [10]

    Kirsti Malterud, Volkert Dirk Siersma, and Ann Dorrit Guassora. 2016. Sample size in qualitative interview studies: guided by information power. Qualitative Health Research, 26, 13, (July 2016), 1753–1760. doi:10.1177/1049732315617444

  4. [11]

    Ernst, Courtney Williams, and Eirini Kalliamvakou

    Margaret-Anne Storey, Neil A. Ernst, Courtney Williams, and Eirini Kalliamvakou. 2020. The who, what, how of software engineering research: a socio-technical framework.Empirical Software Engineering, 25, 5, (Aug. 2020), 4097–4129. doi:10.1007/s10664-020-09858-z

  5. [12]

    Vivek Arora, Enrique Larios Vargas, Maurício Aniche, and Arie van Deursen. 2021. Secure software engineering in the financial services: a practitioners’ perspective.arXiv preprint arXiv:2104.03476

  6. [13]

    Shih-Chieh Dai, Aiping Xiong, and Lun-Wei Ku. 2023. LLM-in-the-loop: leveraging large language model for thematic analysis. InFindings of the Association for Computational Linguistics: EMNLP 2023. Houda Bouamor, Juan Pino, and Kalika Bali, (Eds.) Association for Computational ...

  7. [14]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InProceedings of the 40th International Conference on Machine Learning, 28492–28518. https://proceedings.mlr.press/v202...

  8. [15]

    Christoph Treude and Hideaki Hata. 2023. She elicits requirements and he tests: software engineering gender bias in large language models. In 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR). IEEE, 624–629

  9. [16]

    open codes

    John Chen, Alexandros Lotsos, Sihan Cheng, Caiyi Wang, Lexie Zhao, Jessica Hullman, Bruce Sherin, Uri Wilensky, and Michael Horn. 2024. A computational method for measuring" open codes" in qualitative analysis.arXiv preprint arXiv:2411.12142

  10. [17]

    Robert M Davison et al. 2024. The ethics of using generative ai for qualitative data analysis.Information Systems Journal, 34, 5, 1433–1439

  11. [18]

    Yvonne Dittrich, Helen Sharp, and Cleidson de Souza. 2024. Teaching and learning ethnography for software engineering contexts. InHandbook on Teaching Empirical Software Engineering. Springer, 593–630

  12. [19]

    Per Lenberg, Robert Feldt, Lucas Gren, Lars Göran Wallgren Tengberg, Inga Tidefors, and Daniel Graziotin. 2024. Qualitative software engineering research: reflections and guidelines.Journal of Software: Evolution and Process, 36, 6, e2607

  13. [20]

    Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025. Can LLMs replace manual annotation of software engineering artifacts? In2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR). IEEE, 526–538

  14. [21]

    Lama Alqazlan, Zheng Fang, Michael Castelle, and Rob Procter. 2025. A novel, human-in-the-loop computational grounded theory framework for big social data.Big Data & Society, 12, 2, 20539517251347598

  15. [22]

    Sebastian Baltes et al. 2025. Evaluation guidelines for empirical studies in software engineering involving llms.arXiv preprint arXiv:2508.15503

  16. [23]

    Cauã Ferreira Barros, Bruna Borges Azevedo, Valdemar Vicente Graciano Neto, Mohamad Kassab, Marcos Kalinowski, Hugo Alexandre D Do Nascimento, and Michelle CGSP Bandeira. 2025. Large language model for qualitative research: a systematic mapping study. In2025 IEEE/ACM Internati...

  17. [24]

    Eva Blondeel, Patricia Everaert, and Evelien Opdecam. 2025. A practical guide to implementing chatgpt as a secondary coder in qualitative research.International Journal of Accounting Information Systems, 56, (Dec. 2025), 100754. doi:10.1016/j.accinf.2025.100754

  18. [25]

    Breno Felix De Sousa, Ronnie de Souza Santos, and Kiev Gama. 2025. Integrating positionality statements in empirical software engineering research. In2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, (Ma...

  19. [26]

    Zackary Okun Dunivin. 2025. Scaling hermeneutics: a guide to qualitative coding with llms for reflexive content analysis.EPJ Data Science, 14, 1, 28

  20. [27]

    Fine Michelle

    Tanisha Jowsey, Virginia Braun, Victoria Clarke, Deborah Lupton, and et al. Fine Michelle. 2025. We reject the use of generative artificial intelligence for reflexive qualitative research.Qualitative Inquiry, (Dec. 2025). doi:10.1177/10778004251401851

  21. [28]

    Daye Kang, Zhuolun Han, Jiahe Tian, Muhan Zhang, and Jeffrey M Rzeszotarski. 2025. Themeviz: understanding the effect of human-ai collaboration in theme development with an llm-enhanced interactive visual system.Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct....

  22. [29]

    Yanheng Li, Da Wang, and Yuping Wang. 2025. Understanding voter fraud misinformation videos during the 2024 taiwan election on youtube. Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct. 2025), 1–40. doi:10.1145/3757697

  23. [31]

    Duc Cuong Nguyen and Catherine Welch. 2025. Generative artificial intelligence in qualitative data analysis: analyzing—or just chatting? Organizational Research Methods, 29, 1, (Sept. 2025), 3–39. doi:10.1177/10944281251377154

  24. [32]

    Kien Nguyen-Trung. 2025. Chatgpt in thematic analysis: can ai become a research assistant in qualitative research?Quality & Quantity, 59, 6, (June 2025), 4945–4978. doi:10.1007/s11135-025-02165-z

  25. [33]

    Shruti Phadke. 2025. Exit stories: using reddit self-disclosures to understand disengagement from problematic communities.Proceedings of the ACM on Human-Computer Interaction, 9, 7, (Oct. 2025), 1–27. doi:10.1145/3757592

  26. [34]

    Carolyn Seaman, Rashina Hoda, and Robert Feldt. 2025. Qualitative research methods in software engineering: past, present, and future.IEEE Transactions on Software Engineering

  27. [35]

    Syed Tauhid Ullah Shah, Mohamad Hussein, Ann Barcomb, and Mohammad Moshirpour. 2025. From inductive to deductive: llms-based qualitative data analysis in requirements engineering.arXiv preprint arXiv:2504.19384

  28. [36]

    Chuanchi Wen, Paul Clough, Rachel Paton, and Rebecca Middleton. 2025. Leveraging large language models for thematic analysis: a case study in the charity sector.AI & Society, (Aug. 2025). doi:10.1007/s00146-025-02487-4

  29. [37]

    Sargam Yadav, Abhishek Kaushik, and Asifa Mehmood Qureshi. 2025. Thematic analysis of expert opinions on the use of large language models in software development.Applied AI Letters, 6, 3, e127

  30. [38]

    Hanmo You, Zan Wang, Bin Lin, and Junjie Chen. 2025. Navigating the testing of evolving deep learning systems: an exploratory interview study. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, (Apr. 2025), 2726–2738. doi:10.1109/icse55347.2025...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.