REVIEW 2 major objections 5 minor 1 cited by
Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Large language models are now used mainly for coding and thematic analysis in qualitative research, with efficiency gains offset by output variability and privacy risks.
desk verdict Competent mapping study of LLM-assisted qualitative analysis; findings are useful but not new, and the corpus issues are real but not fatal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the systematic mapping study protocol. A search string and manual searches across software engineering venues and methodology journals produced 2,574 candidate records, which three researchers reduced by title and abstract screening under five exclusion criteria and one inclusion criterion to 20 full papers. The extracted data were then synthesized by thematic analysis, producing the classification tables that carry the argument: qualitative methods, LLM tools, techniques, data types, benefits, and limitations. Those tables are the machinery because every conclusion in the paper is read off them.
What would settle it
A replication of the search that relaxes the five-page minimum, includes non-English work, and exports more than the top 1,000 ACM results would falsify the mapping if it produced a substantial set of additional primary studies whose applications, benefits, or limitations differ from the clusters reported here.
Extended reading notes
Core claim
The central discovery is a descriptive map, not a new technique. Analyzing 20 primary studies, the authors find that LLMs are being used for open and deductive coding, thematic analysis, grounded theory, screening, topic modeling, content analysis, vignette analysis, and critical review, with coding and thematic analysis the most common. The dominant tools are ChatGPT, GPT-3.5, and GPT-4, and the dominant technique is prompt engineering, with fine-tuning and few-shot prompting appearing less often. Against those applications, the reported benefits cluster into theme and pattern identification, efficiency, coding support, autonomy for beginners, enhanced collaboration, and triangulation; the reported limitations cluster into consistency and hallucination problems, shallow high-level comprehension, ethics, privacy and transparency deficits, technical constraints such as token limits, and dependency risks. The paper's conclusion is that this evidence supports a division of labor where LLMs accelerate labor-intensive analysis but human expertise remains responsible for interpretation, and that preliminary recommendations, disclose the model and version, experiment with prompts, protect data, keep humans in the loop, balance with traditional methods, follow AI ethics guidelines, and discuss validity threats contextually, should guide integration into software engineering research.
Load-bearing premise
The conclusions assume that the 20 papers surviving the screening are representative of the wider literature on LLM-assisted qualitative analysis.
Editorial extensions
If this is right
- For empirical software engineering, LLMs can absorb the most time-consuming parts of qualitative analysis, initial open coding, categorization, and screening, letting researchers spend their effort on interpretation.
- The dominant practice of prompt engineering becomes a methodological skill: the type of prompt, instructional, persona-based, or chain-of-thought, changes the quality of the analysis, so reporting prompt strategies should become part of study design.
- Because output variability and hallucinations are persistent, LLM-assisted analyses need an explicit validation step, such as triangulation or comparison with human coding.
- Data protection has to be designed into the workflow before any participant or company data is entered into a proprietary model, not added afterward.
- The recommendations imply that reporting standards for LLM-assisted qualitative studies should include the model and version, prompt choices, and validity-threat discussion.
Reading between the lines
- A testable extension of the triangulation finding: treat repeated runs of the same LLM on the same data as an ensemble of pseudo-coders and measure inter-run agreement; if agreement is low, a reliability threshold could be set before human review.
- The privacy limitation has an engineering consequence the authors do not draw: on-premises or open-weight models with strict logging controls would preserve the efficiency benefits while removing the main data-exposure risk for proprietary software-engineering data.
- A direct corollary for guidelines is that, since model versions change rapidly, a reporting checklist covering model, version, date, prompts, output samples, and human edits could let later readers judge reproducibility, a concern implicit in the paper's disclosure recommendation.
- If coding quality becomes comparable to human coders for well-defined tasks, the bottleneck in qualitative software engineering research shifts from coding effort to construct validity, whether themes found by either humans or LLMs actually answer the research question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic mapping study of how large language models (LLMs) are used to support qualitative analysis, with the stated aim of informing empirical software engineering research. The authors searched ACM, IEEE, Scopus, and selected venues; after title/abstract screening and full-text review they included 20 studies (the abstract says 21). They extract data on qualitative analysis methods, LLM tools, prompting techniques, data types, reported benefits, and reported limitations, and synthesize these into tables (Tables II-VII). The central findings are that LLMs are primarily used for coding, thematic analysis, and data categorization; benefits include efficiency gains and support for novice researchers; limitations include output variability, limited interpretive depth, and ethical/privacy concerns. The paper ends with a set of recommendations for researchers and a threats-to-validity discussion.
Significance. If the corpus is representative, the paper offers a useful early map of an emerging area and a set of plausible, though preliminary, recommendations for software engineering researchers. Methodological strengths include adherence to established systematic-review guidelines, three-researcher screening, two-researcher extraction, an openly available dataset, and an honest threats-to-validity section. The main weakness is that the descriptive map and the derived recommendations rest on a small corpus that is subject to acknowledged but unquantified selection limitations, especially the ACM export cap and title/abstract screening. The paper does not claim formal predictive power, and no circularity issue arises because the findings are obtained through thematic synthesis rather than by fitting a model to its own outputs.
major comments (2)
- [Section V-C (External Validity); Section III (Search Strategy)] The acknowledged ACM export cap (approximately 3,000 results, only the top 1,000 exportable) means that a substantial portion of ACM hits was never screened. Because the paper's central contribution is a descriptive map of LLM-assisted qualitative analysis, this blind spot could systematically exclude studies with different application areas or different reported limitations, which would directly affect the synthesis in Tables II-VII and the recommendations in Section V-B. The authors acknowledge the risk but do not provide a quantitative mitigation. Please add a PRISMA-style flow diagram with the number of studies excluded at each criterion, state how many included studies came from each database and from the manual searches, and discuss whether any identified themes are supported by only one database or one source. This information is necessary for readers to assess the representativeness of the 20-study corpus.
- [Section III (Selection Process)] The screening was based solely on titles and abstracts before applying the inclusion criterion, which is a known risk in a field where LLM use may be described only in the full text. The authors acknowledge this risk in Section V-C but do not report any calibration, such as a full-text check of a random sample of excluded papers or an inter-reviewer agreement measure. Please add a quantitative or at least systematic qualitative account of screening reliability (for example, Cohen's kappa or a description of disagreements and resolutions). Without this, the reliability of the selection step is difficult to assess, and the load-bearing assumption that the 20 included studies represent the literature is left unsupported.
minor comments (5)
- [Abstract; Section III; Section IV; Table I] The abstract states that 21 relevant studies were analyzed, but the body, the method section, and Table I consistently report 20 studies. Please correct the abstract to match the actual corpus.
- [Table VII] In the last row of Table VII, the article ID is given as "A20" while all other entries use the zero-padded format "A020". Please make this consistent.
- [Section I] The sentence "This approach is not aimed at deeply exploring specific aspects of a research problem" appears to be a typo; it contradicts the surrounding discussion of qualitative research. It should likely read "This approach is aimed at deeply exploring..." or "This approach is not limited to...".
- [Section V-B (Recommendations)] Recommendations 2 and 3 both state essentially the same data-protection advice: researcher should anonymize data and avoid direct input of sensitive information into LLMs. Please merge these into a single recommendation or differentiate them clearly to avoid redundancy.
- [Section VII (Dataset Availability)] The dataset link is described only as "here" with no actual URL or persistent identifier. Please provide a working DOI or URL so that the promised artifact is accessible.
Circularity Check
No significant circularity: the paper is a descriptive mapping study whose findings summarize the included literature rather than deriving predictions from fitted inputs.
full rationale
This paper is a systematic mapping study, not a derivational or predictive study. There are no fitted parameters, no equations, and no quantities predicted from inputs. The central claims—that LLMs are used for coding, thematic analysis, and data categorization, with certain benefits and limitations—are synthesized from the 20 included primary studies through thematic analysis. The synthesis does not reduce to the selection criteria by construction: although the search string includes the term 'prompt engineering,' the inclusion criterion INC1 requires papers to focus on the use of LLMs in qualitative research methodologies generally, and the specific reported benefits and limitations (Tables VI and VII) are not forced by the search or inclusion criteria. The recommendations in Section V-B are author-derived guidelines based on identified limitations and established qualitative-research practices, not mathematical implications of the input data. No self-citation is load-bearing: the authors do not cite their own prior work to justify the study's premise or conclusions. The paper self-reports threats to validity, including the ACM export limit and title/abstract screening, which are legitimate concerns about representativeness but are not circularity. The abstract's mention of '21 relevant studies' versus the 20 studies reported in the method and findings is an internal inconsistency, not a circularity issue. Overall, the derivation chain is descriptive and externally grounded in the reviewed literature, so no circular step can be identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The 20 included primary studies accurately report their use of LLMs, and the authors' data extraction faithfully captures those reports.
- domain assumption The search and selection strategy produced a corpus representative of research on LLM-assisted qualitative analysis in software engineering.
- domain assumption Thematic analysis as described by Cruzes and Dyba is an appropriate synthesis method for mapping-study data.
Cite this review
Pith. "Pith review of Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering." pith.science (2026). https://pith.science/paper/ID5ANE6J
@misc{pith2026241206564,
author = {Pith},
title = {Pith review of: Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering},
year = {2026},
howpublished = {\url{https://pith.science/paper/ID5ANE6J}},
note = {Machine review of arXiv:2412.06564}
}
read the original abstract
The use of large language models (LLMs) for qualitative analysis is gaining attention in various fields, including software engineering, where qualitative methods are essential for understanding human and social factors. This study aimed to investigate how LLMs are currently used in qualitative analysis and their potential applications in software engineering research, focusing on the benefits, limitations, and practices associated with their use. A systematic mapping study was conducted, analyzing 21 relevant studies to explore reported uses of LLMs for qualitative analysis. The findings indicate that LLMs are primarily used for tasks such as coding, thematic analysis, and data categorization, offering benefits like increased efficiency and support for new researchers. However, limitations such as output variability, challenges in capturing nuanced perspectives, and ethical concerns related to privacy and transparency were also identified. The study emphasizes the need for structured strategies and guidelines to optimize LLM use in qualitative research within software engineering, enhancing their effectiveness while addressing ethical considerations. While LLMs show promise in supporting qualitative analysis, human expertise remains crucial for interpreting data, and ongoing exploration of best practices will be vital for their successful integration into empirical software engineering research.
Forward citations
Cited by 1 Pith paper
-
Large Language Model for Qualitative Research -- A Systematic Mapping Study
A systematic map of eight studies shows LLM-assisted qualitative analysis is mostly comparable to manual methods, with prompt dependence and hallucination as recurring limitations.
Reference graph
Works this paper leans on
-
[1]
Qualitative methods: what are they and why use them?
S. Sofaer, “Qualitative methods: what are they and why use them?” Health services research, vol. 34, no. 5 Pt 2, p. 1101, 1999
work page 1999
-
[2]
C. B. Seaman, “Qualitative methods,” in Guide to advanced empirical software engineering. Springer, 2008, pp. 35–62. 7
work page 2008
-
[3]
Corbin and A
J. Corbin and A. Strauss, Basics of qualitative research . sage, 2015, vol. 14
2015
-
[4]
Qualitative analysis: What it is and how to begin,
M. Sandelowski, “Qualitative analysis: What it is and how to begin,” Research in nursing & health , vol. 18, no. 4, pp. 371–375, 1995
work page 1995
-
[5]
K. Charmaz, “Grounded theory,” Qualitative psychology: A practical guide to research methods , vol. 3, pp. 53–84, 2015
work page 2015
-
[6]
Grounded theory in software engineering research: a critical review and guidelines,
K.-J. Stol, P. Ralph, and B. Fitzgerald, “Grounded theory in software engineering research: a critical review and guidelines,” in Proceedings of the 38th International conference on software engineering , 2016, pp. 120–131
work page 2016
- [7]
-
[8]
The role of ethnographic studies in empirical software engineering,
H. Sharp, Y . Dittrich, and C. R. De Souza, “The role of ethnographic studies in empirical software engineering,” IEEE Transactions on Soft- ware Engineering, vol. 42, no. 8, pp. 786–804, 2016
work page 2016
Show all 39 references
-
[9]
Action research,
D. E. Avison, F. Lau, M. D. Myers, and P. A. Nielsen, “Action research,” Communications of the ACM , vol. 42, no. 1, pp. 94–97, 1999
1999
-
[10]
Action research can swing the balance in experimental software engineering,
P. S. M. Dos Santos and G. H. Travassos, “Action research can swing the balance in experimental software engineering,” in Advances in computers. Elsevier, 2011, vol. 83, pp. 205–276
2011
-
[11]
Ezzy, Qualitative analysis
D. Ezzy, Qualitative analysis. Routledge, 2013
2013
-
[12]
Carrying out qualitative analysis,
J. Ritchie, L. Spencer, W. O’Connor et al. , “Carrying out qualitative analysis,” Qualitative research practice: A guide for social science students and researchers, vol. 2003, pp. 219–62, 2003
2003
-
[13]
Qualitative software engineering research: Reflections and guidelines,
P. Lenberg, R. Feldt, L. Gren, L. G. Wallgren Tengberg, I. Tidefors, and D. Graziotin, “Qualitative software engineering research: Reflections and guidelines,” Journal of Software: Evolution and Process , vol. 36, no. 6, p. e2607, 2024
2024
-
[14]
Qualitative methods in empirical studies of software engineering,
C. B. Seaman, “Qualitative methods in empirical studies of software engineering,” IEEE Transactions on software engineering, vol. 25, no. 4, pp. 557–572, 1999
1999
-
[15]
Qualitative analysis,
G. R. Gibbs, “Qualitative analysis,” Qualitative Data Analysis , p. 277, 2014
2014
-
[16]
Using qualitative analysis software to facilitate qualita- tive data analysis,
V . Talanquer, “Using qualitative analysis software to facilitate qualita- tive data analysis,” in Tools of chemistry education research . ACS Publications, 2014, pp. 83–95
2014
-
[17]
Computer software and qualitative analysis: A reassess- ment,
R. TESCH, “Computer software and qualitative analysis: A reassess- ment,” New Technology in Sociology: Practical Applications in Research and Work, 2019
2019
-
[18]
Learn for yourself: The self-learning tools for qualitative analysis software packages,
F. Freitas, J. Ribeiro, C. Brand ˜ao, L. P. Reis, F. N. de Souza, and A. P. Costa, “Learn for yourself: The self-learning tools for qualitative analysis software packages,” Digital Education Review , no. 32, pp. 97– 117, 2017
2017
-
[19]
Tools for analyzing qualitative data: The history and relevance of qualitative data analysis software,
L. S. Gilbert, K. Jackson, and S. Di Gregorio, “Tools for analyzing qualitative data: The history and relevance of qualitative data analysis software,” Handbook of research on educational communications and technology, pp. 221–236, 2014
2014
-
[20]
Artificial intelligence and qualitative research: The promise and perils of large language model (llm)‘assistance’,
J. Roberts, M. Baker, and J. Andrew, “Artificial intelligence and qualitative research: The promise and perils of large language model (llm)‘assistance’,” Critical Perspectives on Accounting , vol. 99, p. 102722, 2024
2024
-
[21]
An examination of the use of large language models to aid analysis of textual data,
R. H. Tai, L. R. Bentley, X. Xia, J. M. Sitt, S. C. Fankhauser, A. M. Chicas-Mosier, and B. G. Monteith, “An examination of the use of large language models to aid analysis of textual data,” International Journal of Qualitative Methods , vol. 23, p. 16094069241231168, 2024
2024
-
[22]
Ai and human reasoning: Qualitative research in the age of large language models,
M. Bano, D. Zowghi, and J. Whittle, “Ai and human reasoning: Qualitative research in the age of large language models,” The AI Ethics Journal, vol. 3, no. 1, 2023
2023
-
[23]
Large language models for qualitative research in software engineering: exploring opportunities and challenges,
M. Bano, R. Hoda, D. Zowghi, and C. Treude, “Large language models for qualitative research in software engineering: exploring opportunities and challenges,” Automated Software Engineering , vol. 31, no. 1, p. 8, 2024
2024
-
[24]
Ai and the transformation of social science research,
I. Grossmann, M. Feinberg, D. C. Parker, N. A. Christakis, P. E. Tetlock, and W. A. Cunningham, “Ai and the transformation of social science research,” Science, vol. 380, no. 6650, pp. 1108–1109, 2023
2023
-
[25]
Chatgpt and other large language models are double-edged swords,
Y . Shen, L. Heacock, J. Elias, K. D. Hentel, B. Reig, G. Shih, and L. Moy, “Chatgpt and other large language models are double-edged swords,” p. e230163, 2023
2023
-
[26]
A survey on evaluation of large language models,
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024
2024
-
[27]
Large language models in qualitative research: Can we do the data justice?
H. Schroeder, M. A. L. Qu ´er´e, C. Randazzo, D. Mimno, and S. Schoenebeck, “Large language models in qualitative research: Can we do the data justice?” arXiv preprint arXiv:2410.07362 , 2024
2024 arXiv
-
[28]
Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis,
S. A. Gebreegziabher, Z. Zhang, X. Tang, Y . Meng, E. L. Glassman, and T. J.-J. Li, “Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–19
2023
-
[29]
Collabcoder: a lower-barrier, rigorous workflow for inductive collabo- rative qualitative analysis with large language models,
J. Gao, Y . Guo, G. Lim, T. Zhang, Z. Zhang, T. J.-J. Li, and S. T. Perrault, “Collabcoder: a lower-barrier, rigorous workflow for inductive collabo- rative qualitative analysis with large language models,” in Proceedings of the CHI Conference on Human Factors in Computing Sys...
2024
-
[30]
Dolma: An open corpus of three trillion tokens for language model pretraining research,
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y . Elazar et al. , “Dolma: An open corpus of three trillion tokens for language model pretraining research,” arXiv preprint arXiv:2402.00159 , 2024
2024 arXiv
-
[31]
Fighting reviewer fatigue or amplifying bias? considerations and recommendations for use of chatgpt and other large language models in scholarly peer review,
M. Hosseini and S. P. Horbach, “Fighting reviewer fatigue or amplifying bias? considerations and recommendations for use of chatgpt and other large language models in scholarly peer review,” Research integrity and peer review, vol. 8, no. 1, p. 4, 2023
2023
-
[32]
The dangers of using proprietary llms for research,
´E. Ollion, R. Shen, A. Macanovic, and A. Chatelain, “The dangers of using proprietary llms for research,” Nature Machine Intelligence, vol. 6, no. 1, pp. 4–5, 2024
2024
-
[33]
Evidence-based soft- ware engineering,
B. A. Kitchenham, T. Dyba, and M. Jorgensen, “Evidence-based soft- ware engineering,” in Proceedings. 26th International Conference on Software Engineering. IEEE, 2004, pp. 273–281
2004
-
[34]
Recommended steps for thematic synthesis in software engineering,
D. S. Cruzes and T. Dyba, “Recommended steps for thematic synthesis in software engineering,” in 2011 international symposium on empirical software engineering and measurement . IEEE, 2011, pp. 275–284
2011
-
[35]
Prompt engineering in medical education,
T. Heston and C. Khun, “Prompt engineering in medical education,” International Medical Education , vol. 2, pp. 198–205, 8 2023
2023
-
[36]
A prompt pattern catalog to enhance prompt engineering with chatgpt,
J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. El- nashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” arXiv preprint arXiv:2302.11382, 2023
2023 arXiv
-
[37]
Guidelines for conducting and reporting case study research in software engineering,
P. Runeson and M. H ¨ost, “Guidelines for conducting and reporting case study research in software engineering,” Empirical software engineering, vol. 14, pp. 131–164, 2009
2009
-
[38]
Threats to validity in software engineering research: A critical reflection,
R. Verdecchia, E. Engstr ¨om, P. Lago, P. Runeson, and Q. Song, “Threats to validity in software engineering research: A critical reflection,” Information and Software Technology , vol. 164, p. 107329, 2023
2023
-
[39]
Assessing the impact of prompting, persona, and chain of thought methods on chatgpt’s arithmetic capabilities,
Y . Chen, C. Wong, H. Yang, J. Aguenza, S. Bhujangari, B. Vu, X. Lei, A. Prasad, M. Fluss, E. Phuong et al. , “Assessing the impact of prompting, persona, and chain of thought methods on chatgpt’s arithmetic capabilities,” arXiv preprint arXiv:2312.15006 , 2023. 8
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.