Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering

T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Large language models are now used mainly for coding and thematic analysis in qualitative research, with efficiency gains offset by output variability and privacy risks.

desk verdict Competent mapping study of LLM-assisted qualitative analysis; findings are useful but not new, and the corpus issues are real but not fatal. read the letter →

arxiv 2412.06564 v4 pith:ID5ANE6J submitted 2024-12-09 cs.SE

classification cs.SE
keywords largelanguagemodelsqualitativeanalysissystematicmappingstudysoftwareengineeringthematicpromptempiricalresearchethics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a systematic mapping study asking how large language models are actually used in qualitative analysis and what that means for empirical software engineering. It argues that the current literature concentrates LLM use on coding, thematic analysis, and data categorization, with ChatGPT-family models and prompt engineering dominating practice. The reported benefits are real but bounded: faster, cheaper, less cognitively demanding analysis and a lower barrier for novice researchers. The reported limitations are equally consistent: outputs vary run to run, models miss subtle context, and privacy and transparency problems are unresolved. The paper concludes that LLMs can support, not replace, human interpretation, and that structured guidelines for disclosure, prompting, data protection, and human oversight are the immediate next step.

What carries the argument

The load-bearing mechanism is the systematic mapping study protocol. A search string and manual searches across software engineering venues and methodology journals produced 2,574 candidate records, which three researchers reduced by title and abstract screening under five exclusion criteria and one inclusion criterion to 20 full papers. The extracted data were then synthesized by thematic analysis, producing the classification tables that carry the argument: qualitative methods, LLM tools, techniques, data types, benefits, and limitations. Those tables are the machinery because every conclusion in the paper is read off them.

What would settle it

A replication of the search that relaxes the five-page minimum, includes non-English work, and exports more than the top 1,000 ACM results would falsify the mapping if it produced a substantial set of additional primary studies whose applications, benefits, or limitations differ from the clusters reported here.

Watch

Extended reading notes

Core claim

The central discovery is a descriptive map, not a new technique. Analyzing 20 primary studies, the authors find that LLMs are being used for open and deductive coding, thematic analysis, grounded theory, screening, topic modeling, content analysis, vignette analysis, and critical review, with coding and thematic analysis the most common. The dominant tools are ChatGPT, GPT-3.5, and GPT-4, and the dominant technique is prompt engineering, with fine-tuning and few-shot prompting appearing less often. Against those applications, the reported benefits cluster into theme and pattern identification, efficiency, coding support, autonomy for beginners, enhanced collaboration, and triangulation; the reported limitations cluster into consistency and hallucination problems, shallow high-level comprehension, ethics, privacy and transparency deficits, technical constraints such as token limits, and dependency risks. The paper's conclusion is that this evidence supports a division of labor where LLMs accelerate labor-intensive analysis but human expertise remains responsible for interpretation, and that preliminary recommendations, disclose the model and version, experiment with prompts, protect data, keep humans in the loop, balance with traditional methods, follow AI ethics guidelines, and discuss validity threats contextually, should guide integration into software engineering research.

Load-bearing premise

The conclusions assume that the 20 papers surviving the screening are representative of the wider literature on LLM-assisted qualitative analysis.

Editorial extensions

If this is right

  • For empirical software engineering, LLMs can absorb the most time-consuming parts of qualitative analysis, initial open coding, categorization, and screening, letting researchers spend their effort on interpretation.
  • The dominant practice of prompt engineering becomes a methodological skill: the type of prompt, instructional, persona-based, or chain-of-thought, changes the quality of the analysis, so reporting prompt strategies should become part of study design.
  • Because output variability and hallucinations are persistent, LLM-assisted analyses need an explicit validation step, such as triangulation or comparison with human coding.
  • Data protection has to be designed into the workflow before any participant or company data is entered into a proprietary model, not added afterward.
  • The recommendations imply that reporting standards for LLM-assisted qualitative studies should include the model and version, prompt choices, and validity-threat discussion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension of the triangulation finding: treat repeated runs of the same LLM on the same data as an ensemble of pseudo-coders and measure inter-run agreement; if agreement is low, a reliability threshold could be set before human review.
  • The privacy limitation has an engineering consequence the authors do not draw: on-premises or open-weight models with strict logging controls would preserve the efficiency benefits while removing the main data-exposure risk for proprietary software-engineering data.
  • A direct corollary for guidelines is that, since model versions change rapidly, a reporting checklist covering model, version, date, prompts, output samples, and human edits could let later readers judge reproducibility, a concern implicit in the paper's disclosure recommendation.
  • If coding quality becomes comparable to human coders for well-defined tasks, the bottleneck in qualitative software engineering research shifts from coding effort to construct validity, whether themes found by either humans or LLMs actually answer the research question.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports a systematic mapping study of how large language models (LLMs) are used to support qualitative analysis, with the stated aim of informing empirical software engineering research. The authors searched ACM, IEEE, Scopus, and selected venues; after title/abstract screening and full-text review they included 20 studies (the abstract says 21). They extract data on qualitative analysis methods, LLM tools, prompting techniques, data types, reported benefits, and reported limitations, and synthesize these into tables (Tables II-VII). The central findings are that LLMs are primarily used for coding, thematic analysis, and data categorization; benefits include efficiency gains and support for novice researchers; limitations include output variability, limited interpretive depth, and ethical/privacy concerns. The paper ends with a set of recommendations for researchers and a threats-to-validity discussion.

Significance. If the corpus is representative, the paper offers a useful early map of an emerging area and a set of plausible, though preliminary, recommendations for software engineering researchers. Methodological strengths include adherence to established systematic-review guidelines, three-researcher screening, two-researcher extraction, an openly available dataset, and an honest threats-to-validity section. The main weakness is that the descriptive map and the derived recommendations rest on a small corpus that is subject to acknowledged but unquantified selection limitations, especially the ACM export cap and title/abstract screening. The paper does not claim formal predictive power, and no circularity issue arises because the findings are obtained through thematic synthesis rather than by fitting a model to its own outputs.

major comments (2)
  1. [Section V-C (External Validity); Section III (Search Strategy)] The acknowledged ACM export cap (approximately 3,000 results, only the top 1,000 exportable) means that a substantial portion of ACM hits was never screened. Because the paper's central contribution is a descriptive map of LLM-assisted qualitative analysis, this blind spot could systematically exclude studies with different application areas or different reported limitations, which would directly affect the synthesis in Tables II-VII and the recommendations in Section V-B. The authors acknowledge the risk but do not provide a quantitative mitigation. Please add a PRISMA-style flow diagram with the number of studies excluded at each criterion, state how many included studies came from each database and from the manual searches, and discuss whether any identified themes are supported by only one database or one source. This information is necessary for readers to assess the representativeness of the 20-study corpus.
  2. [Section III (Selection Process)] The screening was based solely on titles and abstracts before applying the inclusion criterion, which is a known risk in a field where LLM use may be described only in the full text. The authors acknowledge this risk in Section V-C but do not report any calibration, such as a full-text check of a random sample of excluded papers or an inter-reviewer agreement measure. Please add a quantitative or at least systematic qualitative account of screening reliability (for example, Cohen's kappa or a description of disagreements and resolutions). Without this, the reliability of the selection step is difficult to assess, and the load-bearing assumption that the 20 included studies represent the literature is left unsupported.
minor comments (5)
  1. [Abstract; Section III; Section IV; Table I] The abstract states that 21 relevant studies were analyzed, but the body, the method section, and Table I consistently report 20 studies. Please correct the abstract to match the actual corpus.
  2. [Table VII] In the last row of Table VII, the article ID is given as "A20" while all other entries use the zero-padded format "A020". Please make this consistent.
  3. [Section I] The sentence "This approach is not aimed at deeply exploring specific aspects of a research problem" appears to be a typo; it contradicts the surrounding discussion of qualitative research. It should likely read "This approach is aimed at deeply exploring..." or "This approach is not limited to...".
  4. [Section V-B (Recommendations)] Recommendations 2 and 3 both state essentially the same data-protection advice: researcher should anonymize data and avoid direct input of sensitive information into LLMs. Please merge these into a single recommendation or differentiate them clearly to avoid redundancy.
  5. [Section VII (Dataset Availability)] The dataset link is described only as "here" with no actual URL or persistent identifier. Please provide a working DOI or URL so that the promised artifact is accessible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a descriptive mapping study whose findings summarize the included literature rather than deriving predictions from fitted inputs.

full rationale

This paper is a systematic mapping study, not a derivational or predictive study. There are no fitted parameters, no equations, and no quantities predicted from inputs. The central claims—that LLMs are used for coding, thematic analysis, and data categorization, with certain benefits and limitations—are synthesized from the 20 included primary studies through thematic analysis. The synthesis does not reduce to the selection criteria by construction: although the search string includes the term 'prompt engineering,' the inclusion criterion INC1 requires papers to focus on the use of LLMs in qualitative research methodologies generally, and the specific reported benefits and limitations (Tables VI and VII) are not forced by the search or inclusion criteria. The recommendations in Section V-B are author-derived guidelines based on identified limitations and established qualitative-research practices, not mathematical implications of the input data. No self-citation is load-bearing: the authors do not cite their own prior work to justify the study's premise or conclusions. The paper self-reports threats to validity, including the ACM export limit and title/abstract screening, which are legitimate concerns about representativeness but are not circularity. The abstract's mention of '21 relevant studies' versus the 20 studies reported in the method and findings is an internal inconsistency, not a circularity issue. Overall, the derivation chain is descriptive and externally grounded in the reviewed literature, so no circular step can be identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims of this mapping study rest on the completeness of the literature search, the fidelity of data extraction from the 20 primary studies, and the appropriateness of thematic synthesis. No free parameters or invented entities are involved because the paper does not fit a model or introduce new constructs.

assumptions (3)
  • domain assumption The 20 included primary studies accurately report their use of LLMs, and the authors' data extraction faithfully captures those reports.
    The synthesis treats the primary studies' self-reports as ground truth; if extraction or reporting is inaccurate, the mapped findings would be distorted. Invoked throughout Section IV.
  • domain assumption The search and selection strategy produced a corpus representative of research on LLM-assisted qualitative analysis in software engineering.
    This is the main external-validity premise. The five exclusion criteria, including a 5-page minimum and download-ability, plus the ACM top-1000 export limit, could systematically exclude relevant studies. See Section III and V-C.
  • domain assumption Thematic analysis as described by Cruzes and Dyba is an appropriate synthesis method for mapping-study data.
    The authors apply an established SE method [34]; if thematic synthesis is unsuitable, the resulting themes and recommendations may not be valid. Section III Extraction, Analysis, and Synthesis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering." pith.science (2026). https://pith.science/paper/ID5ANE6J

@misc{pith2026241206564,
  author       = {Pith},
  title        = {Pith review of: Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ID5ANE6J}},
  note         = {Machine review of arXiv:2412.06564}
}
read the original abstract

The use of large language models (LLMs) for qualitative analysis is gaining attention in various fields, including software engineering, where qualitative methods are essential for understanding human and social factors. This study aimed to investigate how LLMs are currently used in qualitative analysis and their potential applications in software engineering research, focusing on the benefits, limitations, and practices associated with their use. A systematic mapping study was conducted, analyzing 21 relevant studies to explore reported uses of LLMs for qualitative analysis. The findings indicate that LLMs are primarily used for tasks such as coding, thematic analysis, and data categorization, offering benefits like increased efficiency and support for new researchers. However, limitations such as output variability, challenges in capturing nuanced perspectives, and ethical concerns related to privacy and transparency were also identified. The study emphasizes the need for structured strategies and guidelines to optimize LLM use in qualitative research within software engineering, enhancing their effectiveness while addressing ethical considerations. While LLMs show promise in supporting qualitative analysis, human expertise remains crucial for interpreting data, and ongoing exploration of best practices will be vital for their successful integration into empirical software engineering research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Model for Qualitative Research -- A Systematic Mapping Study

    cs.CL 2024-11 conditional novelty 3.0 of 10

    A systematic map of eight studies shows LLM-assisted qualitative analysis is mostly comparable to manual methods, with prompt dependence and hallucination as recurring limitations.

Reference graph

Works this paper leans on

39 extracted references · 26 canonical work pages · cited by 1 Pith paper

  1. [1]

    Qualitative methods: what are they and why use them?

    S. Sofaer, “Qualitative methods: what are they and why use them?” Health services research, vol. 34, no. 5 Pt 2, p. 1101, 1999

  2. [2]

    Qualitative methods,

    C. B. Seaman, “Qualitative methods,” in Guide to advanced empirical software engineering. Springer, 2008, pp. 35–62. 7

  3. [3]

    Corbin and A

    J. Corbin and A. Strauss, Basics of qualitative research . sage, 2015, vol. 14

  4. [4]

    Qualitative analysis: What it is and how to begin,

    M. Sandelowski, “Qualitative analysis: What it is and how to begin,” Research in nursing & health , vol. 18, no. 4, pp. 371–375, 1995

  5. [5]

    Grounded theory,

    K. Charmaz, “Grounded theory,” Qualitative psychology: A practical guide to research methods , vol. 3, pp. 53–84, 2015

  6. [6]

    Grounded theory in software engineering research: a critical review and guidelines,

    K.-J. Stol, P. Ralph, and B. Fitzgerald, “Grounded theory in software engineering research: a critical review and guidelines,” in Proceedings of the 38th International conference on software engineering , 2016, pp. 120–131

  7. [7]

    Brewer, Ethnography

    J. Brewer, Ethnography. McGraw-Hill Education (UK), 2000

  8. [8]

    The role of ethnographic studies in empirical software engineering,

    H. Sharp, Y . Dittrich, and C. R. De Souza, “The role of ethnographic studies in empirical software engineering,” IEEE Transactions on Soft- ware Engineering, vol. 42, no. 8, pp. 786–804, 2016

Show all 39 references
  1. [9]

    Action research,

    D. E. Avison, F. Lau, M. D. Myers, and P. A. Nielsen, “Action research,” Communications of the ACM , vol. 42, no. 1, pp. 94–97, 1999

  2. [10]

    Action research can swing the balance in experimental software engineering,

    P. S. M. Dos Santos and G. H. Travassos, “Action research can swing the balance in experimental software engineering,” in Advances in computers. Elsevier, 2011, vol. 83, pp. 205–276

  3. [11]

    Ezzy, Qualitative analysis

    D. Ezzy, Qualitative analysis. Routledge, 2013

  4. [12]

    Carrying out qualitative analysis,

    J. Ritchie, L. Spencer, W. O’Connor et al. , “Carrying out qualitative analysis,” Qualitative research practice: A guide for social science students and researchers, vol. 2003, pp. 219–62, 2003

  5. [13]

    Qualitative software engineering research: Reflections and guidelines,

    P. Lenberg, R. Feldt, L. Gren, L. G. Wallgren Tengberg, I. Tidefors, and D. Graziotin, “Qualitative software engineering research: Reflections and guidelines,” Journal of Software: Evolution and Process , vol. 36, no. 6, p. e2607, 2024

  6. [14]

    Qualitative methods in empirical studies of software engineering,

    C. B. Seaman, “Qualitative methods in empirical studies of software engineering,” IEEE Transactions on software engineering, vol. 25, no. 4, pp. 557–572, 1999

  7. [15]

    Qualitative analysis,

    G. R. Gibbs, “Qualitative analysis,” Qualitative Data Analysis , p. 277, 2014

  8. [16]

    Using qualitative analysis software to facilitate qualita- tive data analysis,

    V . Talanquer, “Using qualitative analysis software to facilitate qualita- tive data analysis,” in Tools of chemistry education research . ACS Publications, 2014, pp. 83–95

  9. [17]

    Computer software and qualitative analysis: A reassess- ment,

    R. TESCH, “Computer software and qualitative analysis: A reassess- ment,” New Technology in Sociology: Practical Applications in Research and Work, 2019

  10. [18]

    Learn for yourself: The self-learning tools for qualitative analysis software packages,

    F. Freitas, J. Ribeiro, C. Brand ˜ao, L. P. Reis, F. N. de Souza, and A. P. Costa, “Learn for yourself: The self-learning tools for qualitative analysis software packages,” Digital Education Review , no. 32, pp. 97– 117, 2017

  11. [19]

    Tools for analyzing qualitative data: The history and relevance of qualitative data analysis software,

    L. S. Gilbert, K. Jackson, and S. Di Gregorio, “Tools for analyzing qualitative data: The history and relevance of qualitative data analysis software,” Handbook of research on educational communications and technology, pp. 221–236, 2014

  12. [20]

    Artificial intelligence and qualitative research: The promise and perils of large language model (llm)‘assistance’,

    J. Roberts, M. Baker, and J. Andrew, “Artificial intelligence and qualitative research: The promise and perils of large language model (llm)‘assistance’,” Critical Perspectives on Accounting , vol. 99, p. 102722, 2024

  13. [21]

    An examination of the use of large language models to aid analysis of textual data,

    R. H. Tai, L. R. Bentley, X. Xia, J. M. Sitt, S. C. Fankhauser, A. M. Chicas-Mosier, and B. G. Monteith, “An examination of the use of large language models to aid analysis of textual data,” International Journal of Qualitative Methods , vol. 23, p. 16094069241231168, 2024

  14. [22]

    Ai and human reasoning: Qualitative research in the age of large language models,

    M. Bano, D. Zowghi, and J. Whittle, “Ai and human reasoning: Qualitative research in the age of large language models,” The AI Ethics Journal, vol. 3, no. 1, 2023

  15. [23]

    Large language models for qualitative research in software engineering: exploring opportunities and challenges,

    M. Bano, R. Hoda, D. Zowghi, and C. Treude, “Large language models for qualitative research in software engineering: exploring opportunities and challenges,” Automated Software Engineering , vol. 31, no. 1, p. 8, 2024

  16. [24]

    Ai and the transformation of social science research,

    I. Grossmann, M. Feinberg, D. C. Parker, N. A. Christakis, P. E. Tetlock, and W. A. Cunningham, “Ai and the transformation of social science research,” Science, vol. 380, no. 6650, pp. 1108–1109, 2023

  17. [25]

    Chatgpt and other large language models are double-edged swords,

    Y . Shen, L. Heacock, J. Elias, K. D. Hentel, B. Reig, G. Shih, and L. Moy, “Chatgpt and other large language models are double-edged swords,” p. e230163, 2023

  18. [26]

    A survey on evaluation of large language models,

    Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024

  19. [27]

    Large language models in qualitative research: Can we do the data justice?

    H. Schroeder, M. A. L. Qu ´er´e, C. Randazzo, D. Mimno, and S. Schoenebeck, “Large language models in qualitative research: Can we do the data justice?” arXiv preprint arXiv:2410.07362 , 2024

  20. [28]

    Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis,

    S. A. Gebreegziabher, Z. Zhang, X. Tang, Y . Meng, E. L. Glassman, and T. J.-J. Li, “Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–19

  21. [29]

    Collabcoder: a lower-barrier, rigorous workflow for inductive collabo- rative qualitative analysis with large language models,

    J. Gao, Y . Guo, G. Lim, T. Zhang, Z. Zhang, T. J.-J. Li, and S. T. Perrault, “Collabcoder: a lower-barrier, rigorous workflow for inductive collabo- rative qualitative analysis with large language models,” in Proceedings of the CHI Conference on Human Factors in Computing Sys...

  22. [30]

    Dolma: An open corpus of three trillion tokens for language model pretraining research,

    L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y . Elazar et al. , “Dolma: An open corpus of three trillion tokens for language model pretraining research,” arXiv preprint arXiv:2402.00159 , 2024

  23. [31]

    Fighting reviewer fatigue or amplifying bias? considerations and recommendations for use of chatgpt and other large language models in scholarly peer review,

    M. Hosseini and S. P. Horbach, “Fighting reviewer fatigue or amplifying bias? considerations and recommendations for use of chatgpt and other large language models in scholarly peer review,” Research integrity and peer review, vol. 8, no. 1, p. 4, 2023

  24. [32]

    The dangers of using proprietary llms for research,

    ´E. Ollion, R. Shen, A. Macanovic, and A. Chatelain, “The dangers of using proprietary llms for research,” Nature Machine Intelligence, vol. 6, no. 1, pp. 4–5, 2024

  25. [33]

    Evidence-based soft- ware engineering,

    B. A. Kitchenham, T. Dyba, and M. Jorgensen, “Evidence-based soft- ware engineering,” in Proceedings. 26th International Conference on Software Engineering. IEEE, 2004, pp. 273–281

  26. [34]

    Recommended steps for thematic synthesis in software engineering,

    D. S. Cruzes and T. Dyba, “Recommended steps for thematic synthesis in software engineering,” in 2011 international symposium on empirical software engineering and measurement . IEEE, 2011, pp. 275–284

  27. [35]

    Prompt engineering in medical education,

    T. Heston and C. Khun, “Prompt engineering in medical education,” International Medical Education , vol. 2, pp. 198–205, 8 2023

  28. [36]

    A prompt pattern catalog to enhance prompt engineering with chatgpt,

    J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. El- nashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” arXiv preprint arXiv:2302.11382, 2023

  29. [37]

    Guidelines for conducting and reporting case study research in software engineering,

    P. Runeson and M. H ¨ost, “Guidelines for conducting and reporting case study research in software engineering,” Empirical software engineering, vol. 14, pp. 131–164, 2009

  30. [38]

    Threats to validity in software engineering research: A critical reflection,

    R. Verdecchia, E. Engstr ¨om, P. Lago, P. Runeson, and Q. Song, “Threats to validity in software engineering research: A critical reflection,” Information and Software Technology , vol. 164, p. 107329, 2023

  31. [39]

    Assessing the impact of prompting, persona, and chain of thought methods on chatgpt’s arithmetic capabilities,

    Y . Chen, C. Wong, H. Yang, J. Aguenza, S. Bhujangari, B. Vu, X. Lei, A. Prasad, M. Fluss, E. Phuong et al. , “Assessing the impact of prompting, persona, and chain of thought methods on chatgpt’s arithmetic capabilities,” arXiv preprint arXiv:2312.15006 , 2023. 8

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.