REVIEW 2 major objections 3 minor 22 references
Large Language Model for Qualitative Research -- A Systematic Mapping Study
T0 review · 2 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Systematic map of eight studies finds LLMs largely match human qualitative coding.
desk verdict A small, honest mapping study whose DE11 prompt-detail filter narrows the scope more than the 'diverse fields' claim admits; worth refereeing with revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the systematic-mapping protocol itself: a search string run across six databases, explicit inclusion and exclusion criteria, and a sixteen-item data-extraction questionnaire (DE1–DE16) mapped to five research questions. The load-bearing criterion is DE11, which required every included study to describe its prompt engineering; this cut the eligible set from 21 to 8 studies and makes the resulting synthesis a map of reproducible LLM-assisted analysis rather than of all LLM-for-qualitative work.
What would settle it
A replication that relaxes the prompt-detail requirement (DE11) and includes the 13 excluded studies would refute the equivalence claim if those studies consistently report LLM outputs worse than human coding or high hallucination rates.
Extended reading notes
Core claim
The paper's central claim is that the eight qualifying primary studies show LLMs are already being applied to qualitative analysis across diverse domains, that in the majority of comparisons they perform equivalently to traditional human analysis, and that the main barriers are dependence on well-structured prompts, occasional hallucinations, and limited contextual sensitivity. The paper further claims that this equivalence is not yet backed by a standardized evaluation metric, and that the absence of software-engineering applications marks a research gap rather than evidence of failure. As direct corollaries, its findings imply that LLMs can shorten coding from weeks to hours and are best used as aids, with human analysts retaining final interpretive authority.
Load-bearing premise
The conclusions rest on the assumption that the eight studies which happened to document their prompt engineering fairly represent all LLM qualitative-analysis research; if studies without prompt details carry different evidence about effectiveness, the equivalence claim would not generalize.
Editorial extensions
If this is right
- Researchers can treat LLM-assisted open coding and theme extraction as a time-saving step that produces results broadly comparable to human coding, while keeping final interpretation with humans.
- The field lacks a standardized evaluation metric, so until one is adopted, comparisons of LLM versus human analysis will remain difficult to aggregate across studies.
- The absence of software-engineering primary studies is a concrete opening: requirements engineering and user-feedback analysis are named as untested applications.
- Prompt engineering is not a peripheral detail but a core part of the method; studies that omit prompt details are currently excluded from evidence syntheses.
Reading between the lines
- Editorial inference: if prompt-detail reporting becomes standard, the map is likely to grow quickly and the equivalence result may shift, since the DE11 filter probably selects for more careful and complete studies.
- Editorial inference: a head-to-head benchmark where the same interview corpus is coded by several LLMs and several human teams under a fixed metric such as Cohen's kappa would directly test whether the apparent equivalence is real or an artifact of heterogeneous evaluations.
- Editorial inference: the reported 'weeks to hours' speed gain, if replicated, changes the cost structure of qualitative research by making much larger corpora feasible, but the risk of unnoticed hallucinated codes grows with scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a systematic mapping study (SMS) of the use of large language models (LLMs) in qualitative research, following the Kitchenham and Charters guidelines. The authors searched five databases and arXiv, applying inclusion/exclusion criteria and a data extraction questionnaire (DE1–DE16) that was operationalized through ChatGPT with manual verification. Eight primary studies, mostly from 2023–2024, were ultimately analyzed across five research questions covering application contexts, models and configurations, analysis techniques, evaluation metrics and outcomes, and limitations. The main reported findings are that LLMs have been applied in healthcare, education, cultural studies, and technology; that effectiveness is mostly equivalent to traditional methods in the included studies; and that key limitations include reliance on prompt engineering, hallucinations, and contextual insensitivity. The paper also identifies a lack of studies in software engineering as a notable gap.
Significance. If the map is accepted as representative, this would be a useful early synthesis of an emerging and fast-moving area, and the authors deserve credit for following a structured protocol, publishing their prompt versions on Zenodo, and transparently reporting some threats to validity. However, the significance is constrained by the small, narrowly scoped corpus: the DE11 criterion (articles must detail prompt engineering) reduces the set from 21 to 8 studies, and the authors themselves acknowledge that this may have restricted the inclusion of early-stage work. Consequently, the paper is best read as a map of studies that explicitly report prompt details, not as a comprehensive map of LLM-for-qualitative-research literature. The claim that no software engineering research exists in this area, and the broader diversity finding, need to be tested against the excluded studies before they can be treated as evidence about the field rather than about the filter.
major comments (2)
- [Section III.E and Section V] The DE11 filter, which requires that each included study detail its prompt engineering, is load-bearing for the map's representativeness but its effect is not analyzed. The corpus is cut from 21 to 8 studies, and Section V acknowledges that this criterion may have restricted inclusion, yet the paper nevertheless draws conclusions about the state of the art, including the absence of software engineering studies (Section III.F). To support these conclusions, the authors should report how many of the 21 full-text studies were excluded specifically because of DE11, and compare the 8 included studies with the excluded ones on at least application domain, LLM model, and reported effectiveness. Without such an analysis, the 'diverse fields' and 'no software engineering' claims could be artifacts of the prompt-reporting filter rather than properties of the literature.
- [Section III.D and Section V] The paper uses ChatGPT for data extraction and then synthesizes these extractions to evaluate the effectiveness of LLMs in qualitative analysis, which introduces a self-referential risk. The authors mention manual review for doubtful cases and a validation with test articles, but they do not provide any quantitative reliability measure, such as agreement between ChatGPT extractions and human extractions on a sample of studies. Given that the extracted data underpin all five research-question results, a reliability statistic (e.g., percentage agreement or Cohen's kappa on a subset) would substantiate the data validity claim and make the extraction process auditable.
minor comments (3)
- [Section III.F, RQ4] The text states that 'All seven included studies compared LLM-assisted qualitative analysis with traditional methods,' but Table III lists eight studies and the same paragraph goes on to describe [20]'s comparison with Atlas.ti Web. This is a numerical inconsistency that should be corrected to 'eight.'
- [Section III.E] The study selection narrative moves from 21 studies to 8 studies without a breakdown by exclusion criterion. A PRISMA-style flow diagram showing how many studies were excluded at each step (duplicates, title/abstract, availability, DE1, DE11, etc.) would improve transparency and make the DE11 effect visible to readers.
- [Section II.C and elsewhere] There are several typographical and phrasing issues, such as 'maintainiong' in Section II.C and 'suggestes' in Section III.F, which should be corrected in a language pass.
Circularity Check
No significant circularity: the mapping's conclusions are descriptive syntheses of external primary studies, and the acknowledged use of ChatGPT for extraction does not define the results.
full rationale
This paper is a systematic mapping study rather than a derivation of predictions or first-principles results. Its central claims—LLMs are applied across diverse fields, effectiveness is mostly comparable to traditional methods, and prompt dependence and inaccuracies remain challenges—are descriptive syntheses of eight external primary studies (S1–S8) collected through a defined search and selection protocol. The data-extraction criteria (DE1–DE16) form an instrument, not a fitted model; no conclusion is defined in terms of the extraction outputs. The paper acknowledges in Section V that using LLMs both for data extraction and as the object of study could create confirmation bias, and it describes manual verification of extracted data. That is a validity threat, not a circular derivation, because the evidence base consists of the included studies' own reported findings, not the authors' ChatGPT outputs. The DE11 inclusion criterion does constructively ensure that every selected study reports prompt engineering details, but the statement that all included studies detail their prompts is presented as a screening consequence, and the headline findings about field diversity, limitations, and evaluation metrics are not forced by that filter. The only author self-citation (Wohlin et al. [21], which includes co-author Kalinowski) supports a limitation about not using snowballing and is not load-bearing for any central claim. No step reduces by construction to its own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption The eight included studies are representative of the state of the art in LLM-based qualitative analysis.
- domain assumption ChatGPT-based data extraction, followed by manual verification, yields accurate answers to the DE1-DE16 questionnaire.
- ad hoc to paper Requiring studies to detail prompt engineering (DE11) does not systematically distort the map of LLM effectiveness.
Cite this review
Pith. "Pith review of Large Language Model for Qualitative Research -- A Systematic Mapping Study." pith.science (2026). https://pith.science/paper/OUCWPNCC
@misc{pith2026241114473,
author = {Pith},
title = {Pith review of: Large Language Model for Qualitative Research -- A Systematic Mapping Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/OUCWPNCC}},
note = {Machine review of arXiv:2411.14473}
}
read the original abstract
The exponential growth of text-based data in domains such as healthcare, education, and social sciences has outpaced the capacity of traditional qualitative analysis methods, which are time-intensive and prone to subjectivity. Large Language Models (LLMs), powered by advanced generative AI, have emerged as transformative tools capable of automating and enhancing qualitative analysis. This study systematically maps the literature on the use of LLMs for qualitative research, exploring their application contexts, configurations, methodologies, and evaluation metrics. Findings reveal that LLMs are utilized across diverse fields, demonstrating the potential to automate processes traditionally requiring extensive human input. However, challenges such as reliance on prompt engineering, occasional inaccuracies, and contextual limitations remain significant barriers. This research highlights opportunities for integrating LLMs with human expertise, improving model robustness, and refining evaluation methodologies. By synthesizing trends and identifying research gaps, this study aims to guide future innovations in the application of LLMs for qualitative analysis.
Figures
Reference graph
Works this paper leans on
-
[1]
S. J. Russell and P. Norvig, Artificial intelligence: a modern approach . Pearson, 2016
work page 2016
-
[2]
P. P. Ray, “ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope,” Internet of Things and Cyber-Physical Systems , vol. 3, pp. 121-154, 2023
work page 2023
-
[3]
OPENAI, ChatGPT (vers ˜ao GPT-4). San Francisco: OpenAI, 2024. [Online]. Available: https://www.openai.com. Accessed: Sep. 24, 2024
work page 2024
-
[4]
A Survey of Large Language Models,
W. X. Zhao, “A Survey of Large Language Models,” arXiv preprint arXiv:2303.18223. Available: https://arxiv.org/abs/2303.18223
-
[5]
A comprehensive overview of large language models,
H. Naveed, et al., “A comprehensive overview of large language models,” arXiv preprint arXiv:2307.06435, 2023
arXiv 2023
-
[6]
J. Corbin and A. Strauss, Basics of qualitative research: Techniques and procedures for developing grounded theory . Sage Publications, 2014
work page 2014
-
[7]
Charmaz, Constructing grounded theory: A practical guide through qualitative analysis
K. Charmaz, Constructing grounded theory: A practical guide through qualitative analysis. Sage, 2006
2006
-
[8]
Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models,
Y . Zenimoto, R. Hasegawa, T. Utsuro, M. Yoshioka, and N. Kando, “Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 4: Student Research Workshop), pp. 242-254, 2024
work page 2024
Show all 22 references
-
[9]
Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,
W. S. Mathis, S. Zhao, N. Pratt, J. Weleff, and S. De Paoli, “Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,” Computer Methods and Programs in Biomedicine , vol. 255, p. 108...
2024
-
[10]
Large language models for qualitative research in software engineering: exploring opportunities and challenges,
M. Bano, R. Hoda, D. Zowghi, and C. Treude, “Large language models for qualitative research in software engineering: exploring opportunities and challenges,” Automated Software Engineering , vol. 31, no. 1, p. 8, 2024, Springer
2024
-
[11]
Screening articles for systematic reviews with ChatGPT,
E. Syriani, I. David, and G. Kumar, “Screening articles for systematic reviews with ChatGPT,” Journal of Computer Languages , vol. 101287, 2024
2024
-
[12]
Extracting accurate materials data from research papers with conversational language models and prompt en- gineering,
M. P. Polak and D. Morgan, “Extracting accurate materials data from research papers with conversational language models and prompt en- gineering,” Nature Communications , vol. 15, p. 1569, 2024. [Online]. Available: https://doi.org/10.1038/s41467-024-45914-8
2024 doi
-
[13]
Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provocation on the limits of the approach,
S. De Paoli, “Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provocation on the limits of the approach,” Social Science Computer Review, vol. 42, no. 4, pp. 997-1019, 2024
2024
-
[14]
Artificial Intelligence and content analysis: the large language models (LLMs) and the automatized cate- gorization,
A. C. Carius and A. J. Teixeira, “Artificial Intelligence and content analysis: the large language models (LLMs) and the automatized cate- gorization,” AI & Society , pp. 1-12, 2024
2024
-
[15]
Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering,
M. de M. Lec ¸a, L. Valenc ¸a, R. Santos, and R. de S. Santos, “Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering,” arXiv preprint, 2024. Available: https://arxiv.org/abs/2412.06564
2024 arXiv
-
[16]
Guidelines for Performing Systematic Literature Reviews in Software Engineering,
B. Kitchenham and S. Charters, “Guidelines for Performing Systematic Literature Reviews in Software Engineering,” Technical Report EBSE 2007-001, Keele University Keele, UK; Durham University: Durham, UK, 2007
2007
-
[17]
Deep Learning Models for Analyzing Social Construction of Knowledge Online,
C. N. Gunawardena, Y . Chen, N. Flor, and D. S ´anchez, “Deep Learning Models for Analyzing Social Construction of Knowledge Online,” Online Learning, vol. 27, no. 4, pp. 69-92, 2023
2023
-
[18]
Exploring Qualitative Research Using LLMs,
M. Bano, D. Zowghi, and J. Whittle, “Exploring Qualitative Research Using LLMs,” arXiv preprint arXiv:2306.13298, 2023
2023 arXiv
-
[19]
LLMusic: Topic Modeling in Song Lyrics Combining LLM, Prompt Engineering, and BERTopic
J. D. Y . Rojas and K. Becker, “LLMusic: Topic Modeling in Song Lyrics Combining LLM, Prompt Engineering, and BERTopic” (LLMusic: Mod- elagem de t ´opicos em letras de m ´usicas combinando LLM, Engenharia de Prompt e BERTopic), in Workshop de Teses e Dissertac ¸˜oes (WTDBD) - ...
2024
-
[20]
CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models,
J. Gao, Y . Guo, G. Lim, T. Zhang, Z. Zhang, T. J.-J. Li, and S. T. Per- rault, “CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models,” arXiv preprint, 2024. Available: https://arxiv.org/abs/2304.07366
2024 arXiv
-
[21]
Suc- cessful combination of database search and snowballing for iden- tification of primary studies in systematic literature studies,
C. Wohlin, M. Kalinowski, K. R. Felizardo, and E. Mendes, “Suc- cessful combination of database search and snowballing for iden- tification of primary studies in systematic literature studies,” Infor- mation and Software Technology , vol. 147, p. 106908, 2022, doi: https://doi...
2022
-
[164]
Available: https://doi.org/10.5753/sbbd estendido.2024
[Online]. Available: https://doi.org/10.5753/sbbd estendido.2024. 243767
2024 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.