Pith. sign in

REVIEW 2 major objections 3 minor 22 references

Large Language Model for Qualitative Research -- A Systematic Mapping Study

T0 review · 2 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Systematic map of eight studies finds LLMs largely match human qualitative coding.

desk verdict A small, honest mapping study whose DE11 prompt-detail filter narrows the scope more than the 'diverse fields' claim admits; worth refereeing with revisions. read the letter →

arxiv 2411.14473 v4 pith:OUCWPNCC submitted 2024-11-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords largelanguagemodelsqualitativeanalysissystematicmappingstudythematiccodingpromptengineeringLLMevaluationhuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish where and how large language models (LLMs) have actually been used to perform qualitative analysis, and whether the results hold up against traditional human coding. Reviewing 354 candidate papers and admitting eight that both apply an LLM to qualitative data and report the prompt engineering used, it finds that most studies judged the LLM's output equivalent to human analysis in education, healthcare, culture, and technology contexts. It also finds that no included study came from software engineering and that no standard evaluation metric exists across the field. The practical upshot is that LLM-assisted coding is plausible for open coding and theme extraction but still depends on carefully engineered prompts and human oversight.

What carries the argument

The carrying mechanism is the systematic-mapping protocol itself: a search string run across six databases, explicit inclusion and exclusion criteria, and a sixteen-item data-extraction questionnaire (DE1–DE16) mapped to five research questions. The load-bearing criterion is DE11, which required every included study to describe its prompt engineering; this cut the eligible set from 21 to 8 studies and makes the resulting synthesis a map of reproducible LLM-assisted analysis rather than of all LLM-for-qualitative work.

What would settle it

A replication that relaxes the prompt-detail requirement (DE11) and includes the 13 excluded studies would refute the equivalence claim if those studies consistently report LLM outputs worse than human coding or high hallucination rates.

Watch

Extended reading notes

Core claim

The paper's central claim is that the eight qualifying primary studies show LLMs are already being applied to qualitative analysis across diverse domains, that in the majority of comparisons they perform equivalently to traditional human analysis, and that the main barriers are dependence on well-structured prompts, occasional hallucinations, and limited contextual sensitivity. The paper further claims that this equivalence is not yet backed by a standardized evaluation metric, and that the absence of software-engineering applications marks a research gap rather than evidence of failure. As direct corollaries, its findings imply that LLMs can shorten coding from weeks to hours and are best used as aids, with human analysts retaining final interpretive authority.

Load-bearing premise

The conclusions rest on the assumption that the eight studies which happened to document their prompt engineering fairly represent all LLM qualitative-analysis research; if studies without prompt details carry different evidence about effectiveness, the equivalence claim would not generalize.

Editorial extensions

If this is right

  • Researchers can treat LLM-assisted open coding and theme extraction as a time-saving step that produces results broadly comparable to human coding, while keeping final interpretation with humans.
  • The field lacks a standardized evaluation metric, so until one is adopted, comparisons of LLM versus human analysis will remain difficult to aggregate across studies.
  • The absence of software-engineering primary studies is a concrete opening: requirements engineering and user-feedback analysis are named as untested applications.
  • Prompt engineering is not a peripheral detail but a core part of the method; studies that omit prompt details are currently excluded from evidence syntheses.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if prompt-detail reporting becomes standard, the map is likely to grow quickly and the equivalence result may shift, since the DE11 filter probably selects for more careful and complete studies.
  • Editorial inference: a head-to-head benchmark where the same interview corpus is coded by several LLMs and several human teams under a fixed metric such as Cohen's kappa would directly test whether the apparent equivalence is real or an artifact of heterogeneous evaluations.
  • Editorial inference: the reported 'weeks to hours' speed gain, if replicated, changes the cost structure of qualitative research by making much larger corpora feasible, but the risk of unnoticed hallucinated codes grows with scale.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. This paper reports a systematic mapping study (SMS) of the use of large language models (LLMs) in qualitative research, following the Kitchenham and Charters guidelines. The authors searched five databases and arXiv, applying inclusion/exclusion criteria and a data extraction questionnaire (DE1–DE16) that was operationalized through ChatGPT with manual verification. Eight primary studies, mostly from 2023–2024, were ultimately analyzed across five research questions covering application contexts, models and configurations, analysis techniques, evaluation metrics and outcomes, and limitations. The main reported findings are that LLMs have been applied in healthcare, education, cultural studies, and technology; that effectiveness is mostly equivalent to traditional methods in the included studies; and that key limitations include reliance on prompt engineering, hallucinations, and contextual insensitivity. The paper also identifies a lack of studies in software engineering as a notable gap.

Significance. If the map is accepted as representative, this would be a useful early synthesis of an emerging and fast-moving area, and the authors deserve credit for following a structured protocol, publishing their prompt versions on Zenodo, and transparently reporting some threats to validity. However, the significance is constrained by the small, narrowly scoped corpus: the DE11 criterion (articles must detail prompt engineering) reduces the set from 21 to 8 studies, and the authors themselves acknowledge that this may have restricted the inclusion of early-stage work. Consequently, the paper is best read as a map of studies that explicitly report prompt details, not as a comprehensive map of LLM-for-qualitative-research literature. The claim that no software engineering research exists in this area, and the broader diversity finding, need to be tested against the excluded studies before they can be treated as evidence about the field rather than about the filter.

major comments (2)
  1. [Section III.E and Section V] The DE11 filter, which requires that each included study detail its prompt engineering, is load-bearing for the map's representativeness but its effect is not analyzed. The corpus is cut from 21 to 8 studies, and Section V acknowledges that this criterion may have restricted inclusion, yet the paper nevertheless draws conclusions about the state of the art, including the absence of software engineering studies (Section III.F). To support these conclusions, the authors should report how many of the 21 full-text studies were excluded specifically because of DE11, and compare the 8 included studies with the excluded ones on at least application domain, LLM model, and reported effectiveness. Without such an analysis, the 'diverse fields' and 'no software engineering' claims could be artifacts of the prompt-reporting filter rather than properties of the literature.
  2. [Section III.D and Section V] The paper uses ChatGPT for data extraction and then synthesizes these extractions to evaluate the effectiveness of LLMs in qualitative analysis, which introduces a self-referential risk. The authors mention manual review for doubtful cases and a validation with test articles, but they do not provide any quantitative reliability measure, such as agreement between ChatGPT extractions and human extractions on a sample of studies. Given that the extracted data underpin all five research-question results, a reliability statistic (e.g., percentage agreement or Cohen's kappa on a subset) would substantiate the data validity claim and make the extraction process auditable.
minor comments (3)
  1. [Section III.F, RQ4] The text states that 'All seven included studies compared LLM-assisted qualitative analysis with traditional methods,' but Table III lists eight studies and the same paragraph goes on to describe [20]'s comparison with Atlas.ti Web. This is a numerical inconsistency that should be corrected to 'eight.'
  2. [Section III.E] The study selection narrative moves from 21 studies to 8 studies without a breakdown by exclusion criterion. A PRISMA-style flow diagram showing how many studies were excluded at each step (duplicates, title/abstract, availability, DE1, DE11, etc.) would improve transparency and make the DE11 effect visible to readers.
  3. [Section II.C and elsewhere] There are several typographical and phrasing issues, such as 'maintainiong' in Section II.C and 'suggestes' in Section III.F, which should be corrected in a language pass.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the mapping's conclusions are descriptive syntheses of external primary studies, and the acknowledged use of ChatGPT for extraction does not define the results.

full rationale

This paper is a systematic mapping study rather than a derivation of predictions or first-principles results. Its central claims—LLMs are applied across diverse fields, effectiveness is mostly comparable to traditional methods, and prompt dependence and inaccuracies remain challenges—are descriptive syntheses of eight external primary studies (S1–S8) collected through a defined search and selection protocol. The data-extraction criteria (DE1–DE16) form an instrument, not a fitted model; no conclusion is defined in terms of the extraction outputs. The paper acknowledges in Section V that using LLMs both for data extraction and as the object of study could create confirmation bias, and it describes manual verification of extracted data. That is a validity threat, not a circular derivation, because the evidence base consists of the included studies' own reported findings, not the authors' ChatGPT outputs. The DE11 inclusion criterion does constructively ensure that every selected study reports prompt engineering details, but the statement that all included studies detail their prompts is presented as a screening consequence, and the headline findings about field diversity, limitations, and evaluation metrics are not forced by that filter. The only author self-citation (Wohlin et al. [21], which includes co-author Kalinowski) supports a limitation about not using snowballing and is not load-bearing for any central claim. No step reduces by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No numeric free parameters are used. The mapping rests on the selected corpus, the reliability of LLM-assisted extraction, and the DE11 requirement; these are domain assumptions rather than mathematical axioms.

assumptions (3)
  • domain assumption The eight included studies are representative of the state of the art in LLM-based qualitative analysis.
    All RQ1-RQ5 patterns are computed from this corpus; if the DE11 filter or the exclusion of studies without full-text access biases the corpus, the map's conclusions change. Section III.E.
  • domain assumption ChatGPT-based data extraction, followed by manual verification, yields accurate answers to the DE1-DE16 questionnaire.
    The results section depends entirely on these extracted values; validation is described qualitatively with 'test articles with known responses' and no quantitative reliability metric is reported. Section III.D and Section V.
  • ad hoc to paper Requiring studies to detail prompt engineering (DE11) does not systematically distort the map of LLM effectiveness.
    DE11 is a nonstandard inclusion criterion chosen by the authors to ensure reproducibility; it reduced the corpus from 21 to 8 studies and is acknowledged as potentially restrictive. Section III.D and Section V.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Model for Qualitative Research -- A Systematic Mapping Study." pith.science (2026). https://pith.science/paper/OUCWPNCC

@misc{pith2026241114473,
  author       = {Pith},
  title        = {Pith review of: Large Language Model for Qualitative Research -- A Systematic Mapping Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OUCWPNCC}},
  note         = {Machine review of arXiv:2411.14473}
}
read the original abstract

The exponential growth of text-based data in domains such as healthcare, education, and social sciences has outpaced the capacity of traditional qualitative analysis methods, which are time-intensive and prone to subjectivity. Large Language Models (LLMs), powered by advanced generative AI, have emerged as transformative tools capable of automating and enhancing qualitative analysis. This study systematically maps the literature on the use of LLMs for qualitative research, exploring their application contexts, configurations, methodologies, and evaluation metrics. Findings reveal that LLMs are utilized across diverse fields, demonstrating the potential to automate processes traditionally requiring extensive human input. However, challenges such as reliance on prompt engineering, occasional inaccuracies, and contextual limitations remain significant barriers. This research highlights opportunities for integrating LLMs with human expertise, improving model robustness, and refining evaluation methodologies. By synthesizing trends and identifying research gaps, this study aims to guide future innovations in the application of LLMs for qualitative analysis.

Figures

Figures reproduced from arXiv: 2411.14473 by the authors.

Figure 1
Figure 1. illustrates the steps followed in the study selection process, starting from the execution of the search string across the databases. A total of 354 studies were retrieved, distributed as follows: 20 from the ACM Digital Library, 30 from IEEExplore, 78 from Web of Science, 32 from SBC Open Lib, 193 from Scopus, and 1 from Arxiv [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 14 canonical work pages

  1. [1]

    S. J. Russell and P. Norvig, Artificial intelligence: a modern approach . Pearson, 2016

  2. [2]

    ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope,

    P. P. Ray, “ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope,” Internet of Things and Cyber-Physical Systems , vol. 3, pp. 121-154, 2023

  3. [3]

    San Francisco: OpenAI, 2024

    OPENAI, ChatGPT (vers ˜ao GPT-4). San Francisco: OpenAI, 2024. [Online]. Available: https://www.openai.com. Accessed: Sep. 24, 2024

  4. [4]

    A Survey of Large Language Models,

    W. X. Zhao, “A Survey of Large Language Models,” arXiv preprint arXiv:2303.18223. Available: https://arxiv.org/abs/2303.18223

  5. [5]

    A comprehensive overview of large language models,

    H. Naveed, et al., “A comprehensive overview of large language models,” arXiv preprint arXiv:2307.06435, 2023

  6. [6]

    Corbin and A

    J. Corbin and A. Strauss, Basics of qualitative research: Techniques and procedures for developing grounded theory . Sage Publications, 2014

  7. [7]

    Charmaz, Constructing grounded theory: A practical guide through qualitative analysis

    K. Charmaz, Constructing grounded theory: A practical guide through qualitative analysis. Sage, 2006

  8. [8]

    Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models,

    Y . Zenimoto, R. Hasegawa, T. Utsuro, M. Yoshioka, and N. Kando, “Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 4: Student Research Workshop), pp. 242-254, 2024

Show all 22 references
  1. [9]

    Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,

    W. S. Mathis, S. Zhao, N. Pratt, J. Weleff, and S. De Paoli, “Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,” Computer Methods and Programs in Biomedicine , vol. 255, p. 108...

  2. [10]

    Large language models for qualitative research in software engineering: exploring opportunities and challenges,

    M. Bano, R. Hoda, D. Zowghi, and C. Treude, “Large language models for qualitative research in software engineering: exploring opportunities and challenges,” Automated Software Engineering , vol. 31, no. 1, p. 8, 2024, Springer

  3. [11]

    Screening articles for systematic reviews with ChatGPT,

    E. Syriani, I. David, and G. Kumar, “Screening articles for systematic reviews with ChatGPT,” Journal of Computer Languages , vol. 101287, 2024

  4. [12]

    Extracting accurate materials data from research papers with conversational language models and prompt en- gineering,

    M. P. Polak and D. Morgan, “Extracting accurate materials data from research papers with conversational language models and prompt en- gineering,” Nature Communications , vol. 15, p. 1569, 2024. [Online]. Available: https://doi.org/10.1038/s41467-024-45914-8

  5. [13]

    Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provocation on the limits of the approach,

    S. De Paoli, “Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provocation on the limits of the approach,” Social Science Computer Review, vol. 42, no. 4, pp. 997-1019, 2024

  6. [14]

    Artificial Intelligence and content analysis: the large language models (LLMs) and the automatized cate- gorization,

    A. C. Carius and A. J. Teixeira, “Artificial Intelligence and content analysis: the large language models (LLMs) and the automatized cate- gorization,” AI & Society , pp. 1-12, 2024

  7. [15]

    Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering,

    M. de M. Lec ¸a, L. Valenc ¸a, R. Santos, and R. de S. Santos, “Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering,” arXiv preprint, 2024. Available: https://arxiv.org/abs/2412.06564

  8. [16]

    Guidelines for Performing Systematic Literature Reviews in Software Engineering,

    B. Kitchenham and S. Charters, “Guidelines for Performing Systematic Literature Reviews in Software Engineering,” Technical Report EBSE 2007-001, Keele University Keele, UK; Durham University: Durham, UK, 2007

  9. [17]

    Deep Learning Models for Analyzing Social Construction of Knowledge Online,

    C. N. Gunawardena, Y . Chen, N. Flor, and D. S ´anchez, “Deep Learning Models for Analyzing Social Construction of Knowledge Online,” Online Learning, vol. 27, no. 4, pp. 69-92, 2023

  10. [18]

    Exploring Qualitative Research Using LLMs,

    M. Bano, D. Zowghi, and J. Whittle, “Exploring Qualitative Research Using LLMs,” arXiv preprint arXiv:2306.13298, 2023

  11. [19]

    LLMusic: Topic Modeling in Song Lyrics Combining LLM, Prompt Engineering, and BERTopic

    J. D. Y . Rojas and K. Becker, “LLMusic: Topic Modeling in Song Lyrics Combining LLM, Prompt Engineering, and BERTopic” (LLMusic: Mod- elagem de t ´opicos em letras de m ´usicas combinando LLM, Engenharia de Prompt e BERTopic), in Workshop de Teses e Dissertac ¸˜oes (WTDBD) - ...

  12. [20]

    CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models,

    J. Gao, Y . Guo, G. Lim, T. Zhang, Z. Zhang, T. J.-J. Li, and S. T. Per- rault, “CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models,” arXiv preprint, 2024. Available: https://arxiv.org/abs/2304.07366

  13. [21]

    Suc- cessful combination of database search and snowballing for iden- tification of primary studies in systematic literature studies,

    C. Wohlin, M. Kalinowski, K. R. Felizardo, and E. Mendes, “Suc- cessful combination of database search and snowballing for iden- tification of primary studies in systematic literature studies,” Infor- mation and Software Technology , vol. 147, p. 106908, 2022, doi: https://doi...

  14. [164]

    Available: https://doi.org/10.5753/sbbd estendido.2024

    [Online]. Available: https://doi.org/10.5753/sbbd estendido.2024. 243767

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.