{"id":"5b9041d9-535b-4868-8ec2-80f3a3c8f4c5","arxiv_id":"2411.14473","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic map of eight studies shows LLM-assisted qualitative analysis is mostly comparable to manual methods, with prompt dependence and hallucination as recurring limitations.","lead":"This paper maps eight studies on how large language models are used in qualitative analysis, covering models, prompts, methods, and evaluation metrics. It is a useful snapshot for researchers deciding whether and how to use AI tools for coding and theme extraction.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DE11 prompt-detail filter likely distorts the systematic map: the missing SE studies and 'diverse fields' conclusion may be artifacts of the 21-to-8 cutoff.","rationale":"The reader's weakest assumption identifies the DE11 filter as the most load-bearing threat. I concur: the filter removes the majority of candidate studies and is explicitly acknowledged as restrictive, yet its correlation with field and publication year is unexamined. The paper's notable finding of 'no works in Software Engineering' is particularly vulnerable, as SE papers may report outcomes without full prompt transcripts. The manual inclusion of a known study (S20) further violates protocol systematicity, and the absence of the extraction dataset precludes verification. Given that the central claim is a descriptive synthesis, selection bias directly undermines its representativeness. The verdict should remain CONDITIONAL: the paper can be accepted if the authors provide the excluded-study analysis or soften the claims about field gaps. No change to the reader's conditional acceptance is needed.","tokens_in":10544,"tokens_out":5389,"duration_ms":57093,"concrete_test":"Retrieve the 13 studies excluded at the DE11 step from the authors' Parsif.al/Zenodo record (or re-run screening without DE11), classify their field, model, and outcome (e.g., using the DE12–DE14 descriptors), and compare distributions and conclusions to the 8-study map. If SE or pre-2024 studies appear with differing effectiveness evidence, the paper's central claims about diversity and the SE gap require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central descriptive claim—that LLMs are used across diverse fields and show potential to automate qualitative analysis—rests on a map of 8 studies. The DE11 inclusion criterion (Section III.E) requires that each study detail its prompt engineering; this cut the candidate set from 21 to 8, a 62% reduction. The authors acknowledge (Section V) that this may have restricted early-stage studies. Because reporting norms vary by field—software engineering papers, for instance, often focus on tool results rather than prompt internals—this filter could systematically exclude exactly the SE studies whose absence is highlighted as a gap (Section III.F). The manual addition of CollabCoder (S20) after the search protocol, and the use of ChatGPT itself for screening/extraction, compound the selection bias. Without the excluded studies' data, the map's claims about field diversity and the SE gap are not representative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic mapping study (SMS) of the use of large language models (LLMs) in qualitative research, following the Kitchenham and Charters guidelines. The authors searched five databases and arXiv, applying inclusion/exclusion criteria and a data extraction questionnaire (DE1–DE16) that was operationalized through ChatGPT with manual verification. Eight primary studies, mostly from 2023–2024, were ultimately analyzed across five research questions covering application contexts, models and configurations, analysis techniques, evaluation metrics and outcomes, and limitations. The main reported findings are that LLMs have been applied in healthcare, education, cultural studies, and technology; that effectiveness is mostly equivalent to traditional methods in the included studies; and that key limitations include reliance on prompt engineering, hallucinations, and contextual insensitivity. The paper also identifies a lack of studies in software engineering as a notable gap.","tokens_in":10705,"tokens_out":4900,"duration_ms":49051,"significance":"If the map is accepted as representative, this would be a useful early synthesis of an emerging and fast-moving area, and the authors deserve credit for following a structured protocol, publishing their prompt versions on Zenodo, and transparently reporting some threats to validity. However, the significance is constrained by the small, narrowly scoped corpus: the DE11 criterion (articles must detail prompt engineering) reduces the set from 21 to 8 studies, and the authors themselves acknowledge that this may have restricted the inclusion of early-stage work. Consequently, the paper is best read as a map of studies that explicitly report prompt details, not as a comprehensive map of LLM-for-qualitative-research literature. The claim that no software engineering research exists in this area, and the broader diversity finding, need to be tested against the excluded studies before they can be treated as evidence about the field rather than about the filter.","major_comments":[{"comment":"The DE11 filter, which requires that each included study detail its prompt engineering, is load-bearing for the map's representativeness but its effect is not analyzed. The corpus is cut from 21 to 8 studies, and Section V acknowledges that this criterion may have restricted inclusion, yet the paper nevertheless draws conclusions about the state of the art, including the absence of software engineering studies (Section III.F). To support these conclusions, the authors should report how many of the 21 full-text studies were excluded specifically because of DE11, and compare the 8 included studies with the excluded ones on at least application domain, LLM model, and reported effectiveness. Without such an analysis, the 'diverse fields' and 'no software engineering' claims could be artifacts of the prompt-reporting filter rather than properties of the literature.","section":"Section III.E and Section V"},{"comment":"The paper uses ChatGPT for data extraction and then synthesizes these extractions to evaluate the effectiveness of LLMs in qualitative analysis, which introduces a self-referential risk. The authors mention manual review for doubtful cases and a validation with test articles, but they do not provide any quantitative reliability measure, such as agreement between ChatGPT extractions and human extractions on a sample of studies. Given that the extracted data underpin all five research-question results, a reliability statistic (e.g., percentage agreement or Cohen's kappa on a subset) would substantiate the data validity claim and make the extraction process auditable.","section":"Section III.D and Section V"}],"minor_comments":[{"comment":"The text states that 'All seven included studies compared LLM-assisted qualitative analysis with traditional methods,' but Table III lists eight studies and the same paragraph goes on to describe [20]'s comparison with Atlas.ti Web. This is a numerical inconsistency that should be corrected to 'eight.'","section":"Section III.F, RQ4"},{"comment":"The study selection narrative moves from 21 studies to 8 studies without a breakdown by exclusion criterion. A PRISMA-style flow diagram showing how many studies were excluded at each step (duplicates, title/abstract, availability, DE1, DE11, etc.) would improve transparency and make the DE11 effect visible to readers.","section":"Section III.E"},{"comment":"There are several typographical and phrasing issues, such as 'maintainiong' in Section II.C and 'suggestes' in Section III.F, which should be corrected in a language pass.","section":"Section II.C and elsewhere"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a software engineering or empirical software engineering venue, and the authors are transparent about the main risks. In my view, the DE11 filter is the most consequential issue: it is an unusual inclusion criterion that directly shapes the map, and the paper needs to either provide an analysis of the excluded studies or explicitly reframe the contribution as a map of studies that report prompt engineering. The self-referential use of ChatGPT for extraction is also a concern, but it is mitigated by manual review and is not by itself grounds for rejection. I would encourage the editor to seek a revised version that addresses the representativeness question rather than rejecting the paper outright, because the underlying topic is timely and the protocol is otherwise sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a competent but small systematic mapping of LLM use in qualitative analysis—eight primary studies, following Kitchenham and Charters, with a distinctive twist: it only includes studies that detail their prompt engineering. That criterion cut the pool from 21 to 8. The paper is honest about what it does and does not show, and it's worth engaging with as a reference point for that subfield.\n\nWhat's genuinely good: the protocol is transparent (prompts and search process on Zenodo), the authors explicitly position their work relative to the complementary Leça et al. mapping, and they are upfront about using ChatGPT for screening and extraction, with manual verification. The descriptive synthesis in the results sections tracks the primary studies accurately—I spot-checked a few—so as a map it is reliable within its narrow scope.\n\nWhere it's soft: the DE11 filter. The stress-test worry that this systematically distorts the map is plausible but not demonstrated. The authors acknowledge in Section V that the criterion may have excluded early-stage studies; the SE gap they highlight in Section III.F could indeed be an artifact of field-specific reporting norms rather than a real absence of work. That should be discussed as a hypothesis, not stated as a finding. Relatedly, the abstract's \"diverse fields\" claim is overreach for eight studies from healthcare, education, culture, and technology, with no SE.\n\nTwo smaller issues. First, RQ4 says \"All seven included studies compared...\" but there are eight included studies; the count is internally inconsistent. Second, the take-away lesson that ChatGPT \"demonstrated superior capabilities\" is not supported by the evidence, which mostly compares LLMs to human analysts, not to each other. That's an overstatement.\n\nThe self-referential use of ChatGPT for extraction is a real methodological tension, but the manual checks mitigate it, and the paper says so. It's not a fatal flaw.\n\nBottom line: this is a useful descriptive snapshot for researchers entering the area, not a landmark. It deserves a serious referee but needs revision: fix the count error, temper the superiority claim, and reframe the SE gap as a possible artifact of the inclusion criteria.\n\nRecommendation: accept for peer review, conditional on those revisions.","headline":"A small, honest mapping study whose DE11 prompt-detail filter narrows the scope more than the 'diverse fields' claim admits; worth refereeing with revisions.","tokens_in":11242,"tokens_out":2258,"would_cite":false,"duration_ms":23079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Systematic map of eight studies finds LLMs largely match human qualitative coding.","keywords":["large language models","qualitative analysis","systematic mapping study","thematic analysis","qualitative coding","prompt engineering","LLM evaluation","human-in-the-loop"],"falsifier":"A replication that relaxes the prompt-detail requirement (DE11) and includes the 13 excluded studies would refute the equivalence claim if those studies consistently report LLM outputs worse than human coding or high hallucination rates.","tokens_in":10378,"feed_emoji":"🤖","tokens_out":6324,"duration_ms":58590,"temperature":0.7,"pith_summary":"This paper tries to establish where and how large language models (LLMs) have actually been used to perform qualitative analysis, and whether the results hold up against traditional human coding. Reviewing 354 candidate papers and admitting eight that both apply an LLM to qualitative data and report the prompt engineering used, it finds that most studies judged the LLM's output equivalent to human analysis in education, healthcare, culture, and technology contexts. It also finds that no included study came from software engineering and that no standard evaluation metric exists across the field. The practical upshot is that LLM-assisted coding is plausible for open coding and theme extraction but still depends on carefully engineered prompts and human oversight.","feed_headline":"Eight-study map: LLMs equal human coding in most cases","feed_subtitle":"In five of eight qualifying studies, LLM output was rated equivalent to traditional analysis; one was better, one worse.","key_machinery":"The carrying mechanism is the systematic-mapping protocol itself: a search string run across six databases, explicit inclusion and exclusion criteria, and a sixteen-item data-extraction questionnaire (DE1–DE16) mapped to five research questions. The load-bearing criterion is DE11, which required every included study to describe its prompt engineering; this cut the eligible set from 21 to 8 studies and makes the resulting synthesis a map of reproducible LLM-assisted analysis rather than of all LLM-for-qualitative work.","core_discovery":"The paper's central claim is that the eight qualifying primary studies show LLMs are already being applied to qualitative analysis across diverse domains, that in the majority of comparisons they perform equivalently to traditional human analysis, and that the main barriers are dependence on well-structured prompts, occasional hallucinations, and limited contextual sensitivity. The paper further claims that this equivalence is not yet backed by a standardized evaluation metric, and that the absence of software-engineering applications marks a research gap rather than evidence of failure. As direct corollaries, its findings imply that LLMs can shorten coding from weeks to hours and are best used as aids, with human analysts retaining final interpretive authority.","pith_inferences":["Editorial inference: if prompt-detail reporting becomes standard, the map is likely to grow quickly and the equivalence result may shift, since the DE11 filter probably selects for more careful and complete studies.","Editorial inference: a head-to-head benchmark where the same interview corpus is coded by several LLMs and several human teams under a fixed metric such as Cohen's kappa would directly test whether the apparent equivalence is real or an artifact of heterogeneous evaluations.","Editorial inference: the reported 'weeks to hours' speed gain, if replicated, changes the cost structure of qualitative research by making much larger corpora feasible, but the risk of unnoticed hallucinated codes grows with scale."],"forward_implications":["Researchers can treat LLM-assisted open coding and theme extraction as a time-saving step that produces results broadly comparable to human coding, while keeping final interpretation with humans.","The field lacks a standardized evaluation metric, so until one is adopted, comparisons of LLM versus human analysis will remain difficult to aggregate across studies.","The absence of software-engineering primary studies is a concrete opening: requirements engineering and user-feedback analysis are named as untested applications.","Prompt engineering is not a peripheral detail but a core part of the method; studies that omit prompt details are currently excluded from evidence syntheses."],"supporting_citations":[{"why":"Supplies the systematic-review guidelines that structure the mapping protocol.","marker":"[16]"},{"why":"Primary study comparing inductive thematic analysis of healthcare interviews by an open-source LLM with traditional methods; reports the weeks-to-hours speed gain and equivalent performance.","marker":"[9]"},{"why":"Primary study probing the limits of inductive thematic analysis with an LLM; supplies evidence of contextual limits and prompt dependence.","marker":"[13]"},{"why":"Primary study on coding open-ended responses with pseudo-response generation; provides comparison metrics and evidence of over-segmentation.","marker":"[8]"},{"why":"Primary study on automated categorization with an LLM; contributes an equivalence result and notes loss of detail.","marker":"[14]"},{"why":"Primary study presenting a collaborative coding workflow with an LLM; provides the user-study comparison and agreement metrics.","marker":"[20]"},{"why":"Primary study comparing LLM-assisted qualitative research with traditional analysis; supplies the one case where the LLM was judged less effective.","marker":"[18]"},{"why":"Primary study on topic modeling in song lyrics; supplies the one case where LLM performance was judged superior.","marker":"[19]"},{"why":"Primary study applying deep-learning models to analyze social construction of knowledge; adds another equivalence result.","marker":"[17]"},{"why":"Cited in the method to justify using a conversational LLM for data extraction; reports high precision and recall for the ChatExtract approach.","marker":"[12]"}],"fun_headline_variants":["LLMs match human qualitative coding in most studies","Systematic map: LLMs equal humans in majority of tests","LLMs speed up qualitative analysis but need careful prompts","LLMs rival humans in coding, yet lack robustness","LLMs in qualitative research: promising but not yet foolproof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions rest on the assumption that the eight studies which happened to document their prompt engineering fairly represent all LLM qualitative-analysis research; if studies without prompt details carry different evidence about effectiveness, the equivalence claim would not generalize.","fun_headline_variants_meta":{"raw":{"variants":["LLMs match human qualitative coding in most studies","Systematic map: LLMs equal humans in majority of tests","LLMs speed up qualitative analysis but need careful prompts","LLMs rival humans in coding, yet lack robustness","LLMs in qualitative research: promising but not yet foolproof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1696,"prompt_tokens":827,"completion_tokens":869,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":790}},"tokens_in":443,"tokens_out":869,"duration_ms":22664,"temperature":1.0,"reasoning_tokens":790,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:56:18.393890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that relaxes the prompt-detail requirement (DE11) and includes the 13 excluded studies would refute the equivalence claim if those studies consistently report LLM outputs worse than human coding or high hallucination rates.","supporting_citations":[{"cited_title":"Guidelines for Performing Systematic Literature Reviews in Software Engineering,","cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review guidelines that structure the mapping protocol."},{"cited_title":"Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,","cited_arxiv_id":null,"evidence_quote":"Primary study comparing inductive thematic analysis of healthcare interviews by an open-source LLM with traditional methods; reports the weeks-to-hours speed gain and equivalent performance."},{"cited_title":"Performing an inductive thematic analysis of semi- structured interviews with a large language model: An exploration and provocation on the limits of the approach,","cited_arxiv_id":null,"evidence_quote":"Primary study probing the limits of inductive thematic analysis with an LLM; supplies evidence of contextual limits and prompt dependence."},{"cited_title":"Coding Open-Ended Responses using Pseudo Response Generation by Large Language Models,","cited_arxiv_id":null,"evidence_quote":"Primary study on coding open-ended responses with pseudo-response generation; provides comparison metrics and evidence of over-segmentation."},{"cited_title":"Artificial Intelligence and content analysis: the large language models (LLMs) and the automatized cate- gorization,","cited_arxiv_id":null,"evidence_quote":"Primary study on automated categorization with an LLM; contributes an equivalence result and notes loss of detail."},{"cited_title":"LLMusic: Topic Modeling in Song Lyrics Combining LLM, Prompt Engineering, and BERTopic","cited_arxiv_id":null,"evidence_quote":"Primary study on topic modeling in song lyrics; supplies the one case where LLM performance was judged superior."},{"cited_title":"Deep Learning Models for Analyzing Social Construction of Knowledge Online,","cited_arxiv_id":null,"evidence_quote":"Primary study applying deep-learning models to analyze social construction of knowledge; adds another equivalence result."}],"review_version":1}