{"id":"7628e7a3-2b8b-4b34-92ff-bee7323afd39","arxiv_id":"2412.06564","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic mapping study of 20 papers shows LLMs are mainly used for coding and thematic analysis, with efficiency benefits but reliability, nuance, and privacy limitations.","lead":"This paper reviews 20 studies that used large language models, like ChatGPT, to help analyze qualitative data such as interview transcripts in software engineering research. It maps the tasks, benefits, and limitations of these tools, and offers recommendations for using them responsibly.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The descriptive map depends on a small, potentially biased corpus; the ACM export cap and title/abstract screening could systematically exclude relevant studies, skewing the reported applications and recommendations.","rationale":"The reader's weakest assumption is indeed the most load-bearing concern. The paper's central contribution is a descriptive synthesis of how LLMs are used in qualitative analysis; if the 20 selected papers are unrepresentative, the synthesis could be misleading. The ACM export limit and reliance on title/abstract screening are specific, acknowledged, and unmitigated. I considered other issues—the 21/20 count discrepancy and the non-verifiable dataset link—but these are minor and do not threaten the central claim. The representativeness concern is substantive, yet the paper explicitly labels its findings as preliminary and includes an honest threats-to-validity section. A CONDITIONAL verdict that requires either quantitative mitigation or a narrowed claim is therefore appropriate, which matches the reader's recommendation. My proposed concrete test would settle whether the concern actually lands, but until it is performed the uncertainty remains, so the verdict should stay unchanged.","tokens_in":11029,"tokens_out":5040,"duration_ms":51908,"concrete_test":"Run the same search string on the ACM Digital Library using the API or by paginating through all ~3,000 results, then apply EXC1–EXC5 and INC1 to the complete set and re-derive Tables II–VII. If the updated set changes any category's article count by more than 20% or adds a new application, benefit, or limitation category, the export cap materially biased the findings. Additionally, randomly sample 100 of the 1,574 discarded papers and screen their full texts (not just abstracts) for INC1; this estimates the false-negative rate of title/abstract screening and can show whether 'primarily used for coding and thematic analysis' is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—a descriptive map of current LLM-assisted qualitative analysis practices and recommendations for software engineering—rests on the representativeness of the 20 included studies. The search process has a known blind spot: the ACM export cap of 1,000 from ~3,000 results, plus the exclusion of short papers, non-downloadable items, and non-English works. More importantly, screening relied on titles and abstracts only, which is risky in a fast-moving area where LLM use may be described only in the full text. The authors acknowledge this in Section V-C but do not quantify the potential bias. If the missing ACM results or abstract-screened false negatives contain studies with different application areas (e.g., discourse analysis, narrative analysis) or different reported limitations, the synthesized tables (II–VII) and the derived recommendations would be skewed. This is load-bearing because the paper's stated contributions are precisely the map and recommendations; a biased sample would make the findings an artifact of the search strategy rather than a reliable description of the literature.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a systematic mapping study of how large language models (LLMs) are used to support qualitative analysis, with the stated aim of informing empirical software engineering research. The authors searched ACM, IEEE, Scopus, and selected venues; after title/abstract screening and full-text review they included 20 studies (the abstract says 21). They extract data on qualitative analysis methods, LLM tools, prompting techniques, data types, reported benefits, and reported limitations, and synthesize these into tables (Tables II-VII). The central findings are that LLMs are primarily used for coding, thematic analysis, and data categorization; benefits include efficiency gains and support for novice researchers; limitations include output variability, limited interpretive depth, and ethical/privacy concerns. The paper ends with a set of recommendations for researchers and a threats-to-validity discussion.","tokens_in":11211,"tokens_out":5120,"duration_ms":54434,"significance":"If the corpus is representative, the paper offers a useful early map of an emerging area and a set of plausible, though preliminary, recommendations for software engineering researchers. Methodological strengths include adherence to established systematic-review guidelines, three-researcher screening, two-researcher extraction, an openly available dataset, and an honest threats-to-validity section. The main weakness is that the descriptive map and the derived recommendations rest on a small corpus that is subject to acknowledged but unquantified selection limitations, especially the ACM export cap and title/abstract screening. The paper does not claim formal predictive power, and no circularity issue arises because the findings are obtained through thematic synthesis rather than by fitting a model to its own outputs.","major_comments":[{"comment":"The acknowledged ACM export cap (approximately 3,000 results, only the top 1,000 exportable) means that a substantial portion of ACM hits was never screened. Because the paper's central contribution is a descriptive map of LLM-assisted qualitative analysis, this blind spot could systematically exclude studies with different application areas or different reported limitations, which would directly affect the synthesis in Tables II-VII and the recommendations in Section V-B. The authors acknowledge the risk but do not provide a quantitative mitigation. Please add a PRISMA-style flow diagram with the number of studies excluded at each criterion, state how many included studies came from each database and from the manual searches, and discuss whether any identified themes are supported by only one database or one source. This information is necessary for readers to assess the representativeness of the 20-study corpus.","section":"Section V-C (External Validity); Section III (Search Strategy)"},{"comment":"The screening was based solely on titles and abstracts before applying the inclusion criterion, which is a known risk in a field where LLM use may be described only in the full text. The authors acknowledge this risk in Section V-C but do not report any calibration, such as a full-text check of a random sample of excluded papers or an inter-reviewer agreement measure. Please add a quantitative or at least systematic qualitative account of screening reliability (for example, Cohen's kappa or a description of disagreements and resolutions). Without this, the reliability of the selection step is difficult to assess, and the load-bearing assumption that the 20 included studies represent the literature is left unsupported.","section":"Section III (Selection Process)"}],"minor_comments":[{"comment":"The abstract states that 21 relevant studies were analyzed, but the body, the method section, and Table I consistently report 20 studies. Please correct the abstract to match the actual corpus.","section":"Abstract; Section III; Section IV; Table I"},{"comment":"In the last row of Table VII, the article ID is given as \"A20\" while all other entries use the zero-padded format \"A020\". Please make this consistent.","section":"Table VII"},{"comment":"The sentence \"This approach is not aimed at deeply exploring specific aspects of a research problem\" appears to be a typo; it contradicts the surrounding discussion of qualitative research. It should likely read \"This approach is aimed at deeply exploring...\" or \"This approach is not limited to...\".","section":"Section I"},{"comment":"Recommendations 2 and 3 both state essentially the same data-protection advice: researcher should anonymize data and avoid direct input of sensitive information into LLMs. Please merge these into a single recommendation or differentiate them clearly to avoid redundancy.","section":"Section V-B (Recommendations)"},{"comment":"The dataset link is described only as \"here\" with no actual URL or persistent identifier. Please provide a working DOI or URL so that the promised artifact is accessible.","section":"Section VII (Dataset Availability)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate mapping study that could become suitable for publication after revision. The main concern is the representativeness of the corpus: the ACM export cap and title/abstract screening are acknowledged but not quantitatively mitigated, and the abstract/body count discrepancy (21 vs. 20) suggests the manuscript needs careful checking. The editors may want to emphasize that the authors should provide a flow diagram and per-database breakdown, since this is standard for mapping studies and will directly address the reviewer concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zeynep, quick take on arXiv:2412.06564. It's a systematic mapping study of LLM use in qualitative analysis, aimed at empirical software engineering. The main contribution is a structured desk-map: 20 studies, tables of tasks, tools, benefits, limitations, recommendations. That's useful but not new in kind; the findings (coding, thematic analysis, efficiency, reliability, privacy) match what's already in the primary studies and adjacent reviews. The SE framing and the recommendations are the modest added value.\n\nThe paper does the mapping-study craft well. Screening by three researchers, extraction by two, thematic synthesis with a second coder, and an honest threats section. It's clearly written, and the recommendations (disclose LLM version, try prompt strategies, data anonymization, human oversight) are sensible and grounded in the extracted evidence.\n\nSoft spots, in order of importance. The abstract says 21 studies; the body and Table I say 20. That's a concrete error to fix. More substantive: the automatic search hit the ACM export cap of 1,000 out of ~3,000, and screening was title/abstract only. The authors flag both in Section V-C, and they supplement with manual searches and other libraries. The stress-test worry that this biases the map is real but limited: the findings are descriptive and align with the broader LLM-qualitative literature, so a few missed studies are unlikely to flip the conclusions. Still, the paper would be stronger with a sensitivity estimate or some analysis of what the cap might exclude. The dataset link is just 'here' with no URL, which is annoying for a reproducibility claim. Sample size of 20 is acknowledged.\n\nWho's this for? Researchers in SE who want a quick overview of how LLMs are being used in qualitative analysis and which pitfalls are documented. It's not a theoretical breakthrough, but it's a competent consolidation. I'd send it to peer review with minor revisions. The references look appropriate, no self-citation issues, and the limitations are stated rather than hidden.\n\nMy recommendation: engage with it, ask for the count fix, dataset link, and a bit more on the export-cap sensitivity, then publish.","headline":"Competent mapping study of LLM-assisted qualitative analysis; findings are useful but not new, and the corpus issues are real but not fatal.","tokens_in":11709,"tokens_out":1673,"would_cite":true,"duration_ms":16769,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models are now used mainly for coding and thematic analysis in qualitative research, with efficiency gains offset by output variability and privacy risks.","keywords":["large language models","qualitative analysis","systematic mapping study","software engineering","thematic analysis","prompt engineering","empirical software engineering","research ethics"],"falsifier":"A replication of the search that relaxes the five-page minimum, includes non-English work, and exports more than the top 1,000 ACM results would falsify the mapping if it produced a substantial set of additional primary studies whose applications, benefits, or limitations differ from the clusters reported here.","tokens_in":10864,"feed_emoji":"🤖","tokens_out":5448,"duration_ms":51458,"temperature":0.7,"pith_summary":"This paper is a systematic mapping study asking how large language models are actually used in qualitative analysis and what that means for empirical software engineering. It argues that the current literature concentrates LLM use on coding, thematic analysis, and data categorization, with ChatGPT-family models and prompt engineering dominating practice. The reported benefits are real but bounded: faster, cheaper, less cognitively demanding analysis and a lower barrier for novice researchers. The reported limitations are equally consistent: outputs vary run to run, models miss subtle context, and privacy and transparency problems are unresolved. The paper concludes that LLMs can support, not replace, human interpretation, and that structured guidelines for disclosure, prompting, data protection, and human oversight are the immediate next step.","feed_headline":"LLMs speed qualitative coding but still need human interpretation","feed_subtitle":"Efficiency gains are real, but output variability, shallow interpretation, and privacy risks keep humans in the loop.","key_machinery":"The load-bearing mechanism is the systematic mapping study protocol. A search string and manual searches across software engineering venues and methodology journals produced 2,574 candidate records, which three researchers reduced by title and abstract screening under five exclusion criteria and one inclusion criterion to 20 full papers. The extracted data were then synthesized by thematic analysis, producing the classification tables that carry the argument: qualitative methods, LLM tools, techniques, data types, benefits, and limitations. Those tables are the machinery because every conclusion in the paper is read off them.","core_discovery":"The central discovery is a descriptive map, not a new technique. Analyzing 20 primary studies, the authors find that LLMs are being used for open and deductive coding, thematic analysis, grounded theory, screening, topic modeling, content analysis, vignette analysis, and critical review, with coding and thematic analysis the most common. The dominant tools are ChatGPT, GPT-3.5, and GPT-4, and the dominant technique is prompt engineering, with fine-tuning and few-shot prompting appearing less often. Against those applications, the reported benefits cluster into theme and pattern identification, efficiency, coding support, autonomy for beginners, enhanced collaboration, and triangulation; the reported limitations cluster into consistency and hallucination problems, shallow high-level comprehension, ethics, privacy and transparency deficits, technical constraints such as token limits, and dependency risks. The paper's conclusion is that this evidence supports a division of labor where LLMs accelerate labor-intensive analysis but human expertise remains responsible for interpretation, and that preliminary recommendations, disclose the model and version, experiment with prompts, protect data, keep humans in the loop, balance with traditional methods, follow AI ethics guidelines, and discuss validity threats contextually, should guide integration into software engineering research.","pith_inferences":["A testable extension of the triangulation finding: treat repeated runs of the same LLM on the same data as an ensemble of pseudo-coders and measure inter-run agreement; if agreement is low, a reliability threshold could be set before human review.","The privacy limitation has an engineering consequence the authors do not draw: on-premises or open-weight models with strict logging controls would preserve the efficiency benefits while removing the main data-exposure risk for proprietary software-engineering data.","A direct corollary for guidelines is that, since model versions change rapidly, a reporting checklist covering model, version, date, prompts, output samples, and human edits could let later readers judge reproducibility, a concern implicit in the paper's disclosure recommendation.","If coding quality becomes comparable to human coders for well-defined tasks, the bottleneck in qualitative software engineering research shifts from coding effort to construct validity, whether themes found by either humans or LLMs actually answer the research question."],"forward_implications":["For empirical software engineering, LLMs can absorb the most time-consuming parts of qualitative analysis, initial open coding, categorization, and screening, letting researchers spend their effort on interpretation.","The dominant practice of prompt engineering becomes a methodological skill: the type of prompt, instructional, persona-based, or chain-of-thought, changes the quality of the analysis, so reporting prompt strategies should become part of study design.","Because output variability and hallucinations are persistent, LLM-assisted analyses need an explicit validation step, such as triangulation or comparison with human coding.","Data protection has to be designed into the workflow before any participant or company data is entered into a proprietary model, not added afterward.","The recommendations imply that reporting standards for LLM-assisted qualitative studies should include the model and version, prompt choices, and validity-threat discussion."],"supporting_citations":[{"why":"Supplies the systematic-review protocol that defines the search, screening, and synthesis steps.","marker":"[33]"},{"why":"Provides the thematic-synthesis method the authors use to group extracted codes into benefit and limitation themes.","marker":"[34]"},{"why":"Primary study in the dataset on using LLMs to aid analysis of textual data, used as evidence for coding support and its limits.","marker":"[21]"},{"why":"Primary study describing CollabCoder, an LLM-based collaborative qualitative coding workflow, evidence for coding support and collaboration benefits.","marker":"[29]"},{"why":"SE-specific exploration of LLM opportunities and challenges in qualitative research, contextualizing the mapping's motivation and ethical concerns.","marker":"[23]"},{"why":"Source of the promise-and-peril framing for LLM assistance in qualitative research, grounding the reported benefits and limitations.","marker":"[20]"},{"why":"Identifies dangers of proprietary LLMs in research, cited to support the ethics, privacy, and transparency limitation.","marker":"[32]"}],"fun_headline_variants":["LLM coding in SE: faster, but human interpretation stays key","Systematic map reveals LLM limits in qualitative SE research","LLMs help code data, not interpret it—review finds","LLM integration in SE analysis: benefits, risks, and guidance","Qualitative SE meets LLMs: efficiency gains, ethical concerns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions assume that the 20 papers surviving the screening are representative of the wider literature on LLM-assisted qualitative analysis.","fun_headline_variants_meta":{"raw":{"variants":["LLM coding in SE: faster, but human interpretation stays key","Systematic map reveals LLM limits in qualitative SE research","LLMs help code data, not interpret it—review finds","LLM integration in SE analysis: benefits, risks, and guidance","Qualitative SE meets LLMs: efficiency gains, ethical concerns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000324,"raw_usage":{"total_tokens":1836,"prompt_tokens":979,"completion_tokens":857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":770}},"tokens_in":595,"tokens_out":857,"duration_ms":8781,"temperature":1.0,"reasoning_tokens":770,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:30:38.597702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication of the search that relaxes the five-page minimum, includes non-English work, and exports more than the top 1,000 ACM results would falsify the mapping if it produced a substantial set of additional primary studies whose applications, benefits, or limitations differ from the clusters reported here.","supporting_citations":[{"cited_title":"An examination of the use of large language models to aid analysis of textual data,","cited_arxiv_id":null,"evidence_quote":"Primary study in the dataset on using LLMs to aid analysis of textual data, used as evidence for coding support and its limits."},{"cited_title":"Collabcoder: a lower-barrier, rigorous workflow for inductive collabo- rative qualitative analysis with large language models,","cited_arxiv_id":null,"evidence_quote":"Primary study describing CollabCoder, an LLM-based collaborative qualitative coding workflow, evidence for coding support and collaboration benefits."},{"cited_title":"Artificial intelligence and qualitative research: The promise and perils of large language model (llm)‘assistance’,","cited_arxiv_id":null,"evidence_quote":"Source of the promise-and-peril framing for LLM assistance in qualitative research, grounding the reported benefits and limitations."},{"cited_title":"The dangers of using proprietary llms for research,","cited_arxiv_id":null,"evidence_quote":"Identifies dangers of proprietary LLMs in research, cited to support the ethics, privacy, and transparency limitation."}],"review_version":1}