{"id":"f46bc5d7-1259-4f99-a8cf-53886e95eb51","arxiv_id":"2603.08951","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"GenAI tools can usefully assist narrow deductive coding and transcription in qualitative SE research, but current evidence does not support autonomous or interpretive use.","lead":"This paper argues that generative AI is not a silver bullet for qualitative research in software engineering, and reviews where current evidence supports its use. It is worth reading because it maps where GenAI can help (deductive coding, transcription, summarization) and where it falls short (interpretive, constructivist methods).","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paper's 'cannot make sense of novel contexts' premise is asserted, not established; if a blind evaluation shows LLM themes capture latent interpretation, the central 'not an autonomous researcher' conclusion weakens.","rationale":"The paper's central claim is explicitly anchored to current evidence: 'in SE studies, current evidence shows that GenAI is useful mainly as a coding or summarization aid... not as an autonomous qualitative researcher' (§5). The empirical scan and cited studies support a narrow reading: most successes are deductive annotation. However, the paper repeatedly strengthens this into a categorical statement about LLM capabilities—'GenAI tools cannot make sense of novel contexts' (§4.2) and 'lacks the epistemological grounding, reflexivity, and ethical responsibility expected of human researchers' (§5). The latter is not a direct empirical claim; it involves contested views of constructivist epistemology. The paper cites one group's rejection of GenAI for reflexive research [27] as though it settles the question, while also citing studies that use LLMs in thematic analysis [13, 30, 36]. This tension is the weakest point: if LLMs can demonstrably participate in meaning-making in a way that constructivist researchers find valid, the central 'no silver bullet' argument loses its strongest footing and would need to be reduced to a 'current evidence is limited' claim. The proposed blind evaluation would test the absolute premise. The convenience sample and LLM screening issue are real but secondary because the central claim does not depend solely on that scan; the philosophical premise does. Given the paper already hedges in §6 and frames many conclusions as 'future plans,' the appropriate verdict remains CONDITIONAL, not REJECT: the authors should either weaken the categorical assertions or provide evidence for them. Thus no change to the reader's verdict.","tokens_in":10609,"tokens_out":6678,"duration_ms":66363,"concrete_test":"Pre-registered blind evaluation: take a fresh SE interview dataset from an under-represented, non-Western context that is out-of-distribution for current LLMs; have a state-of-the-art LLM produce themes and interpretations under a constructivist/reflexive prompt; have experienced qualitative researchers rate the outputs for latent interpretation, contextual sensitivity, and reflexivity against human-generated themes, blinded to source. If LLM themes are rated at or above human level on latent interpretation, the premise fails; if not, the premise is supported. Re-running this on the Montes et al. [30] dataset with original authors as raters would provide a direct point of comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—GenAI is useful only as a coding/summarization aid and not as an autonomous qualitative researcher—stands or falls with the asserted premise in §4.2 that 'GenAI tools cannot make sense of novel contexts' and the §2.1 claim that a 'statistical neural network' cannot participate in the co-construction of meaning. These are not derived from the cited evidence; they are philosophical/empirical assertions, and 'sense-making' is never operationally defined. The paper itself relies on studies (Montes et al. [30], Dai et al. [13]) in which LLMs do produce themes and codes in interpretive workflows, and it concedes LLMs can summarize and translate novel texts, which involves some form of contextual understanding. Because the premise is used to rule out LLMs for constructivist/interpretivist studies (§4.2, §5), a single demonstration that an LLM can generate context-sensitive, latent interpretations in a novel SE setting would falsify the strongest version of the claim. The paper would still be defensible as a cautiously worded review of current evidence, but not as a categorical limitation. This internal tension is the most load-bearing soft spot.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a position/reflection piece for the software-engineering community. It argues that claims that GenAI can automate qualitative analysis overgeneralize from narrow successes. The authors distinguish research strategies (respondent, field, lab, data) and dimensions of qualitative work (epistemology, coding strategy, granularity, iteration and researcher roles), then review a convenience sample of recent ICSE, CHASE, and CSCW papers and synthesize the emerging literature on LLM-assisted coding, summarization, and thematic analysis. Their central conclusion is that GenAI is useful mainly as a coding or summarization aid—particularly for deductive, low-context annotation—but is not an autonomous qualitative researcher, especially for constructivist and interpretivist studies. The paper closes with quality considerations and a five-item research agenda for benchmarking, interpretive methods, human-AI workflows, standards, and paradigm reconciliation.","tokens_in":10873,"tokens_out":6380,"duration_ms":64626,"significance":"If accepted, the paper provides a timely counterweight to enthusiasm about GenAI in qualitative SE research and offers useful guidance on where automated assistance is currently appropriate. It is heavily cited, carefully hedged in most of its formulations, and explicitly discloses the preliminary nature of its empirical scan, which is accompanied by a replication package—a strength. However, the paper is primarily a synthesis and argument; its only new empirical contribution is a small, convenience-sampled scan whose loading on the central conclusion is limited but still present. The broadest claim—that GenAI tools cannot make sense of novel contexts and are therefore unsuitable for constructivist work—is asserted rather than established; it is a falsifiable empirical hypothesis, but no direct test is offered. If that claim is weakened to 'current evidence does not show,' the paper remains a valuable scoping review and roadmap for the community.","major_comments":[{"comment":"The statement 'GenAI tools cannot make sense of novel contexts' is load-bearing: it underlies the conclusion that GenAI cannot be an autonomous qualitative researcher and is epistemologically incompatible with constructivist methods. Yet 'sense making' is never operationally defined, and no evidence is given for the categorical impossibility. The paper itself cites LLMs successfully summarizing/translating novel texts and producing themes that are 'less aware of latent interpretations' [30]—a graded, empirical failure mode, not an absence of sense-making. This tension weakens the central claim. Please either (a) define sense-making and specify a falsifiable test, or (b) replace the categorical 'cannot' with 'current evidence does not show,' which is the paper's own formulation in §5. Without this change, the philosophical claim overreaches the evidence.","section":"§4.2 and §2.1"},{"comment":"The screening pass uses ChatGPT without reporting precision, recall, or validation against a hand-checked sample; the keyword search and single-author manual check are not reliability-assessed. The claim that 'none of the ICSE or CHASE papers' used GenAI and that CSCW prevalence is 7/209 (3.3%) rests on this unvalidated pipeline. There is also a numeric inconsistency: the text later says '7 CSCW (2.7%)' without a denominator; 7/312 is 2.2% and 7/209 is 3.3%. Since the section is explicitly preliminary, this is fixable by reporting the ChatGPT accuracy on a small gold-standard subset, clarifying the denominator, and labeling the counts as indicative. I would not reject the paper over this, but the numbers need cleaning.","section":"§3.1, Table 1"},{"comment":"The paper conflates an empirical claim about current LLM performance with a principled/epistemological claim about the nature of a 'statistical neural network'. §4.2 says LLMs 'lack socially embedded sense making' and are 'fundamentally at odds' with constructivist approaches, while §5 summarizes the argument as 'current evidence shows.' These have different scopes. If the stronger claim is intended, the paper should engage with the possibility that constructivist researchers might use GenAI in reflexive workflows and explain why the weak-but-real contextual sensitivity demonstrated in translation/summarization tasks fails to count. If the weaker claim is intended, the §4.2 wording should be aligned. Distinguishing the two would make the central argument more precise and defensible.","section":"§4.2, §5"}],"minor_comments":[{"comment":"Use one denominator for all percentages. The same seven papers are described as 3.3% of 209 and 2.7% of an unstated base; 7/312 is 2.2%.","section":"Table 1"},{"comment":"Specify whether the keyword search was applied to full text or only to titles/abstracts, and state the exact search string so the filtering step is independently replicable.","section":"§3.1"},{"comment":"The sentence 'This will dramatically scale up the number of artifacts such studies can analyze' reads as promotional; consider a neutral phrasing consistent with the paper's cautious stance.","section":"§4.1"},{"comment":"The aside that some researchers question rigor criteria is interesting but not integrated into the quality-criteria discussion; a sentence linking it to the argument would help.","section":"Footnote 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful position piece, and the empirical scan is honestly disclosed; I would not reject it. The main issue is the categorical 'cannot' in §4.2, which is load-bearing for the central claim and needs either operationalization or hedging to a 'current evidence' claim. The percentage inconsistency in §3.1 is a small but embarrassing error. Heavy reliance on the authors' own earlier work is understandable given their prior contributions, and I see no further citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a thoughtful reflection piece that should be read by anyone thinking about using GenAI in qualitative SE research. The authors do a good job of mapping the landscape, but their own new data (the venue scan) is too slight to be more than a preliminary observation.\n\nThe paper's real contribution is the synthesis. They bring together recent studies (Ahmed et al., Shah et al., Montes et al., Barros et al.) and organize them around a few useful axes: epistemological orientation, coding strategy, data granularity, and iteration. That framework in §2 is worth taking into the design of future mixed human-AI studies. The discussion of reliability, validity, and reflexivity in §5 is also a pragmatic guide for researchers who need to decide what to report.\n\nTable 1 is the only new empirical piece: a convenience sample of three 2025 venues, screened by ChatGPT with no reported accuracy, with single-author manual checks. The authors disclose all of this and still present the scan as \"limited.\" That transparency earns them some credit, but it means the scan cannot carry much weight. The 3.3% prevalence figure at CSCW is an anecdote, not a measurement.\n\nThe larger concern is a categorical tone that shows up in a few places. In §4.2 they say GenAI tools cannot make sense of novel contexts. In §2.1 they suggest a statistical neural network cannot participate in the co-construction of meaning. These are asserted, not argued, and they sit uneasily beside the paper's own acknowledgment that LLMs can translate and summarize novel texts, and that Montes et al. found a slight preference for GenAI themes in a blind trial. If those categorical claims were load-bearing, the paper would have a problem. But they are not load-bearing. The central thesis, stated clearly in §5, is about current evidence: GenAI is useful mainly as a coding or summarization aid, not as an autonomous qualitative researcher. That claim is well supported by the literature the authors cite, and they hedge it appropriately.\n\nThe citation pattern is fine; the self-citations point to real, published work. The missing artifact link is a minor annoyance, easily fixed.\n\nWho should read this? Empirical SE researchers and anyone designing or reviewing qualitative studies with GenAI components. It is not a breakthrough empirical result, but it is a useful, honest map of where the field is. I would send it to peer review and ask the authors to soften the categorical statements and either provide the replication package or label the scan explicitly as exploratory. The central message will survive that revision, and the paper will be more defensible.","headline":"A useful, well-hedged synthesis arguing that GenAI is not a silver bullet for qualitative SE research, but the paper's own empirical scan is too thin to carry much weight and a few categorical claims are stronger than the evidence justifies.","tokens_in":11356,"tokens_out":2870,"would_cite":true,"duration_ms":27802,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Autonomous GenAI cannot replace qualitative researchers in software engineering; the current evidence supports only narrow, low-context tasks like deductive coding, summarization, and transcription.","keywords":["qualitative research","software engineering","generative AI","large language models","thematic analysis","grounded theory","deductive coding","research methods"],"falsifier":"Conduct a blinded, preregistered study in which a zero-shot large language model—given no codebook and no examples—analyzes interview transcripts from a novel, previously unresearched organizational context, and have experienced grounded-theory researchers, blind to the source, rate whether the model's themes capture latent interpretations and contextual nuance as well as human-generated themes; if the model consistently matches or exceeds human ratings, the paper's claim that GenAI cannot make sense of novel contexts and is unsuitable for interpretive work would be refuted.","tokens_in":10494,"feed_emoji":"🤖","tokens_out":3405,"duration_ms":33426,"temperature":0.7,"pith_summary":"This paper tries to establish that claims of generative AI automating qualitative research in software engineering are overgeneralized from narrow successes. It argues that GenAI support must be tailored not only to the data but to the specific research strategy, epistemological stance, coding approach, and researcher role. Reviewing recent evidence, it finds GenAI useful mainly as a transcription, summarization, translation, or deductive-coding aid, not as an autonomous qualitative analyst. The paper matters because it gives software engineering researchers a structured way to decide where GenAI can help and where it risks undermining interpretive depth and constructivist values.","feed_headline":"GenAI helps code, not think, in qualitative SE studies","feed_subtitle":"Evidence supports GenAI for transcription and deductive coding—not for interpretive, constructivist analysis.","key_machinery":"The central organizing device is a two-dimensional map of qualitative research: research strategies (respondent, field, lab, data) crossed with dimensions of epistemology (postpositivist vs. constructivist), coding strategy (inductive, deductive, hybrid), data granularity and type, and iteration/researcher roles. This map does the argumentative work of predicting where GenAI fits—postpositivist, deductive, low-context, fine-grained annotation tasks—and where it does not: constructivist, inductive, context-rich, reflexive interpretation. The paper also reworks standard quality criteria (reliability, validity, reflexivity, ethics) as lenses for evaluating GenAI-assisted studies.","core_discovery":"The paper's central claim is that current evidence supports GenAI only for low-context, deductive annotation tasks and for acceleration aids such as transcription and summarization, while genuinely interpretive work—thematic synthesis, grounded theory, ethnography, and other constructivist analyses—remains outside GenAI's demonstrated capabilities and is in epistemological tension with it. The authors ground this in a review of the spectrum of qualitative research strategies, a small empirical scan of 2025 conference papers showing minimal adoption in software engineering venues (zero of the ICSE and CHASE papers reviewed reported GenAI use in coding, while seven of 209 CSCW qualitative pape","pith_inferences":["Beyond the paper: if constructive, interpretive research remains resistant to GenAI, the field may split into scalable, artifact-heavy deductive analyses that increasingly use GenAI and a smaller, more human-centered interpretive tradition whose value rises precisely because it cannot be automated.","Beyond the paper: the paper's framework yields a testable prediction—GenAI performance in coding tasks should track the clarity of the codebook and the contextual dependence of the labels; a benchmark that varies both dimensions could draw a capability frontier.","Beyond the paper: the tentative evidence that adoption is driven more by venue timing than by discipline resistance could be tested by tracking whether ICSE and CHASE papers catch up to CSCW in GenAI use in subsequent years.","Beyond the paper: the epistemological mismatch might narrow if the community develops standards for documenting a model's positionality, such as reporting training-data provenance and prompt choices, in much the way human positionality statements are used."],"forward_implications":["GenAI can be used safely as a fast second coder for well-defined deductive codebooks, as a transcription and translation tool, and as a source of candidate themes that humans must validate.","Inductive and interpretive studies, including grounded theory, ethnography, and reflexive thematic analysis, should not treat GenAI output as analysis; at most it can surface patterns for human interpretation.","Software engineering venues should require disclosure of any GenAI involvement in qualitative coding, including subtle tool autocomplete features, along with prompts, model versions, and parameter settings.","Reliability metrics such as inter-coder agreement are insufficient on their own; they must be paired with reflexivity, positionality statements, and meaningful member checking when GenAI participates.","Future benchmarking should compare human-human, human-AI, and AI-AI agreement across diverse software engineering artifact types to map where GenAI can substitute for multiple coders and where it cannot."],"fun_headline_variants":["GenAI automates coding, not insight, in qualitative SE","Qualitative SE: GenAI works for deduction, not interpretation","GenAI aids transcription, but not interpretive analysis","Evidence: GenAI helps low-context coding, not deep insight"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that GenAI tools, being statistical neural networks, cannot make sense of novel contexts and cannot participate in the co-construction of meaning that constructivist qualitative research requires; the paper asserts this without direct empirical evidence, and if a future model demonstrably did either, the central argument would need substantial revision.","fun_headline_variants_meta":{"raw":{"variants":["GenAI automates coding, not insight, in qualitative SE","Qualitative SE: GenAI works for deduction, not interpretation","GenAI aids transcription, but not interpretive analysis","Evidence: GenAI helps low-context coding, not deep insight"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2609,"prompt_tokens":700,"completion_tokens":1909,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":1849}},"tokens_in":444,"tokens_out":1909,"duration_ms":13012,"temperature":1.0,"reasoning_tokens":1849,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T18:30:04.386458+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a blinded, preregistered study in which a zero-shot large language model—given no codebook and no examples—analyzes interview transcripts from a novel, previously unresearched organizational context, and have experienced grounded-theory researchers, blind to the source, rate whether the model's themes capture latent interpretations and contextual nuance as well as human-generated themes; if the model consistently matches or exceeds human ratings, the paper's claim that GenAI cannot make sense of novel contexts and is unsuitable for interpretive work would be refuted.","supporting_citations":[],"review_version":1}