{"id":"2bdee42e-eb2c-4df7-9ced-92ca8adb0bec","arxiv_id":"2607.24991","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GUEST gives SE researchers process recommendations for planning, conducting, and reporting GenAI-supported SLRs and independent GenAI tool evaluations under mandatory human oversight.","lead":"The paper proposes GUEST, preliminary process guidelines for using and evaluating generative AI in systematic literature reviews in software engineering. It argues GenAI cannot run unsupervised SLRs and needs human oversight, while offering help on repetitive or validation tasks.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"GUEST's claim to \"enable trustworthy SLRs and rigorous evaluations\" across all major SLR tasks rests on the authors' own two companion studies plus a cross-validation that is not independent of its development inputs.","rationale":"The reader identified the right region of weakness — generalizable task rules grounded in a single-source rapid review, thought experiments, and two companion studies, with evidence openly thin for extraction/RoB/SoE. I agree with that substance but locate the load-bearing point slightly differently: the problem is not only the thinness of the evidence but its provenance and the structure of the validation. The two empirical pillars are the authors' own studies; the §4.5 cross-check against Baltes et al. is not independent because Baltes extends Wagner et al. [97], an explicit input to the recommendation development, and RAISE was itself an input. So the paper's main external corroboration partially re-derives its own premises. That said, this does not warrant a harsher verdict than the reader's CONDITIONAL. The paper is unusually candid: it labels the guidelines preliminary, marks RoB/SoE/extraction recommendations as provisional, devotes §7 to the exact open questions a skeptic would raise, classifies recommendations by novelty kind (Table 4), and its central normative claim (human oversight required; GenAI not capable of unsupervised systematic studies) is modest and well-argued from the nature of GenAI tools (§2.1) rather than from the thin empirical base. The risk is adoption drift — readers treating GUEST's task-level rules as validated standards rather than placeholders — which the CONDITIONAL verdict, with the proviso that under-evidenced task rules stay explicitly provisional, already captures. Hence UNCHANGED with partial agreement.","tokens_in":43287,"tokens_out":2482,"duration_ms":69281,"concrete_test":"Have a team with no connection to the authors apply GUEST end-to-end to one prospective SE SLR: run the screening validation protocol (ScrRev1–3) with WMCC, and attempt to operationalize Ext2–Ext7 and A1–A3 as written. Record (a) whether the WMCC-based tool-adoption decision is stable across a plausible range of FN:FP cost ratios (recompute the [61] reanalysis rankings at, say, 5:1, 10:1, 20:1), and (b) how many of Ext5–7/A1–A3 require author interpretation to execute. If the adoption decision flips with the cost ratio, or the under-evidenced task rules cannot be applied without consulting the authors, the \"enables\" claim should be narrowed to screening and synthesis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that GUEST enables trustworthy GenAI-assisted SLRs and rigorous independent evaluations, with task-level rules spanning screening, extraction, qualitative synthesis, RoB, and SoE. The empirical backbone, by the paper's own account (§4.3, Table 3), covers only two tasks: screening (via the authors' own [61]) and qualitative synthesis (via the authors' own [79], autoethnographic trials without an independent extraction stage, as conceded in §C.4.2). For extraction, RoB, and SoE the recommendations (Ext1–7, A1–A3, RevA1–2) derive solely from thought experiments and briefing notes, with direct evidence described as \"limited\" or \"contradictory\" (§7). The one quantitative anchor — WMCC choosing a tool that misses 5.8% of relevant studies vs 63.3% for accuracy — comes from a self-cited reanalysis and depends on a cost ratio whose value is not empirically justified in this paper. The completeness check in §4.5 is weaker than it appears: Baltes et al. [3] \"refines and extends\" Wagner et al. [97], which was an explicit development input (§3.1), so the \"independent\" baseline shares lineage with the construction set; the paper concedes this but still treats the comparison as validation. RAISE [90] was also an input, so neither arm of §4.5 is a fully external test of whether the thought experiments missed important issues. This does not make the recommendations wrong — they are largely conservative process controls, honestly labelled preliminary, and §7 lists the right open questions — but it means the strongest claim's scope (\"all major SLR tasks\") is supported by evidence that is self-generated for two tasks, absent for three, and corroborated by a partially circular comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper presents GUEST, a preliminary set of process recommendations for software engineering researchers who either (a) evaluate GenAI tools on SLR tasks or (b) use GenAI tools while conducting SLRs. The method has three strands: a rapid review of existing GenAI-use guidelines (single-source SCOPUS, searches closed 29/03/2025, documented in Appendix A with search strings and known-item validation); role-based thought experiments across planning/conduct/reporting stages, supported by per-task briefing notes (Appendix C); and a synthesis of empirically reported failure modes mapped to recommendations (Table 3). The output is a set of planning/conduct/reporting recommendations (Table 5) and task-level recommendations for screening, data extraction, qualitative synthesis, and risk-of-bias/strength-of-evidence assessment (Table 6), each annotated by \"kind\" (SLR-specific vs. general practice, Table 4). The central claims are modest: GenAI requires human oversight and cannot currently conduct unsupervised systematic studies; the recommendations are explicitly preliminary. Empirical grounding is strong for screening and qualitative synthesis (via the authors' peer-reviewed companion studies [61, 79]) and admittedly thin for extraction, RoB, and SoE, which §7 flags as open research questions. A cross-check against Baltes et al. and RAISE is reported in §4.5.","tokens_in":43641,"tokens_out":5185,"duration_ms":189126,"significance":"Timely and useful. As GenAI-assisted SLRs proliferate in SE, the absence of agreed evaluation methodology is a real problem, and the paper's MMRE/crossover analogies for how bad metrics entrench are apt. Strengths worth naming: the rapid review is documented to a reproducible standard (search strings in Tables 9–10, known-item validation in Tables 7–8); Table 3 maps empirically reported failure modes to the recommendations that address them; Table 4's four-kind classification of recommendations — separating SLR-specific contributions from restated good practice — is a methodological contribution other guideline efforts could adopt; the empirical backbone rests on peer-reviewed companion studies rather than self-referential definitions; and §7 lists concrete, falsifiable open questions. The guidance is honest about its limits (§5, A.7) and complementary to rather than duplicative of RAISE. If the framing issues in the major comments are fixed, this will be a standard reference for SE researchers evaluating or using GenAI in SLRs.","major_comments":[{"comment":"§4.5 is titled and framed as 'Validation', and §5.1 states the recommendations 'were corroborated by an independent guideline set', describing Baltes et al. [3] as 'the independent guidelines ... which we deliberately did not use while developing our recommendations'. But §4.5 itself notes Baltes et al. 'refines and extends' Wagner et al. [97], which §3.1 lists as a direct development input, and RAISE [90] is also a stated input (§3.1). Neither arm of §4.5 is therefore independent of the construction set; the exercise is a completeness/consistency check, and it is the paper's only systematic test that the thought experiments did not miss important issues — exactly what §5.1's construct-validity argument rests on. The comparison itself is careful and valuable; the framing overstates it. Suggest retitling/reframing §4.5 as a completeness check, softening 'independent' in §5.1, and adding a","section":"§4.5, §5.1"},{"comment":"The empirical backbone (Table 3, §4.3) covers only two SLR tasks: screening (via [61]) and qualitative synthesis (via [79]). For data extraction, RoB, and SoE, recommendations Ext1–Ext7, A1–A3, Syn1–Syn3, and RevA1–RevA2 rest on thought experiments and briefing notes alone, with direct evidence described as 'limited' or 'contradictory' (§7, Appendix C.3, C.4.2). The paper concedes this honestly, but Tables 5–6 present all recommendations with uniform force, distinguished only by 'kind' (i)–(iv), which classifies novelty, not evidential support. Since the tables are the deliverable practitioners will consult, please annotate each recommendation with its evidence status (companion-study grounded / external empirical / thought-experiment only), paralleling the kind labels, and add one scoping sentence in the abstract or conclusion stating which tasks are evidence-grounded. This is a present","section":"Tables 5–6, §4.3, §7"},{"comment":"The paper's one quantitative anchor — the accuracy-ranked LLM missing 63.3% of relevant studies vs 5.8% for the WMCC-ranked one — comes from a reanalysis of a single 9,695-article evaluation reported in the authors' own [61], and the WMCC ranking depends on the FN:FP cost ratio, whose value is neither justified nor subjected to sensitivity analysis in this paper. Scr2 is a kind-(i) recommendation, i.e., part of the paper's distinctive contribution, yet a reader of this paper alone cannot judge how robust the WMCC mandate is to the chosen ratio. A short paragraph stating the ratio used, its rationale, and the sensitivity of the tool ranking would suffice; alternatively, scope Scr2 the way Ext3 is already scoped ('unless there is an explicit justification for assuming differential costs'). Ext3 shows the authors can write exactly this kind of caveat.","section":"§4.3, Table 6 (Scr2), Appendix C.1"}],"minor_comments":[{"comment":"Eligibility criterion E2 excludes self-published (arXiv-only) papers from the rapid review, yet Baltes et al. [3] — cited as an arXiv preprint (arXiv:2508.15503) — is used as the §4.5 cross-check baseline. Given how fast this literature moves, the exception is defensible, but the asymmetry should be acknowledged in one sentence.","section":"Appendix A.3 (E2), §4.5"},{"comment":"Candidate-study selection was performed by a single researcher (A.5), acknowledged as a limitation in A.7. A second-screener audit of even a small sample of excluded candidates, or a reported agreement check, would materially strengthen the RR at low cost.","section":"Appendix A.5, A.7"},{"comment":"Searches closed 29/03/2025. The narrative tracking of later guidance (Baltes et al., Farotimi et al., Nahar et al.; §3.2) is a reasonable substitute, but the justification currently appears only in A.7; please state the search horizon and the tracking strategy briefly in §3.2 where the RR is first summarized.","section":"§3.2, Appendix A.7"},{"comment":"Among the dropped RAISE items, 'reviewers need to be aware that research papers may be written by AI' is dismissed as having unclear implications. Since AI-authored primary studies would directly threaten SLR validity (screening, RoB), one sentence explaining why this is out of scope — or a pointer to §7 as an open question — would close the gap.","section":"§4.5"},{"comment":"Table 3's qualitative-synthesis row could cross-reference the stated limitation of [79] in Appendix C.4.2 (no independent data-extraction stage; full papers given to the tool), so readers see the caveat where the failure modes are summarized, not only in the appendix.","section":"Table 3, Appendix C.4.2"},{"comment":"Typographical: 'chatbotprompts' (§2.1, item 6); 'Its search processes conforms' (A.2); 'GenAi tool performance' (B.1.3); 'usingvsevaluating' (Table 1, missing spaces); reference [2] carries an '[n. d.]' date; 'SLRS tasks' (B.2.1).","section":"Various"},{"comment":"Figure 2 is information-dense; the 'validate & refine prompts' / 're-validate on new issues' feedback arrows are easy to miss. Consider slightly larger labels or a caption sentence noting the iterative loop, which §B.2.5 describes well in prose.","section":"Figure 2, §B.2.5"}],"recommendation":"minor_revision","confidential_remarks":"The empirical backbone leans on two of the authors' own companion studies ([61], [79]); this is disclosed prominently (§2.3, §4.3) and both are peer-reviewed, so I do not regard it as problematic, but the editor may wish to note the citation pattern. The literature searches closed 29/03/2025, well over a year before submission; the authors track later developments narratively and their argument that re-running the searches would not change the recommendations is plausible for a guidelines paper, but it is worth an editor's awareness. Fit with the journal seems good: this is methodology guidance for SE secondary studies, complementary to (and explicitly positioned against) RAISE and Baltes et al. rather than duplicative of them."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline: this is practical infrastructure for empirical SE, not a new theory of GenAI. Kitchenham et al. package GUEST—role-split (evaluator vs reviewer), planning/conduct/reporting rules, and task-level advice—so people stop treating chatbot summaries as systematic reviews and stop ranking screeners on accuracy under severe imbalance.\n\nWhat is actually new is the SE-specific assembly. Table 1 positions cleanly against RAISE, Farotimi, Baltes, and Wagner. The screening stance is the sharpest part: MCC/WMCC, full confusion matrices, and the reanalysis where metric choice swings missed relevant studies from ~63% down to ~5.8%. That lands, and it is the piece I would actually use. The failure-mode map (Table 3), the process extensions in the appendices, and the explicit “GenAI as extra validator, not unsupervised synthesizer” line for qualitative/RoB/SoE work are clear and conservative in the right way. They also list the open questions in §7 instead of papering over them.\n\nSoft spots, in proportion. The empirical backbone is two companion papers (screening + qualitative synthesis); extraction, RoB, and SoE rest on thought experiments and briefing notes, which the authors admit. The §4.5 check is only partly external—Baltes extends Wagner, and RAISE was a development input—so treat it as consistency, not independent proof. The rapid review is single-source SCOPUS with a March 2025 cutoff and one screener; fine for theme coverage, not for claiming completeness. None of that sinks the paper: the claims are labelled preliminary, the central oversight argument holds, and the recommendations mostly restate good process control under GenAI constraints.\n\nWho it is for: SE groups running or reviewing GenAI-assisted SLRs, and editors tired of contaminated benchmarks and vanity metrics. Math/data load is light by design (process guidance). Citation pattern is appropriate; self-cites to the companions are load-bearing inputs, not circular definitions.\n\nI would engage with it, cite the screening metric advice, and send it to referees as preliminary guidelines. Expect them to push provisional language on under-evidenced tasks—that is fair, not a desk reject.","headline":"Solid preliminary SE process guidance for GenAI-in-SLRs; useful packaging and honest limits, not a fully evidenced rule set for every SLR task.","tokens_in":44423,"tokens_out":549,"would_cite":true,"duration_ms":9778,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"GenAI cannot run unsupervised systematic literature reviews; GUEST process rules keep humans in charge while using it for help and checks.","keywords":["Evaluation studies","Large Language Models (LLMs)","Generative AI (GenAI) tools","Systematic Review","GenAI tool use","GenAI tool evaluation","Software Engineering (SE)","Guidelines"],"falsifier":"A prospective, contamination-controlled comparison on a full software-engineering SLR pipeline in which a GenAI workflow without continuous human validation matches dual-human process controls on validity, traceability, evidence qualification, and auditability—or clearly fails those same checks under the paper’s own metrics.","tokens_in":44186,"feed_emoji":"📋","tokens_out":854,"duration_ms":15356,"temperature":0.7,"pith_summary":"Systematic literature reviews need validity, traceability, qualification of evidence strength, and verifiability that plain expert summaries do not provide. Generative AI tools can summarize text and may cut cost on repetitive work, but they also hallucinate, give inconsistent answers, leak or contaminate test data, hide their reasoning, and can inherit bias. This paper argues they therefore cannot replace human researchers on full systematic reviews today. It offers GUEST: concrete process recommendations for two roles—people who evaluate GenAI on review tasks, and people who run reviews with GenAI help—covering planning, conduct, reporting, and task-level rules for screening, extraction, qualitative synthesis, risk of bias, and strength of evidence. The aim is trustworthy GenAI-assisted reviews and rigorous independent evaluations rather than unsupervised automation.","feed_headline":"GenAI cannot run unsupervised systematic reviews","feed_subtitle":"GUEST process rules keep humans accountable while allowing help on repetitive SLR tasks","key_machinery":"GUEST (GenAI Use and Evaluation in SLR Tasks): role-split process recommendations (evaluators vs reviewers) for planning, conduct, and reporting, plus task-level rules—especially cost-sensitive screening metrics that treat lost relevant studies as worse than extra work, and treating GenAI as a validator rather than sole agent on qualitative synthesis, risk-of-bias, and strength-of-evidence work.","core_discovery":"GenAI requires human oversight and is not currently capable of unsupervised systematic studies; it can give cost-effective help on some repetitive tasks and extra validation on some complex ones. The authors package that stance as GUEST process recommendations so software engineering researchers can both run and report trustworthy SLRs that use GenAI and produce rigorous independent evaluations of GenAI on SLR tasks.","pith_inferences":["Journals and conferences that require GUEST-style disclosure will make GenAI-assisted SLR claims easier to audit and meta-analyze.","The same human-oversight and contamination logic likely extends to other secondary-research genres that borrow SLR process controls.","Until SE-specific risk-of-bias and strength-of-evidence instruments stabilize, GenAI scores on those tasks will stay hard to compare across studies.","Multi-tool ensembles may cut idiosyncratic errors but do not remove shared training-data contamination and raise reported cost."],"forward_implications":["SLR reports that use GenAI must document prompts, model versions, parameters, human checks, and full confusion-matrix counts where screening is automated.","Screening evaluations should prefer metrics that penalize missed relevant studies over raw accuracy or unweighted scores under class imbalance.","On qualitative synthesis, risk-of-bias, and strength-of-evidence tasks, GenAI should run as an extra checker beside humans, not as the sole decision maker, until stronger evidence exists.","Prospective case studies built into live SLRs become the preferred way to gather evaluation evidence and limit data contamination.","End-to-end agentic review platforms remain open research unless each stage’s intermediate outputs can be inspected and validated."],"fun_headline_variants":["GenAI needs human oversight for systematic reviews","GUEST rules: GenAI helps SLRs but never unsupervised","GenAI cannot run SLRs alone—humans stay accountable","GenAI aids repetitive SLR tasks under human control","GUEST keeps GenAI assistive, not autonomous, in SLRs"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"Role-based thought experiments, a single-library rapid review with one screener, and two companion studies are enough to ground general process rules for every major SLR task, including those where the paper itself says direct evidence is still thin or contradictory.","fun_headline_variants_meta":{"raw":{"variants":["GenAI needs human oversight for systematic reviews","GUEST rules: GenAI helps SLRs but never unsupervised","GenAI cannot run SLRs alone—humans stay accountable","GenAI aids repetitive SLR tasks under human control","GUEST keeps GenAI assistive, not autonomous, in SLRs"]},"model":"grok-4.5","effort":"low","cost_usd":0.004292,"raw_usage":{"total_tokens":1301,"prompt_tokens":824,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":42924000,"prompt_tokens_details":{"text_tokens":824,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":412,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":824,"tokens_out":65,"duration_ms":7302,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T04:10:00.352137+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A prospective, contamination-controlled comparison on a full software-engineering SLR pipeline in which a GenAI workflow without continuous human validation matches dual-human process controls on validity, traceability, evidence qualification, and auditability—or clearly fails those same checks under the paper’s own metrics.","supporting_citations":[],"review_version":1}