{"id":"89c734c8-74de-4048-8eec-840e5f05e43e","arxiv_id":"2504.14822","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An interactive, multi-agent LLM-based system with visual oversight lets one clinician complete a systematic review in about 1.5 hours while reaching about 80% of human-written review quality.","lead":"InsightAgent is an interactive AI system that helps clinicians complete a systematic review in about 1.5 hours instead of months, using a visual map of the literature and multiple agents that read and summarize papers under real-time human guidance. In a study with 9 medical professionals, its output reached roughly 80% of human-written review quality, and it found relevant studies more accurately than two fully automatic baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central measurement problem: the 79.7% figure is an extended-abstract rubric score, not systematic-review quality; the paper's own Limitations state that full-text reading, risk-of-bias assessment, and quantitative synthesis are absent, so reports can score ~80/100 while failing core PRISMA…","rationale":"I read the paper as a systems contribution: a multi-agent pipeline with semantic partitioning, an RSS map, provenance tracking, and three interaction modes. The design is coherent, and the release of code and data is a meaningful plus. The reader's CONDITIONAL verdict remains appropriate, but my stress-test pushes the measurement problem one level deeper than inter-rater reliability. Even a perfectly reliable score on the Appendix E rubric would not establish systematic-review quality, because the rubric was designed for extended abstracts and the human reference is an abstract. The paper's own Limitations section states that full-text reading, quantitative/statistical synthesis, and evidence weighting are absent; these are basic requirements for a high-quality systematic review in the biomedical domain. Thus the central assertion of completing a high-quality systematic review in 1.5 hours is not what was actually measured. The 27.2% interaction gain may still be a real and useful effect for abstract-level evidence synthesis, so I am not recommending rejection. I would keep the conditional verdict and add a condition: evaluate against a full systematic-review standard, including PRISMA-informed criteria, full-text screening, risk-of-bias assessment, and quantitative synthesis where appropriate, or explicitly reframe the claims as being about abstract-level reviews. Blinding and inter-rater reliability remain important secondary conditions, but the construct-validity issue is more load-bearing because it questions whether the headline quantity was measured at all.","tokens_in":28335,"tokens_out":10335,"duration_ms":98410,"concrete_test":"Score the 15 generated reports (or, as a minimal check, the Appendix H remdesivir report) against the PRISMA 2020 checklist, focusing on items that require full-text retrieval, risk-of-bias assessment, and quantitative synthesis, such as PRISMA items 8, 13, 21, and 22. Have two methodologists who did not design the paper's rubric rate the same reports with both the Appendix E rubric and the PRISMA checklist. If reports score roughly 80/100 on the paper rubric while satisfying only a small fraction of the PRISMA items, the 79.7% figure is an artifact of the abstract-oriented rubric and the claim should be reframed as 'high-quality extended abstract.' If the reports actually satisfy the relevant PRISMA items, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the headline claim, the rubric in Section 4.1/Appendix E would have to measure systematic-review quality. It does not. Appendix E explicitly tailors the rubric to 'agent-based summarization of a biomedical systematic review in extended-abstract format,' and the human reference in the worked example (Appendix H, Table 14) is the original published abstract, not a completed systematic review. Section 4.1 also says peer-reviewed human systematic reviews are treated as 'ground truth (100 points)' without showing that the same full-text human reviews were scored with this rubric. The Limitations section then concedes that only titles/abstracts are read, that the system cannot extract or synthesize numerical statistics or effect sizes, cannot perform meta-analysis, and does not weight evidence by study design. Each of these is a core component of a high-quality healthcare systematic review. Consequently, the 27.2% interaction improvement and the 79.7% 'of human-written quality' figure can be valid for abstract-level summaries while still not supporting the claim that a clinician can 'complete a high-quality systematic review in about 1.5 hours.' The load-bearing condition that fails is construct validity: the measured outcome is not the claimed outcome.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents InsightAgent, a human-in-the-loop LLM-agent system for biomedical literature review. InsightAgent maps a retrieved corpus into a radial relevance-and-similarity layout (RSS map), partitions it with K-means, dispatches parallel agents to screen titles/abstracts and incrementally synthesize evidence, and exposes a provenance/dependency tree so users can verify claims. Users can intervene through path navigation (dragging an agent to an article), chat, and direct instruction, with the agent entering a reflection phase after each intervention. The evaluation covers 15 published systematic reviews and 9 medical experts, reporting that InsightAgent (GPT-4o) reaches 98.5% recall and 88.2% F1 in record screening; that interaction improves generated-review quality by 27.2% (p = 3.43e-7) and user satisfaction by 34.4% (p = 1.89e-6); and that the interactive system scores 79.7 out of 100 rubric points, characterized as 79.7% of human-written quality, in sessions averaging about 1.5 hours.","tokens_in":28506,"tokens_out":15741,"duration_ms":124120,"significance":"The system contribution is substantial: the multi-agent partitioning is compared against a single-agent variant (Appendix B); the provenance-tree mechanism directly addresses the untraceable-summary problem identified for ChatCite and AutoSurvey; code and data are publicly released; and the evaluation against 15 real published reviews with domain experts is far more grounded than typical LLM-summarization evaluations. The screening results, particularly the interactive 98.5% recall, are striking, and Appendix I's error taxonomy (quantitative evidence, evidence weighting, heterogeneity, faithfulness) is a useful analytic contribution. However, the significance as currently framed depends on the claim that the system completes a high-quality systematic review, and the measured outcome — an extended-abstract rubric score anchored on human abstracts — does not support that construct. With a re-scoped claim (abstract-level evidence synthesis) and additional measurement reporting (inter-rater reliability, corpus-reconstruction validation), the paper would be a strong contribution to human-centered AI for evidence synthesis.","major_comments":[{"comment":"The abstract's central claim that a clinician can 'complete a high-quality systematic review' in about 1.5 hours at 79.7% of human-written quality is not supported by the paper's own measurement instrument. Appendix E states that the rubric is tailored to 'agent-based summarization of a biomedical systematic review in extended-abstract format'; Category 2 item 2.1 ('Agent-Based Summarization Description') is definitionally inapplicable to a human-written review; and the human reference used in the worked example (Appendix H, Table 14) is the published abstract of Tan-Lim and Esteban-Ipac, not the completed systematic review. The Limitations section concedes that only titles and abstracts are read, that the system cannot extract or synthesize numerical statistics or effect sizes, cannot perform meta-analysis, and does not weight evidence by study design — each a core component of a high-quality healthcare systematic review. The 79.7% figure is therefore a quality score for extended abstracts against an abstract-level anchor, and neither the abstract's nor Section 5's 'complete a high-quality systematic review' claim follows from it. The authors should either re-frame the headline claims to abstract-level evidence synthesis support, or validate the rubric against full systematic reviews and report scores for the human references on the same rubric.","section":"Section 4.1; Appendix E; Appendix H; Limitations"},{"comment":"The statistical claims rest on an unreported measurement foundation. Each review is 'independently scored by two experts,' but the paper reports no inter-rater reliability (Cohen's kappa or per-item agreement) and no rater calibration; the paired t-test p-values (3.43e-7 for quality, 1.89e-6 for satisfaction) cannot be assessed without knowing the agreement structure or whether the paired observations are per-review, per-rater, or per-session. The same nine users who operated the systems also rated the outputs, with no stated blinding to system identity or to the auto/interactive condition, and no reported counterbalancing of presentation order. Additionally, Table 3 reports interface-usability Likert items (e.g., drag-and-drop, citation tracing) for the InsightAgentauto condition, in which the interactive interface is not in use; it is unclear what participants were rating in that condition. Please report inter-rater statistics, the blinding and counterbalancing protocol (or explicitly acknowledge its absence), and a precise description of the autonomous-condition questionnaire.","section":"Section 4.1; Section 4.2.3; Table 3"},{"comment":"The reconstructed corpora are not validated against the published search strategies. Appendix D states that each corpus was replicated from PubMed, but the 15 reviews include Cochrane reviews (e.g., Hyun et al., Wu et al., Sharrad et al.) whose published searches cover multiple databases and trial registries; a PubMed-only reconstruction will likely omit records those reviews screened. The paper does not report what fraction of each published review's included studies is present in the reconstructed corpus, and Table 6 shows inclusion rates as low as 0.37% (Wu et al.), so even a small reconstruction gap can cap the achievable recall. If the ground-truth included set is not fully present, the reported 98.5% interactive recall is an artifact of corpus construction rather than a measure of screening ability; if the corpus contains extra irrelevant records, reported precision reflects reconstruction noise. Please report, per review, the coverage of the published included set and a sensitivity analysis of the screening metrics to reconstruction completeness.","section":"Section 4.1; Appendix D; Table 6"},{"comment":"The '1.5 hours versus months' comparison conflates session length with the duration of the claimed workflow. Section 4.1 reports that participants finished a session in 1.5 hours while 'no time limit' was set, and the session produces an extended abstract from a pre-retrieved title/abstract corpus. The months-long human pipeline being compared against includes full-text screening, dual independent review, risk-of-bias assessment, and quantitative synthesis, all of which the Limitations state the system does not perform. The time comparison should be scoped to the synthesis stage on a provided corpus, and the measured time should be decomposed into corpus preparation, screening, interaction, and report generation, with the comparison made against the corresponding stages of the manual process.","section":"Section 4.1; Conclusion; Abstract"},{"comment":"The units of analysis in the interaction and error analyses are not derivable from the stated study design. Section 4.2.4 says 50 interim-synthesis pairs were sampled per interaction type (150 pairs total), but the study involves 9 users and 15 reviews, and it is unclear where the 150 pairs come from, who scored them with the rubric, and how the before/after scores in Figure 4 relate to the two-expert rubric used in Table 2. Likewise, Table 5 reports R=59 for InsightAgent (GPT-4o) and R=30 for InsightAgent (Llama-3.3), while the stated design (15 reviews × 2 raters) yields 30 reports per condition; the provenance of R=59 should be explained so that the error rates in Appendix I are auditable.","section":"Section 4.2.4, Figure 4; Table 5"}],"minor_comments":[{"comment":"The sentence reporting precision as '(62.4%/41.9% vs 20.4%)' uses 41.9%, which matches no entry in Table 1; the GPT-4o autonomous precision is 51.9%.","section":"Section 4.2.1 (text near Table 1)"},{"comment":"The sentence 'We report the details of our interface design in appendix 6' should reference Appendix C instead.","section":"Section 3.2"},{"comment":"The sentence 'We provide an example final synthesis template in Appendix A' is inaccurate: Appendix A contains the four operational prompts (Retrieve, Read, Synthesize, Reflect), and the final-synthesis template appears only inside the Synthesize prompt; please correct the cross-reference.","section":"Section 3.3"},{"comment":"The caption reads 'LLamA 3.3 70B' and should read 'Llama 3.3 70B'.","section":"Figure 3 caption"},{"comment":"The sentence 'To make is suitable for our evaluation setting' contains a typo; it should read 'To make it suitable.'","section":"Appendix E"},{"comment":"The InsightAgent example report states that 11 studies met the inclusion criteria while the human review included 9; this discrepancy should be discussed as an inclusion-fidelity issue rather than left to the evaluator's comment in Table 19.","section":"Table 15 vs. Table 14"},{"comment":"The claim of a '47% (F1 points)' improvement in article identification is misleading: the ratio 88.2/60.0 is a 47% relative increase in F1, not 47 F1 points (the absolute gain is 28.2 points); please state the metric unambiguously.","section":"Introduction; Section 4.2.2"},{"comment":"The statement that InsightAgent is 'the first system to be practically useful for domain experts' is a qualitative judgment not established by the reported experiments; please remove it or support it with the scoped evidence.","section":"Section 4.2.2"}],"recommendation":"major_revision","confidential_remarks":"I agree with the reader's conditional assessment: the measurement issue is the load-bearing weakness, but it is addressable within the manuscript's scope and should not by itself trigger rejection provided the headline claims are re-scoped. I would direct the authors to report Cohen's kappa and a blinding/counterbalancing statement even if the study was in fact unblinded; transparency here would materially change the force of the interaction-improvement claims. The manuscript also sits somewhat between cs.HC and cs.CL; for an HCI journal, the evaluation of the visualization and interaction design (nine users, no kappa) is thin, while for a systems venue the measured outcome should be aligned with the claim. The authors' candid Limitations section and unusually complete appendices are strengths worth acknowledging."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the system is genuinely interesting and the screening results look strong, but the headline claim that a clinician can complete a high-quality systematic review in 1.5 hours is not supported by what they measured. The evaluation scores an extended abstract with a rubric built for that format, not a completed systematic review.\n\nWhat's actually new: they integrate a semantic document map, parallel reading agents, a provenance tree, and live steering into one working tool, and they release code and data. That is a real engineering contribution. The record screening numbers are impressive—98.5% recall and 88.2% F1 with GPT-4o, well above the top-100 baselines. The interaction effect is also large: +27.2% quality score, +34.4% satisfaction, with small p-values. Even accounting for the small sample, that is a plausible signal that human steering helps.\n\nThe soft spots are the usual ones, but there is one load-bearing issue. The rubric (Appendix E) explicitly adapts itself to 'agent-based summarization of a biomedical systematic review in extended-abstract format,' and the human reference in the worked example is the published abstract, not the full review. So the 79.7% figure is '79.7% of a human-written abstract,' not of a full systematic review. That is a genuine construct validity gap, and the title/abstract overstate it. The authors' own Limitations section confirms the system reads only titles/abstracts and cannot do meta-analysis or evidence weighting—so the measured outcome is not the claimed outcome. On the measurement side, nine users and fifteen reviews is small; no inter-rater reliability is reported; the raters were not blinded. These are disclosed or at least not hidden, and they don't kill the screening result, but they weaken the quality-score comparisons.\n\nThe paper is worth engaging. It is well-written, the system is described in enough detail to reproduce, and the authors are honest about limitations. The related work is fair and they cite the systems they build on. The right fix is to recalibrate the language: 'high-quality systematic review' becomes 'a structured, abstract-level synthesis with high screening recall,' and the abstract/headline should match. Also worth asking for raw per-rater scores and inter-rater reliability, and some validation that the reconstructed corpora match the published search strategies.\n\nI'd send this to peer review. The direction is solid, the artifact is useful, and the flaws are fixable in revision, not fatal.","headline":"Impressive screening and a genuinely useful system, but the '79.7% of human quality' headline overclaims—the rubric scores an extended abstract, not a systematic review.","tokens_in":29122,"tokens_out":3678,"would_cite":true,"duration_ms":29836,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A clinician steering an interactive AI agent can complete a systematic review in about 1.5 hours and reach 79.7 percent of human-written quality.","keywords":["systematic review","large language models","interactive AI agents","human-in-the-loop","visual analytics","evidence synthesis","record screening","multi-agent system"],"falsifier":"Run the same 15 reviews through a preregistered, blinded study in which independent raters evaluate system outputs without knowing which system produced them, and report per-item agreement between raters; if the 27.2% interaction gain and the 79.7% figure shrink to noise, the central claim collapses. Separately, compare each reconstructed PubMed search's retrieved set against the published PRISMA flow diagram to check corpus fidelity.","tokens_in":28027,"feed_emoji":"🤖","tokens_out":8517,"duration_ms":69294,"temperature":0.7,"pith_summary":"InsightAgent combines a semantic map of the literature, parallel LLM agents that read and synthesize each cluster, and a visual interface that lets a domain expert steer the agents in real time. The paper's central claim is that this combination lets a single clinician complete a systematic review in roughly 1.5 hours rather than months, reaching 79.7% of human-written quality. In user studies with 9 medical professionals and 15 published biomedical systematic reviews, letting the expert interact with the agent raised generated-review quality by 27.2% (p = 3.43e-7) and user satisfaction by 34.4% (p = 1.89e-6), and interactive screening reached 98.5% recall and 88.2% F1. If true, it would matter because record screening and evidence synthesis are the main bottlenecks of evidence-based medicine, and the result suggests human oversight rather than full automation is the key to usable AI-generated reviews.","feed_headline":"Interactive AI agents finish a systematic review in 1.5 hours","feed_subtitle":"With expert steering, it scores 79.7 of 100 versus human reviews and catches 98.5 percent of relevant studies","key_machinery":"The load-bearing mechanism is the relevance-and-similarity map (RSS map) combined with multi-agent partition and live human steering. The RSS map places articles in a plane where relevance to the research question grows toward the center and semantically similar articles form nearby clusters; K-means then splits the corpus into about nine clusters, each assigned to an agent that reads from the center outward, selecting among the eight nearest neighbors at each step. New evidence is folded into prior memory through the merge $M_{k+1}=f(M_k,S_j)$, and a dependency tree records which article summaries support each interim synthesis. Users can drag agents to missed articles (path navigation), issue natural-language directives (chat navigation), or edit criteria and summarization requirements (instruction navigation); each intervention triggers a reflection pass that updates the reading strategy and reconciles conflicted memory. The provenance tree is what makes the output auditable, and the receptive-field rule is what keeps exploration systematic.","core_discovery":"The discovery is that a human-in-the-loop agent architecture, rather than a fully autonomous LLM pipeline, is what lifts generated systematic reviews toward usable quality. InsightAgent maps the corpus onto a radial relevance-and-similarity layout, partitions it into semantic clusters, and dispatches an agent per cluster; each agent reads from the center outward with a constant receptive field of eight neighbors, stores summaries in a local memory that is merged incrementally, and records every merge in a provenance tree. Expert users can intervene by dragging an agent's path, chatting natural-language directives, or editing inclusion and summarization parameters, after which the agent reflects and reconciles its memory. Across 15 reconstructed biomedical corpora, the interactive GPT-4o variant reached 98.5% recall and 88.2% F1 in screening, scored 79.7/100 on a rubric where human reviews anchor at 100, and improved on its autonomous counterpart by 27.2% in quality and 34.4% in satisfaction; the same result holds qualitatively, with experts reporting more comprehensive, traceable, and sometimes more detailed syntheses than the human-written reference.","pith_inferences":["Because the same experts rated both conditions and were not blind to which output came from InsightAgent, part of the observed 27.2% gain could reflect expectations about interactivity; a blinded replication would test this.","If the effect is real, the interaction benefit may transfer to other high-stakes synthesis tasks — clinical guidelines, policy reviews, regulatory dossiers — wherever a domain expert can spot-check an agent's reading path.","The corpus reconstruction step deserves validation: comparing each reconstructed PubMed search's retrieved set against the original review's PRISMA flow diagram would test whether the 15 corpora are faithful test beds.","A natural next experiment is an ablation separating visualization-only, chat-only, and path-only conditions with a larger expert sample, to identify which interaction channel carries the quality gain."],"forward_implications":["A single domain expert, not a team, can produce a systematic review in about 1.5 hours with quality near human output, sharply lowering the months-long cost of evidence synthesis.","Human interaction is the decisive factor: disabling it drops generated-review quality by 27.2%, so fully autonomous LLM summarizers are the wrong target for this task.","Near-perfect recall (98.5%) in record screening means relevant studies are rarely missed when an expert can correct the agent's path.","Traceable provenance improves expert confidence: users reported that visualizing exploration paths and checking evidence trees made them trust the output.","The system's remaining gaps — abstract-only reading, no pooled statistics or evidence weighting — define the next step toward closing the 20-point gap to human-written quality."],"supporting_citations":[{"why":"Supplies the Cochrane Handbook grounding for multi-reviewer strategies and systematic review standards that motivate the multi-agent design.","marker":"(Chandler et al., 2019)"},{"why":"Provides the RSS map layout that turns the corpus into a relevance-preserving visual space shared by agents and users.","marker":"(Qiu et al., 2024)"},{"why":"ChatCite is one of the fully autonomous baselines whose quality (47.1) provides the lower comparison point.","marker":"(Li et al., 2024)"},{"why":"AutoSurvey defines the fully autonomous baseline (54.0) whose top-100 retrieval and survey generation set the comparison point.","marker":"(Wang et al., 2024)"},{"why":"GPT-4o is the main backbone model for the interactive system and baselines.","marker":"(Hurst et al., 2024)"},{"why":"Llama 3.3 70B supplies the open-weight backbone used to show the design helps beyond one proprietary model.","marker":"(Grattafiori et al., 2024)"},{"why":"BM25 supplies the classical retrieval reference point for record-screening recall and precision.","marker":"(Robertson and Walker, 1994)"},{"why":"PRISMA supplies the reporting framework used to reconstruct the 15 review corpora and shape the evaluation rubric.","marker":"(Page et al., 2021)"},{"why":"One of the 15 evaluation reviews; its human report anchors the qualitative comparison and the 100-point ground truth.","marker":"(Tan-Lim and Esteban-Ipac, 2024)"}],"fun_headline_variants":["Human-guided AI finishes systematic reviews in 1.5 hours","Expert-steered AI hits 98.5% recall on systematic reviews","Interactive AI scores 79.7/100 on systematic review quality","From months to 1.5 hours: human-in-the-loop AI for systematic reviews"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes the 100-point expert rating scale really measures review quality even though the two raters were not blind to which system made each review and the paper never reports whether the raters agreed with each other; it also assumes the 15 rebuilt search results match the original reviews' screening decisions.","fun_headline_variants_meta":{"raw":{"variants":["Human-guided AI finishes systematic reviews in 1.5 hours","Expert-steered AI hits 98.5% recall on systematic reviews","Interactive AI scores 79.7/100 on systematic review quality","From months to 1.5 hours: human-in-the-loop AI for systematic reviews"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000374,"raw_usage":{"total_tokens":2022,"prompt_tokens":997,"completion_tokens":1025,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":946}},"tokens_in":613,"tokens_out":1025,"duration_ms":9319,"temperature":1.0,"reasoning_tokens":946,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:39:04.879197+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 15 reviews through a preregistered, blinded study in which independent raters evaluate system outputs without knowing which system produced them, and report per-item agreement between raters; if the 27.2% interaction gain and the 79.7% figure shrink to noise, the central claim collapses. Separately, compare each reconstructed PubMed search's retrieved set against the published PRISMA flow diagram to check corpus fidelity.","supporting_citations":[],"review_version":1}