{"id":"6c70c970-c825-402d-aa32-9028df32b715","arxiv_id":"2607.25489","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review of 557 studies finds medical agentic AI is technically promising but overwhelmingly validated on benchmarks, simulations, and retrospective data rather than in real clinical workflows.","lead":"A medical-AI review maps 557 studies of \"agentic\" systems that plan, use tools, and coordinate multiple models. It finds the evidence base is dominated by benchmarks, simulations, and retrospective data rather than real clinical workflows, and lays out an evaluation agenda for clinical translation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that the evidence base is dominated by low-maturity validation is asserted but never quantified: no counts of the 557 studies by validation level or evaluation dimension appear in the paper.","rationale":"The reader's weakest assumption focuses on the completeness and correctness of the 557-study evidence set (Europe PMC screening, single-author extraction, accurate agentic classification). Those are legitimate process concerns, and the paper itself acknowledges them. However, I identify a more directly load-bearing issue: the paper never actually presents the distributional synthesis that its central claim depends on. The review claims to perform systematic evidence mapping, but no quantitative summary of the map appears in the main text or supplement. Figure 4 is explicitly representative, not comprehensive, and Table 1 only lists review stages. Therefore, even if every one of the reader's completeness concerns were resolved, the central assertion would remain unsupported by the manuscript. This does not destroy the likely truth of the conclusion—prior narrower reviews and the inherent recency of the field make the 'emerging, not mature' framing credible—but it makes the empirical basis uncheckable. The appropriate response is the same as the reader's: request the auditable data and the distributional evidence before treating the claim as established. Hence I do not change the verdict; I agree partially because the reader also noted missing extraction data but did not make it the central assumption.","tokens_in":24129,"tokens_out":6879,"duration_ms":72724,"concrete_test":"Request the authors' full extraction dataset: one row per included study with columns for study ID, validation level (one of the five in §2.6), architecture, and evaluation dimensions. Independently compute counts and percentages per validation level and per evaluation dimension. The central claim is substantiated if, for example, the combined proportion at 'static benchmark' plus 'agent benchmark' is clearly the majority and 'prospective workflow' is near zero. If the distribution shows a different pattern, the 'dominated by' claim must be revised. This single check settles whether the absence of reported distributional evidence hides a meaningful risk.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's central empirical premise is a distributional claim: the evidence base for medical agentic AI is 'dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation.' Yet the manuscript never reports the marginal distribution of the 557 included studies across the five validation levels defined in §2.6 (static benchmark, agent benchmark, retrospective real data, external evaluation, prospective workflow), nor across the six evaluation dimensions of §3.7. Figure 4 plots only a handful of 'representative' systems and explicitly disclaims ranking; no table or supplementary file lists the included studies with their assigned classifications. Consequently, the 'dominated by' assertion is unfalsifiable from the paper alone. The completeness concerns raised in §4.4 (unscreened Europe PMC records, single-author screening) are secondary: even if the 557-set were complete and accurately classified, the reader cannot confirm the distributional claim that drives the conclusion 'emerging technical paradigm rather than mature clinical solution.' This is load-bearing because the paper's whole translational message is that research priorities should shift away from benchmark capability toward prospective validation—a recommendation that rests on the claimed distribution. If the actual distribution were different (e.g., a substantial share of retrospective or external evaluations), the strength of that message would change materially.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This scoping review with systematic evidence mapping characterizes the emerging field of agentic AI in medicine. After searching five electronic sources, the authors screened 1,649 exportable records and provisionally included 557 unique studies. The paper proposes a taxonomy of four architectural families (single-agent tool-use, multi-agent collaboration, knowledge-augmented, multimodal medical), defines a five-level clinical-validation maturity scale, and maps representative systems by agentic complexity and validation level. The central claim is that the evidence base is dominated by public benchmarks, simulated settings, retrospective data, and small-scale expert evaluation, and therefore medical agentic AI should be regarded as an emerging technical paradigm rather than a mature clinical solution. The discussion and conclusions call for prospective, workflow-integrated validation, structured evidence traceability, and clearer reporting standards.","tokens_in":24444,"tokens_out":2605,"duration_ms":31692,"significance":"If the distributional claim is correct, the paper provides a useful and timely synthesis that could help align research priorities with clinical translation needs. The review is explicitly cautious: the agentic complexity score (0–8) is repeatedly disclaimed as descriptive rather than validated, and the validation-maturity categories are presented as an analytical framing rather than a measured scale. The inclusion of formal concepts such as the POMDP formulation (Eq. 1), expected calibration error (Eq. 2), and decision curve analysis (Eq. 3) strengthens the conceptual grounding. The paper also gives credit to prior reviews and distinguishes enabling components (e.g., vision–language models, segmentation tools) from complete agentic workflows. However, the central empirical premise—the distribution of evidence across validation levels—is never quantified from the 557-study set, and the review's reproducibility is limited by unscreened Europe PMC records and single-author screening. These issues are acknowledged in the text but are load-bearing for the main translational message.","major_comments":[{"comment":"The central claim that the evidence base is 'dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation' is never supported by counts of the 557 included studies across the five validation levels defined in §2.6 or the six evaluation dimensions of §3.7. Figure 4 plots only a handful of representative systems and explicitly disclaims ranking; no table or supplementary file lists the included studies with their assigned classifications. As a result, the distributional premise that drives the conclusion 'emerging technical paradigm rather than a mature clinical solution' is not auditable from the paper alone. Please add a study-level evidence map or at least aggregate counts by validation level and evaluation dimension.","section":"§3.1, §3.2, Fig. 4"},{"comment":"The unscreened Europe PMC records (n=445) are acknowledged as a limitation, but their potential impact on the 557-study set and on the distributional claim is not assessed. If these records disproportionately contain retrospective or benchmark studies, the 'dominated by' conclusion might be robust; if they disproportionately contain more mature evaluations, the conclusion could weaken. Please report the source-specific record counts, and either provide a sensitivity analysis based on title-level screening of Europe PMC records or justify why overlap with PubMed makes the unscreened set negligible.","section":"§2.4, §4.4"},{"comment":"Study screening and data extraction were performed by a single author. This creates a risk of systematic misclassification of validation level and agentic complexity, which directly affects the evidence map. Even a small random sample independently screened by a second reviewer, with inter-rater agreement reported, would substantially improve confidence in the classification. The current description does not include any reliability check.","section":"Author Contributions, §2.6"},{"comment":"Eligibility and classification rely on each primary study's self-reported workflow. The paper notes that under-reporting may influence assigned categories, but it does not state how missing information was handled (e.g., whether authors were contacted, or whether a 'not reported' category was used). This matters because a study that omits details of memory or feedback mechanisms could be misclassified as less agentic, biasing the complexity–validation map. Please specify the coding rules for missing or ambiguous information and report how many studies were affected.","section":"§2.3, §4.4"}],"minor_comments":[{"comment":"CXR-Agent is included in the figure as a contextual example but was explicitly not part of the formal evidence-mapping set (§3.6.1). To avoid confusion, the figure caption should state this more prominently, or the point should be visually distinguished from the included studies.","section":"Fig. 4"},{"comment":"Table 1 refers to a 'five-stage validation setting' while the text (§2.6) describes 'five ordered levels.' Use consistent terminology to prevent ambiguity.","section":"Table 1"},{"comment":"In the ECE definition, the notation acc(B_m) and conf(B_m) is defined, but the formula would benefit from a brief note that the bins are usually equal-width confidence intervals. This is implied by the reference to binning, but it would aid readers not familiar with the literature.","section":"Eq. (2)"},{"comment":"Decision curve analysis: the equation uses TP(pt) and FP(pt) as counts, which is clear, but the sentence 'The term pt/(1−pt) represents the relative harm...' is a bit terse. Consider adding one sentence that the net benefit is plotted against threshold probability and compared with 'treat all' and 'treat none' strategies, as stated later.","section":"Eq. (3)"},{"comment":"The data availability statement says 'No new datasets were generated or analyzed.' However, the systematic evidence-mapping set (557 studies with assigned classifications) is a new dataset that should be made available as a supplementary table to support reproducibility and future updates.","section":"Data availability"},{"comment":"The search syntax is described conceptually but the exact queries for each source are not reported. For a reproducible scoping review, provide the full search strings for at least one representative database, or include them as an appendix.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's conceptual framework and cautious interpretation are likely sound and would be a useful contribution to the medical AI literature. The main issue is that the systematic evidence-mapping claim—the backbone of the translational message—is not quantitatively supported in the manuscript. Adding study-level classifications or aggregate counts, a sensitivity note on the Europe PMC records, and a reliability check for single-author screening should be feasible within the scope of a revision. If the authors can provide those, the paper would become substantially more rigorous and impactful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a genuinely useful consolidation, but the paper sells a systematic evidence map it never actually shows. The central claim—that the evidence base is dominated by benchmarks, simulations, retrospective data, and small-scale expert review—is repeated in the abstract, results, and conclusion, but nowhere do the authors report the distribution of their 557 included studies across the five validation levels or the six evaluation dimensions they define. That is a load-bearing omission: if, say, a third of studies had already reached external or prospective evaluation, the 'emerging technical paradigm' message would need substantial qualification. Figure 4 plots only a handful of named systems and explicitly disclaims ranking. No supplementary table lists the included studies with their assigned classifications.\n\nWhat's good: the architectural taxonomy (single-agent tool-use, multi-agent, knowledge-augmented, multimodal) is clear enough to use, and the six-dimension evaluation framework plus the five-level validation ladder are sensible organizing devices. The paper carefully separates model-level capability from agentic behavior, and its final recommendations—more prospective, workflow-integrated validation; traceability of tool calls and retrieved evidence; staged reporting guidelines—are reasonable. The authors also openly acknowledge some of the process gaps: Europe PMC records were never exported or screened, screening and extraction were done by a single author, and they disclaim the agentic-complexity score as unvalidated. Those are real limitations but they are stated, not hidden.\n\nThe soft spot is that the systematic evidence-mapping claim goes beyond what the manuscript supports. There is no search syntax, no PRISMA-ScR checklist, no included-study list, and no extraction data. Even setting aside completeness, the reader cannot check the one empirical claim that carries the review's argument. A referee should require either the underlying classification tables or a softening of the 'dominated by' language to reflect an impressionistic reading.\n\nFor whom: researchers and funders looking for a structured orientation to medical agentic AI will get real value from the taxonomy and evaluation checklist. It is a good starting point, not a definitive evidence synthesis.\n\nMy recommendation: send it to peer review, but with the explicit requirement that the evidence map be made auditable before publication. The framework deserves to be in the literature; the distributional claim needs to earn its keep.","headline":"A useful synthesis whose central distributional claim is asserted, not shown—worth refereeing, but only after the evidence map becomes auditable.","tokens_in":24880,"tokens_out":2518,"would_cite":true,"duration_ms":25906,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scoping review of 557 studies concludes that medical agentic AI is an emerging technical paradigm rather than a mature clinical solution, because most current evidence comes from benchmarks, simulations, retrospective data, and small expe","keywords":["agentic AI","medical AI","large language models","multi-agent systems","clinical validation","evidence mapping","scoping review","clinical translation"],"falsifier":"Re-run the selection with the 445 unscreened Europe PMC records included and check whether the distribution of validation levels shifts; if a meaningful share of those studies had prospective or external clinical validation, the conclusion that the evidence base is dominated by low-maturity evaluations would weaken. A second check: if a multi-institutional prospective study of an agentic medical system demonstrated clear workflow benefit and safety, the claim that agentic AI is not a mature clinical solution would need to be updated.","tokens_in":23995,"feed_emoji":"🏥","tokens_out":5576,"duration_ms":50274,"temperature":0.7,"pith_summary":"Medical agentic AI systems — which plan, use tools, retrieve knowledge, remember context, and coordinate specialized agents around a clinical goal — are often discussed as if they were ready for the clinic. This review tries to establish what the evidence actually supports. It screened 1,649 exportable records and provisionally mapped 557 studies, classifying them by architecture, application, and how close each evaluation comes to real clinical use. The central finding is that most studies sit at low validation maturity: they rely on public benchmarks, simulated environments, retrospective datasets, or small-scale expert judgment, while process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are measured inconsistently. The sympathetic reading is that the field shows genuine technical promise, especially in medical imaging, but should be treated as an emerging research paradigm rather than a deployable clinical solution.","feed_headline":"557 studies: medical agentic AI is an emerging idea, not a mature tool","feed_subtitle":"Benchmark and simulated evaluations dominate; prospective validation is scarce, so research priorities must shift.","key_machinery":"The analytical engine is the evidence map itself. Eligibility was defined functionally — a system counted as agentic only if it showed goal-directed multistep execution plus at least one explicit agentic mechanism such as planning, tool use, environment interaction, memory, feedback-based refinement, or multi-agent collaboration — rather than by whether it used the word 'agent.' Each included study was then classified along two axes: an agentic complexity score (0–8, counting planning, tool use, retrieval, memory, reflection, multi-agent collaboration, multimodal processing, and workflow integration) and a five-level clinical validation maturity scale (static benchmark, agent benchmark, retr","core_discovery":"The paper's central claim is that current medical agentic AI is better understood as a system-level orchestration paradigm than as a mature clinical technology. On its own terms, the review demonstrates this by mapping 557 studies across four architectural families — single-agent tool use, multi-agent collaboration, knowledge-augmented agents, and multimodal medical agents — and ranking each study's validation setting on a five-level maturity ladder from static benchmarks to prospective workflows. The decisive observation is that the evidence concentrates in the lower rungs: public datasets, simulated environments, retrospective records, and small expert evaluations dominate, while prospecti","pith_inferences":["The mapping implies a sharper regulatory hypothesis than the authors state: if benchmark-level evidence cannot support clinical claims, then an agentic system's 'indication' should be defined by its workflow and oversight boundaries, not by the underlying model's benchmark score.","The finding that multi-agent benchmarks do not show consistent superiority over strong single-model baselines suggests that the field may currently over-invest in multi-agent collaboration; a testable extension is to compare single-agent and multi-agent versions of the same task with matched compute and measurement of error propagation.","A practical extension of the five-level maturity scale would be to require a minimal reporting checklist — tool invocations, retrieved sources, abstention decisions, clinician overrides, failure cases — before a study can be assigned to a validation level; this would make future evidence maps more reproducible.","As simulated EHR environments and virtual hospitals become more realistic, they could serve as an inexpensive screening stage that selects only the most promising agentic systems for prospective clinical studies, reducing the cost and risk of clinical translation."],"forward_implications":["Public-benchmark accuracy should no longer be treated as evidence of clinical readiness; evaluation must also report tool-selection reliability, code-execution success, retrieval relevance, error recovery, abstention and escalation behavior, and human revision burden.","Because the four architecture families are not mutually exclusive and many systems combine them, the reliability of an agentic system depends on the whole execution chain, not on the underlying foundation model alone.","In medical imaging, future systems must demonstrate visual grounding and cross-modal consistency — generated text must be traceable to identifiable image findings — and be validated across institutions and acquisition protocols.","Clinical translation should proceed through staged evidence generation: external retrospective validation, prospective silent testing, controlled workflow studies, and post-deployment surveillance, with reporting guided by study-design-appropriate standards.","Current medical agentic systems should be designed as supervised workflows with bounded action spaces, evidence provenance records, uncertainty-based abstention, and mandatory clinician review, not as autonomous decision-makers."],"fun_headline_variants":["Medical agentic AI: 557 studies, mostly benchmarks and sims","Agentic AI in medicine: 557 studies, clinical proof scarce","Agentic AI in medicine: 557 studies, still an emerging idea","Medical agentic AI: 557 studies, but real-world evidence thin","Agentic AI in medicine: 557 studies, few validated in clinics"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The review's map is only as reliable as its assumption that the 557 included studies were correctly identified and classified — which depends on the 445 Europe PMC records that could not be exported or screened not biasing the evidence pool, and on screening and extraction performed largely by a single author being accurate.","fun_headline_variants_meta":{"raw":{"variants":["Medical agentic AI: 557 studies, mostly benchmarks and sims","Agentic AI in medicine: 557 studies, clinical proof scarce","Agentic AI in medicine: 557 studies, still an emerging idea","Medical agentic AI: 557 studies, but real-world evidence thin","Agentic AI in medicine: 557 studies, few validated in clinics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001075,"raw_usage":{"total_tokens":4333,"prompt_tokens":736,"completion_tokens":3597,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":3516}},"tokens_in":480,"tokens_out":3597,"duration_ms":24333,"temperature":1.0,"reasoning_tokens":3516,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:15:26.337868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the selection with the 445 unscreened Europe PMC records included and check whether the distribution of validation levels shifts; if a meaningful share of those studies had prospective or external clinical validation, the conclusion that the evidence base is dominated by low-maturity evaluations would weaken. A second check: if a multi-institutional prospective study of an agentic medical system demonstrated clear workflow benefit and safety, the claim that agentic AI is not a mature clinical solution would need to be updated.","supporting_citations":[],"review_version":1}