Pith. sign in

REVIEW 4 cited by

How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09529 v3 pith:FAF5LGIR submitted 2024-12-12 cs.CV

classification cs.CV
keywords radiologyautomatedllmstaskagentclinicalcomplexcores
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce RadA-BenchPlat, an evaluation platform that benchmarks the performance of large language models (LLMs) act as agent cores in radiology environments using 2,200 radiologist-verified synthetic patient records covering six anatomical regions, five imaging modalities, and 2,200 disease scenarios, resulting in 24,200 question-answer pairs that simulate diverse clinical situations. The platform also defines ten categories of tools for agent-driven task solving and evaluates seven leading LLMs, revealing that while models like Claude-3.7-Sonnet can achieve a 67.1% task completion rate in routine settings, they still struggle with complex task understanding and tool coordination, limiting their capacity to serve as the central core of automated radiology systems. By incorporating four advanced prompt engineering strategies--where prompt-backpropagation and multi-agent collaboration contributed 16.8% and 30.7% improvements, respectively--the performance for complex tasks was enhanced by 48.2% overall. Furthermore, automated tool building was explored to improve robustness, achieving a 65.4% success rate, thereby offering promising insights for the future integration of fully automated radiology applications into clinical practice. All of our code and data are openly available at https://github.com/MAGIC-AI4Med/RadABench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ABRA: Agent Benchmark for Radiology Applications

    cs.CV 2026-05 unverdicted novelty 8.0 of 10

    ABRA shows radiology agents excel at tool execution (89%+) but struggle with outcomes (0-25%), with oracle perception raising outcomes to 69-100%, identifying perception as the primary bottleneck.

  2. MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

    cs.CV 2026-03 conditional novelty 8.0 of 10

    MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate ...

  3. CT-PrepAgent: Bounded Policy and Controlled Execution for Adaptive CT Data Preparation

    cs.AI 2026-08 conditional novelty 6.0 of 10

    An LLM agent constrained to a fixed CT preprocessing menu, with deterministic execution and verification, raised verified output yield and matched baseline utility on three public and two private cohorts.

  4. Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A multimodal framework using pulmonary findings, LLM-generated reasoning, and cardiac subvolume features achieves AUC 0.919 for CVD screening and 0.838 for CVD mortality from LDCT on NLST.

Pith tools