AutoMedBench evaluates AI agents on long-horizon medical workflows across five stages and finds validation and submission as dominant failure points based on thousands of runs.
Fhir-agentbench: Benchmarking llm agents for realistic interoperable ehr question answering
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
EHR-Complex is a new interactive benchmark on MIMIC-IV with 52K tasks averaging 31.93 SQL components, where top LLMs achieve 62.3% accuracy and exhibit SQL logic, medical-code, and semantic failures.
ClinSeekAgent automates active multimodal evidence seeking for clinical reasoning, improving LLM performance on raw EHR and CXR tasks while enabling distillation into smaller models.
D2MDT uses department-aware multi-agent consultation with residual deliberation to improve EHR-based mortality prediction and efficiency.
citing papers explorer
-
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench evaluates AI agents on long-horizon medical workflows across five stages and finds validation and submission as dominant failure points based on thousands of runs.
-
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
EHR-Complex is a new interactive benchmark on MIMIC-IV with 52K tasks averaging 31.93 SQL components, where top LLMs achieve 62.3% accuracy and exhibit SQL logic, medical-code, and semantic failures.
-
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
ClinSeekAgent automates active multimodal evidence seeking for clinical reasoning, improving LLM performance on raw EHR and CXR tasks while enabling distillation into smaller models.
-
D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction
D2MDT uses department-aware multi-agent consultation with residual deliberation to improve EHR-based mortality prediction and efficiency.
- GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents