Pith. sign in

REVIEW 45 cited by

Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02957 v3 pith:EV3WXX23 submitted 2024-05-05 cs.AI

Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

classification cs.AI
keywords agentsmedicalsimulacrumagenthospitalllmstreatingautonomous
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The recent rapid development of large language models (LLMs) has sparked a new wave of technological revolution in medical artificial intelligence (AI). While LLMs are designed to understand and generate text like a human, autonomous agents that utilize LLMs as their "brain" have exhibited capabilities beyond text processing such as planning, reflection, and using tools by enabling their "bodies" to interact with the environment. We introduce a simulacrum of hospital called Agent Hospital that simulates the entire process of treating illness, in which all patients, nurses, and doctors are LLM-powered autonomous agents. Within the simulacrum, doctor agents are able to evolve by treating a large number of patient agents without the need to label training data manually. After treating tens of thousands of patient agents in the simulacrum (human doctors may take several years in the real world), the evolved doctor agents outperform state-of-the-art medical agent methods on the MedQA benchmark comprising US Medical Licensing Examination (USMLE) test questions. Our methods of simulacrum construction and agent evolution have the potential in benefiting a broad range of applications beyond medical AI.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 45 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

    cs.HC 2024-05 conditional novelty 8.0

    AgentClinic is a multimodal agent benchmark demonstrating that LLM diagnostic accuracy on MedQA drops to below one-tenth in sequential clinical simulations, with Claude-3.5 leading and large tool-use differences acros...

  2. LegalWorld: A Life-Cycle Interactive Environment for Legal Agents

    cs.CL 2026-06 unverdicted novelty 7.0

    LegalWorld is a life-cycle interactive environment modeling Chinese civil litigation as five causally connected stages grounded in 75,309 judgments, paired with LongJud-Bench for cross-stage agent evaluation.

  3. Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory

    cs.AI 2026-06 unverdicted novelty 7.0

    SkeMex distills agent trajectories into value-aware skills organized in general/task/action branches and evolves them via a closed-loop Read-Write-Assess-Govern process, outperforming prior memory agents on clinical tasks.

  4. Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases

    cs.CL 2026-06 unverdicted novelty 7.0

    MedSP1000 benchmark shows top LLMs complete at most 60.4% of expert rubric items during multi-turn standardized patient simulations.

  5. ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

    cs.AI 2026-06 unverdicted novelty 7.0

    ClinicalMC is a benchmark of 1,275 Chinese and 5,804 English multi-course clinical samples across four stages, evaluated via a multi-agent framework on closed-source, open-source, and medical LLMs in static and dynami...

  6. Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

    cs.AI 2026-05 unverdicted novelty 7.0

    A survey that unifies prior work on multi-agent LLM systems via the LIFE framework, mapping dependencies across collaboration, failure attribution, and autonomous self-evolution while identifying cross-stage challenges.

  7. Active Learning for Communication Structure Optimization in LLM-Based Multi-Agent Systems

    cs.MA 2026-05 unverdicted novelty 7.0

    An ensemble-based information-theoretic active learning method with ensemble Kalman inversion selects valuable tasks to optimize communication structures in LLM multi-agent systems under constrained budgets.

  8. VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing

    cs.MA 2026-04 unverdicted novelty 7.0

    VERITAS is a multi-agent system for verifiable hypothesis testing on multimodal clinical MRI datasets that achieves 81.4% verdict accuracy with frontier models and introduces an epistemic evidence labeling framework.

  9. Emergent Coordination in Multi-Agent Language Models

    cs.MA 2025-10 unverdicted novelty 7.0

    Multi-agent LLM systems can be steered via prompt design from mere aggregates to higher-order collectives with identity-linked differentiation and goal-directed complementarity, as measured by partial information deco...

  10. SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

    cs.CV 2025-05 conditional novelty 7.0

    Presents SpatialScore benchmark for MLLM spatial reasoning, evaluates 49 models showing large human gap, and supplies SpatialCorpus plus SpatialAgent to improve performance.

  11. MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

    cs.CL 2026-07 conditional novelty 6.0

    MedGame converts static clinical case reports into structured interactive storytelling games and shows that fine-tuned open-source LLMs approach commercial performance on a new 5,000-case benchmark.

  12. Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning

    cs.LG 2026-07 conditional novelty 6.0

    An agent with structured CT evidence packages and guideline-guided control iteratively escalates imaging phases only when current evidence is judged insufficient for diagnosis.

  13. Mandol: An Agglomerative Agent Memory System for Long-Term Conversations

    cs.DB 2026-06 unverdicted novelty 6.0

    Mandol unifies memory storage and retrieval into an agglomerative semantic graph architecture with quantitative query mechanisms, reporting best accuracy on LoCoMo and LongMemEval plus 5.4x retrieval and 4.8x insertio...

  14. MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

    cs.AI 2026-06 unverdicted novelty 6.0

    MedEvoEval is an executable longitudinal evaluation framework that converts medical cases into action-gated simulated episodes to track how doctor agents evolve decision-making, resource use, and experience across mul...

  15. MetaPS: Adaptive Programmatic Strategy Selection for Market Agents

    cs.AI 2026-06 unverdicted novelty 6.0

    MetaPS trains models via simulation rollouts to select from programmatic strategy libraries for market agents, yielding better performance than fixed or direct LLM baselines across model sizes.

  16. The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models

    cs.MA 2026-06 unverdicted novelty 6.0

    Epi-LLM integrates LLMs as agents in ABM epidemic simulations, finding reduced peak infections, 58-65% quarantine compliance, and perceived severity as top predictor with pseudo-R² 0.055 comparable to human data.

  17. Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

    cs.AI 2026-06 unverdicted novelty 6.0

    Traj-Evolve combines non-parametric experience retrieval and multi-agent RL with a leave-one-out unification strategy to outperform baselines on lung cancer prediction from up to five years of multimodal EHRs, includi...

  18. Enhancing LLM Metacognition via Cognitive Pairwise Training

    cs.LG 2026-05 unverdicted novelty 6.0

    CPT is introduced as a pairwise reasoning-trace comparison stage that improves the reasoning-metacognition trade-off over standard SFT+RL pipelines across model scales.

  19. EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

    cs.AI 2026-05 unverdicted novelty 6.0

    EHRBench uses an EHR-LLM-KB pipeline to automatically create 960,067 reliable QA items spanning diagnosis, treatment, and prognosis for large-scale LLM evaluation in clinical decision making.

  20. Active Learning for Communication Structure Optimization in LLM-Based Multi-Agent Systems

    cs.MA 2026-05 unverdicted novelty 6.0

    An ensemble-based information-theoretic active learning method using ensemble Kalman inversion selects valuable tasks to optimize communication structures in LLM multi-agent systems more reliably than random sampling ...

  21. DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents

    cs.AI 2026-05 unverdicted novelty 6.0

    DTap is a new red-teaming platform for AI agents that uses autonomous exploration across realistic simulations to discover vulnerabilities and creates a verifiable benchmark dataset.

  22. VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing

    cs.MA 2026-04 conditional novelty 6.0

    A four-phase multi-agent co-scientist tests natural-language hypotheses on cardiac and glioma MRI and labels outcomes Supported, Refuted, Underpowered, or Invalid with an executable evidence trail.

  23. Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography

    cs.CV 2025-11 conditional novelty 6.0

    A multimodal framework using pulmonary findings, LLM-generated reasoning, and cardiac subvolume features achieves AUC 0.919 for CVD screening and 0.838 for CVD mortality from LDCT on NLST.

  24. TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

    cs.MA 2025-07 accept novelty 6.0

    TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.

  25. Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning

    cs.CL 2025-05 unverdicted novelty 6.0

    DoctorAgent-RL trains a Qwen2.5-7B doctor agent via multi-agent RL on the new MTMedDialog dataset to conduct dynamic, question-driven consultations, reaching 70% exact diagnostic match in real-patient trials.

  26. HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

    cs.CL 2024-12 unverdicted novelty 6.0

    HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.

  27. GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning

    cs.LG 2024-06 unverdicted novelty 6.0

    GuardAgent safeguards LLM agents by generating task plans from safety requests and mapping them to executable guardrail code, achieving over 98% accuracy on a healthcare access-control benchmark and 83% on a web safet...

  28. Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application

    cs.CL 2026-06 unverdicted novelty 5.0

    This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environm...

  29. MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis

    cs.MA 2026-06 unverdicted novelty 5.0

    MeDxAgent multi-agent system improves interactive medical diagnosis accuracy by 10.3% on a new 4,421-case benchmark by using specific design choices like collecting demographics first and passing summarized dialogues.

  30. Beyond Isolated Behaviors: Hierarchical User Modeling for LLM Personalization

    cs.CL 2026-06 unverdicted novelty 5.0

    PHF applies Bourdieu's Theory of Practice to create hierarchical user models for LLM personalization and reports consistent gains on the LaMP benchmark.

  31. SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning

    cs.CL 2026-05 unverdicted novelty 5.0

    SEMA-RAG assigns clinical schema interpretation, sufficiency-driven retrieval, and evidence adjudication to three agents in a self-evolving multi-agent RAG system, reporting +6.46 average accuracy gains over baselines...

  32. Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

    cs.AI 2026-05 conditional novelty 5.0

    The survey proposes the LIFE framework to unify fragmented research on collaboration, failure attribution, and self-evolution in LLM multi-agent systems into a progression toward self-organizing intelligence.

  33. A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications

    cs.IR 2026-05 unverdicted novelty 5.0

    A survey that taxonomizes agent skills for LLM-based agents across representation, acquisition, retrieval, and evolution stages while reviewing methods, resources, and open challenges.

  34. OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

    cs.AI 2026-03 conditional novelty 5.0

    OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.

  35. InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

    cs.CL 2025-10 conditional novelty 5.0

    Rubric-based incremental RL with LLM-generated case-specific checklists lifts Qwen3-4B's HealthBench-Hard score from 7.0 to 27.5 with 2k samples, and improves InfoBench instruction-following from 42.0 to 82.9.

  36. Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics

    cs.AI 2025-10 reject novelty 5.0

    An LLM-powered multi-agent school with dual experience/knowledge memory increasingly reproduces an expert-curated classroom script, with the full memory configuration scoring highest.

  37. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

  38. Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

    cs.CV 2026-07 conditional novelty 4.0

    A scoping review of 557 studies finds medical agentic AI is technically promising but overwhelmingly validated on benchmarks, simulations, and retrospective data rather than in real clinical workflows.

  39. From Question Answering to Task Completion: A Survey on Agent System and Harness Design

    cs.AI 2026-06 unverdicted novelty 4.0

    Survey framing LLM agents as model-plus-harness systems, decomposing harness responsibilities, mapping them to tasks, and highlighting open challenges in evaluation, safety, and co-evolution.

  40. SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning

    cs.CL 2026-05 unverdicted novelty 4.0

    SEMA-RAG is a three-agent self-evolving RAG system that reports an average 6.46-point accuracy gain over the strongest baseline across five medical QA benchmarks and five LLM backbones.

  41. A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications

    cs.IR 2026-05 unverdicted novelty 4.0

    The paper surveys agent skills for LLM agents, organizing the literature into a four-stage lifecycle of representation, acquisition, retrieval, and evolution while highlighting their role in system scalability.

  42. A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

    cs.AI 2025-07 accept novelty 4.0

    The paper delivers the first systematic review of self-evolving agents, structured around what components evolve, when adaptation occurs, and how it is implemented.

  43. From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review

    cs.AI 2025-04 accept novelty 4.0

    A survey consolidating benchmarks, agent frameworks, real-world applications, and protocols for LLM-based autonomous agents into a proposed taxonomy with recommendations for future research.

  44. A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications

    cs.IR 2026-05 unverdicted novelty 3.0

    A survey that defines agent skills as reusable procedural artifacts and reviews methods, resources, and applications across their representation, acquisition, retrieval, and evolution stages.

  45. Large Language Model Agent: A Survey on Methodology, Applications and Challenges

    cs.CL 2025-03 accept novelty 3.0

    A survey that deconstructs LLM agent systems via a methodology-centered taxonomy linking design principles to emergent behaviors, applications, and challenges.