Pith. sign in

REVIEW 8 cited by

AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09742 v4 pith:R4DUNTF4 submitted 2024-02-15 cs.CL

classification cs.CL
keywords llmsmedicalemphclinicalhospitalinteractionsdiagnosticgaps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial intelligence has significantly advanced healthcare, particularly through large language models (LLMs) that excel in medical question answering benchmarks. However, their real-world clinical application remains limited due to the complexities of doctor-patient interactions. To address this, we introduce \textbf{AI Hospital}, a multi-agent framework simulating dynamic medical interactions between \emph{Doctor} as player and NPCs including \emph{Patient}, \emph{Examiner}, \emph{Chief Physician}. This setup allows for realistic assessments of LLMs in clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and NPCs to evaluate LLMs' performance in symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance diagnostic accuracy through iterative discussions. Despite improvements, current LLMs exhibit significant performance gaps in multi-turn interactions compared to one-step approaches. Our findings highlight the need for further research to bridge these gaps and improve LLMs' clinical diagnostic capabilities. Our data, code, and experimental results are all open-sourced at \url{https://github.com/LibertFan/AI_Hospital}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    HealthAgentBench is a new benchmark of 54 healthcare agent tasks where even the strongest frontier AI agent reaches only about 42% success rate on end-to-end clinical workflows.

  2. RE-MCDF: Closed-Loop Multi-Expert LLM Reasoning for Knowledge-Grounded Clinical Diagnosis

    cs.AI 2026-02 unverdicted novelty 6.0 of 10

    RE-MCDF introduces a generation-verification-revision closed-loop multi-expert LLM architecture guided by a medical knowledge graph to enforce inter-disease logical constraints and outperform baselines on neurology EM...

  3. Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation

    cs.CL 2025-10 conditional novelty 6.0 of 10

    DyReMe generates fresh, confounded medical consultation cases to expose weaknesses in state-of-the-art LLMs across accuracy, veracity, helpfulness, and consistency beyond static benchmarks.

  4. TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

    cs.MA 2025-07 accept novelty 6.0 of 10

    TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.

  5. Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning

    cs.CL 2025-05 unverdicted novelty 6.0 of 10

    DoctorAgent-RL trains a Qwen2.5-7B doctor agent via multi-agent RL on the new MTMedDialog dataset to conduct dynamic, question-driven consultations, reaching 70% exact diagnostic match in real-patient trials.

  6. R2MED: A Benchmark for Reasoning-Driven Medical Retrieval

    cs.IR 2025-05 accept novelty 6.0 of 10

    R2MED is the first benchmark for reasoning-driven medical retrieval, where even top models reach only 41.4 nDCG@10 on queries requiring inference beyond lexical or semantic overlap.

  7. A Survey of Scaling in Large Language Model Reasoning

    cs.AI 2025-04 unverdicted novelty 3.0 of 10

    A survey categorizing scaling in LLM reasoning across input size, steps, rounds, training, and future directions, noting that scaling can negatively affect performance.

  8. Large Language Model Agent: A Survey on Methodology, Applications and Challenges

    cs.CL 2025-03 accept novelty 3.0 of 10

    A survey that deconstructs LLM agent systems via a methodology-centered taxonomy linking design principles to emergent behaviors, applications, and challenges.

Pith tools