Pith. sign in

REVIEW 9 cited by

Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08495 v4 pith:2PQ5EM6D submitted 2024-03-13 cs.CL

Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator

classification cs.CL
keywords llmsmedicalclinicallanguagelargemodelspatientapplication
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) have demonstrated remarkable proficiency in human interactions, yet their application within the medical field remains insufficiently explored. Previous works mainly focus on the performance of medical knowledge with examinations, which is far from the realistic scenarios, falling short in assessing the abilities of LLMs on clinical tasks. In the quest to enhance the application of Large Language Models (LLMs) in healthcare, this paper introduces the Automated Interactive Evaluation (AIE) framework and the State-Aware Patient Simulator (SAPS), targeting the gap between traditional LLM evaluations and the nuanced demands of clinical practice. Unlike prior methods that rely on static medical knowledge assessments, AIE and SAPS provide a dynamic, realistic platform for assessing LLMs through multi-turn doctor-patient simulations. This approach offers a closer approximation to real clinical scenarios and allows for a detailed analysis of LLM behaviors in response to complex patient interactions. Our extensive experimental validation demonstrates the effectiveness of the AIE framework, with outcomes that align well with human evaluations, underscoring its potential to revolutionize medical LLM testing for improved healthcare delivery.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

    cs.HC 2024-05 conditional novelty 8.0

    AgentClinic is a multimodal agent benchmark demonstrating that LLM diagnostic accuracy on MedQA drops to below one-tenth in sequential clinical simulations, with Claude-3.5 leading and large tool-use differences acros...

  2. Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

    cs.HC 2026-05 unverdicted novelty 7.0

    PWP generates realistic, diverse virtual patients via HEXACO personality parametrization, achieving clinician-rated realism close to human actors, wider behavioral range than baselines, and reduced oversharing while a...

  3. Simulating Couple Conflict: Designing A Multi-Agent System for Therapy Training and Practice

    cs.CY 2026-01 conditional novelty 7.0

    A stateful multi-agent system simulates demand-withdraw couple conflicts across six stages for therapist training and outperforms prompt-based baselines in realism and state detection.

  4. MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support

    cs.LG 2026-04 unverdicted novelty 6.0

    MoBayes separates LLM language parsing from Bayesian probabilistic reasoning in conversational clinical decision support and reports performance gains over standalone frontier LLMs across multiple knowledge bases and ...

  5. MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support

    cs.LG 2026-04 conditional novelty 6.0

    Separating language from Bayesian reasoning lets cheap LLM sensors beat larger standalone LLM doctors on conversational diagnosis with controllable abstention and lower cost.

  6. ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

    cs.CL 2026-03 conditional novelty 6.0

    ClinConsensus, a Chinese medical benchmark with physician-calibrated thresholded rubric coverage, shows LLMs achieve 39.6–52.1% rubric accuracy but only 17.8–32.9% CACS@10, exposing a large coverage gap.

  7. A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid

    cs.CL 2026-02 conditional novelty 6.0

    A patient simulator integrating medical, linguistic, and behavioral profiles exposes a monotonic performance decline in an antidepressant decision aid as simulated health literacy decreases.

  8. Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support

    cs.AI 2026-05 unverdicted novelty 5.0

    Multi-turn evidence seeking reduces LLM diagnostic accuracy by 12.75% and supporting-evidence quality by 24.36% versus full-context evaluation in a new OSCE-inspired benchmark across 468 cases and 15 models.

  9. MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support

    cs.LG 2026-04 unverdicted novelty 5.0

    BMBE separates LLM language handling from a standalone Bayesian diagnostic engine, producing calibrated selective diagnosis, a performance gap over frontier LLMs, and robustness to adversarial inputs.