Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

This paper claims that a conversational AI antidepressant decision aid performs markedly worse for patients with limited health literacy, and that this gap is measurable and reproducible through a controllable patient simulation framework.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:08 UTC pith:OU3665GE

load-bearing objection A solid, NIST-aligned patient simulator that measures a health-literacy performance gradient in an antidepressant decision aid, but the headline gradient is partly built into the prompt design and external validity is untested. the 3 major comments →

arxiv 2602.11391 v4 pith:OU3665GE submitted 2026-02-11 cs.CL

A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid

classification cs.CL
keywords patient simulationconversational AIhealth literacyantidepressant decision supportAI risk assessmentAI equityretrieval-augmented generationLLM evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper builds a patient simulator that varies three dimensions at once: medical history (drawn from a large EHR cohort), language style (a graded health-literacy scale plus condition-specific speech patterns), and conversational behavior (cooperative, distracted, adversarial). After running 500 simulated intake conversations through a conversational antidepressant decision aid, it finds that the aid's retrieval of the correct medical concepts degrades as health literacy drops: rank-1 concept retrieval falls from 81.9% for proficient speakers to 47.6% for limited-literacy speakers, and downstream antidepressant recommendations get worse. The central claim is that this monotonic gradient is a concrete equity risk in deployed-style conversational clinical AI, and that the same simulation framework can test mitigations before real patients are involved.

Core claim

On the paper's own terms, the discovery is a monotonic performance gradient: a conversational AI decision aid that retrieves medication, diagnosis, and procedure concepts from free-text patient responses succeeds far more often when patients use precise clinical terminology than when they express the same underlying facts in plain, vague, or affect-heavy language. The gradient persists into the final antidepressant recommendation. Because the simulator holds medical content fixed while varying only expression and engagement, the paper attributes the difference to the aid's reliance on canonical vocabulary rather than to differences in clinical histories.

What carries the argument

The load-bearing mechanism is the composition of three controlled profiles inside a single prompt-driven simulator. Medical profiles come from the MAGI algorithm, a risk-ratio-gated feature selector that picks outcome-relevant, statistically independent clinical features from EHR data and then adds a bounded amount of correlated residual diversity. Linguistic profiles enforce a five-level health-literacy continuum plus condition-specific styles (depression, illness anxiety). Behavioral profiles simulate cooperative, distracted, and adversarial engagement. A chain-of-thought prompting pipeline makes every simulated turn traceable: facts are referenced by index, style-transferred, and emitted

Load-bearing premise

That the simulated linguistic profiles—especially the 'limited health literacy' voice with slang, vague quantities, and fragmented sentences—are a faithful stand-in for how real low-literacy patients actually talk to a machine; if real patients express the same facts through different patterns, the measured gradient may not transfer to actual clinical encounters.

What would settle it

Record real outpatient intake conversations from patients whose health literacy has been measured with a validated instrument, run the same AI decision aid on those transcripts, and compare rank-1 concept retrieval across literacy strata; if the real-world gradient is flat, reversed, or much smaller than 47.6% to 81.9%, the paper's central claim would be falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Conversational clinical AI that relies on precise terminology will likely perform inequitably across health-literacy levels; designers should treat health literacy as a first-order risk factor and test mitigations such as clarification prompts or terminology normalization.
  • The measured gradient offers a concrete benchmark for intake systems: rank-1 concept retrieval from 47.6% (limited) to 69.6% (functional) to 81.9% (proficient), with downstream recommendation F1 dropping from 0.73 to 0.48.
  • Because the simulator keeps medical content fixed while varying language and behavior, failures can be localized to specific intake stages—retrieval versus downstream recommendation—rather than attributed to patient acuity.
  • LLM judge agreement with human annotators (0.78 kappa overall, comparable to human-human 0.73 kappa) supports using automated judges for large-scale screening of simulated conversations, with human review reserved for subtle semantic errors.
  • The controlled error-injection method (perturbing concepts such as hypertension to prehypertension) provides a way to validate annotator and judge sensitivity, not just their agreement.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural next step is to test whether a clarification or normalization preprocessing stage shrinks the health-literacy gap; the simulator makes such an ablation cheap and measurable before any clinical trial.
  • Because the degradation is monotonic, the framework could be used to estimate a safe operating range for intake systems: the maximum jargon level at which retrieval accuracy stays above an acceptable threshold for a given population.
  • The finding likely generalizes beyond antidepressants to any conversational intake system that normalizes free text against a canonical vocabulary—symptom checkers, triage bots, history-taking agents—so the same simulation approach may expose similar equity gaps elsewhere.
  • The paper's stated limitation about ecological validity implies that the highest-value next step is recording real patient intakes to calibrate the linguistic profiles; until then, the measured 47.6% floor should be read as a simulation-based estimate, not an observed clinical rate.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents a patient simulation framework for evaluating conversational healthcare AI, implemented as a GPT-4.1-based simulated patient with medical, linguistic, and behavioral profiles. Medical profiles are generated from All of Us EHR data via a risk-ratio-gated feature selection algorithm (MAGI); linguistic profiles span a health-literacy gradient and condition-specific patterns; behavioral profiles include cooperative, distracted, and adversarial styles. The simulator is validated by human annotation and an LLM judge, then used to evaluate an AI Decision Aid for antidepressant selection across 500 simulated conversations. The main empirical claim is a monotonic performance gradient: Rank-1 concept retrieval falls from 81.9% for proficient health literacy to 47.6% for limited health literacy, with corresponding declines in antidepressant recommendation F1. The paper concludes that health literacy is a primary equity risk for conversational clinical AI.

Significance. If the reported gradient reflects real patient communication patterns, the finding is practically important: it would identify a concrete failure mode for deployed-style intake systems and motivate MANAGE-function interventions. The paper's methodological strengths include a public data/code commitment, a controlled perturbation protocol for annotation validation, and a dual human/LLM annotation pipeline with generally high agreement. The MAGI algorithm is transparent in its feature lineage. However, the headline causal claim currently rests on simulated utterances whose ecological validity is explicitly untested, and the profiles and recommender share the same underlying algorithm and dataset. These issues do not invalidate the framework as a stress-testing tool, but they do prevent the paper from supporting the broad equity conclusion as stated.

major comments (3)
  1. [§3.3.3 (Table 2), §4.6 (Table 11), §5.5] The headline gradient (Rank-1 retrieval 47.6% vs 81.9%) is a compound-treatment effect. The Limited profile bundles 'everyday terms, slang, vague quantities' and 'short, fragmented sentences; frequent fillers', while Proficient bundles precise terminology and 'multi-clause, logically sequenced sentences'. Vagueness, fragmentation, and disfluency are all plausible causes of degraded embedding-based retrieval; the design does not isolate health literacy. The paper's own §5.5 concedes that ecological validity is untested. To support the causal attribution to health literacy, an ablation is needed (e.g., vary vocabulary complexity alone, sentence structure alone, and quantity specificity alone), or a validation against real patient utterances stratified by measured health literacy. Otherwise the monotonic degradation should be described as a prompt-design sensitivity result, not a health-lit
  2. [§3.3.2, §3.5, §4.6] The simulator's medical profiles are generated by the MAGI algorithm (risk-ratio gating and independence screening) and the AI Decision Aid's analytical recommendation engine is also 'based on the Medical Artificial General Intelligence Algorithm (MAGI)' trained on the same All of Us dataset. This shared ancestry means the downstream F1 gradient is not an external, black-box assessment: the simulated patients and the recommender are coupled by construction. The paper should demonstrate that the gradient persists with an independent recommendation model (e.g., logistic regression on held-out data or an external cohort), or explicitly reframe the results as an internal stress test of a specific MAGI-based pipeline.
  3. [§4.6, Tables 11, Figure 9] The central claim of monotonic degradation is presented without uncertainty quantification or statistical tests. Table 11 and Figure 9 report point estimates only; with 500 conversations and per-profile sample sizes of 440–1023 concepts, sampling variability and conversation-level clustering are not addressed. Report bootstrap confidence intervals or profile-comparison tests (e.g., clustered bootstrap by conversation) for Rank-1 retrieval and recommendation F1 before drawing clinical-equity conclusions.
minor comments (4)
  1. [Appendix 9.1.2] The risk-ratio acceptance condition is stated inconsistently with the main text. §3.3.2 requires 1/1.5 < RR ≤ high (high=7), while the appendix first repeats this but later says candidates with RR >1.5 or RR <1/1.5 are eligible and defines independence as [1/1.5,1.5]. This should be reconciled for reproducibility.
  2. [§3.7] The three evaluation settings total 630 conversations (300+180+150); explain more explicitly how shared profiles reduce this to 500 unique conversations.
  3. [Data availability footnotes] The footnotes say the GitHub link will be provided upon acceptance. If the manuscript claims public release, provide a persistent repository link in the submitted version.
  4. [Table 11] The 'Not Retrieved' percentages are not complementary to Rank-1 or Top-20, and the definition of Top-20 (inclusive of Rank-1 or not) is ambiguous. Clarify the metric definitions so the row values can be interpreted precisely.

Circularity Check

2 steps flagged

Downstream 'recommendation accuracy' reduces to MAGI-vs-MAGI self-agreement; the headline literacy gradient is also partly an artifact of the linguistic profile definitions.

specific steps
  1. self definitional [§3.2 Data; §3.5 AI Decision Aid; Appendix 9.1.2 (PredictResp); Fig. 9 caption]
    "Both the patient simulator medical profiles and the AI Decision Aid prediction models are derived from the All of Us Research Program Registered Tier v8 dataset ... The analytical advice system is based on the Medical Artificial General Intelligence Algorithm (MAGI) [Citation removed to preserve anonymity] ... PredictResp(S, e₀): Returns the probability of response to the antidepressant ... computed using the MAGI algorithm."

    The 'reference' recommendation used for F1 is produced by MAGI from the full medical profile, while the AI Decision Aid's recommendation is produced by the same MAGI algorithm from the retrieved concept subset. When retrieval is perfect, the two MAGI inputs are identical, so F1 = 1 by construction. Thus the reported 'declines in antidepressant recommendation accuracy' are not validated against any external clinical outcome; they measure MAGI's self-agreement under information loss, i.e., the downstream 'prediction' is defined in terms of the same algorithm that generates the target.

  2. renaming known result [§3.3.3 Table 2; §4.6 Table 11; §5.3]
    "Limited ... Vocab: Everyday terms, slang, vague quantities. Structure: Short, fragmented sentences; frequent fillers. ... Proficient ... Vocab: Technical terms; qualifiers such as “likely” or “seems improved.” Structure: Multi-clause, logically sequenced sentences. ... This gradient reflects the decision aid’s reliance on precise terminology and structured expression for accurate retrieval. Vague or affect-heavy language reduces ranking accuracy."

    The independent variable 'health literacy' is operationalized directly as vague/fragmented language versus precise/structured language, and retrieval ranking is driven by semantic similarity to canonical medical vocabulary. The monotonic Rank-1 gradient (47.6% Limited to 81.9% Proficient) is therefore a near-immediate consequence of the profile definitions and the retrieval mechanism, not an independent empirical discovery about health literacy. The paper's own §5.5 concedes that 'ecological validity remains untested,' which underscores that the causal label is imposed on the manipulated linguistic style rather than derived from real patient-language data.

full rationale

The paper is not globally circular: medical profiles are grounded in All of Us risk-ratio calculations, linguistic/behavioral profiles are controlled prompt manipulations, and the Rank-1 concept-retrieval result is a measured outcome of an embedding-based RAG system rather than a fitted parameter. However, two load-bearing reductions qualify the headline claims. First, the downstream 'recommendation accuracy' (Fig. 9) compares the AI Decision Aid's MAGI-based recommendation with a reference generated by the same MAGI algorithm from the same profile; this is a self-consistency metric and reduces by construction once the retrieved concept set equals the profile concept set. Second, the 'health literacy gradient' is encoded into the simulator's linguistic profiles (vague/fragmented vs. precise/structured wording), so the observed degradation is substantially an artifact of the operationalization rather than evidence about real-world health-literacy variation. The paper acknowledges the latter in §5.5, but the abstract and §5.3 present the gradient as a discovered equity risk. These issues do not invalidate the simulator framework as a controlled stress-testing tool, but they mean the central equity conclusion is partially circular: the simulator is instructed to produce the very communication pattern that the AI Decision Aid is then shown to handle poorly. Score 6 reflects this partial, construction-level circularity.

Axiom & Free-Parameter Ledger

5 free parameters · 7 axioms · 0 invented entities

The central claim rests on several hand-set thresholds, an undisclosed self-cited MAGI model shared by the simulator and the system under test, and the assumption that LLM-generated profile language faithfully represents real patient communication. No new physical entities are introduced.

free parameters (5)
  • high (risk-ratio upper threshold) = 7
    Set by hand in §3.3.2; controls how strongly features may be coupled before being excluded from medical profiles. Affects profile composition and downstream evaluation.
  • max_residual_additions = 3–5
    Hand-chosen number of residual features added in Stage 4 (§3.7); influences profile diversity and concept recall.
  • K (top-K predictor count) = 500
    Number of outcome-relevant features retained in Stage 1; arbitrary and not tuned against a benchmark (§3.3.2).
  • low (risk-ratio lower bound) = 1/1.5
    Symmetric threshold from mRMR-style reasoning; hand-set in §3.3.2.
  • MAGI internal parameters = not disclosed
    The MAGI algorithm (citation removed) is trained on All of Us and used to predict response probabilities in both profile generation (PredictResp) and the decision aid; its parameters are not specified in this paper.
axioms (7)
  • domain assumption All of Us Registered Tier v8 data are representative of patients with MDD and their antidepressant response patterns.
    Used throughout §3.2 for cohort, features, and response distributions; no external validation.
  • domain assumption Taking an antidepressant ≥10 weeks without switching or augmenting is a valid surrogate for response.
    Definition from ref [2] (same group's prior work), applied in §3.2.
  • ad hoc to paper The risk-ratio gate with thresholds 1/1.5 and 7 produces clinically coherent, independent feature sets.
    Justified only by analogy to mRMR (§3.3.2), not by clinical or statistical benchmarking.
  • domain assumption The five linguistic profiles and three behavioral profiles faithfully operationalize real patient communication styles.
    Based on literature (§3.3.3–3.3.4) but not validated against real patients; §5.5 concedes ecological validity is untested.
  • domain assumption GPT-4.1 with the provided CoT prompt instantiates the specified profiles accurately.
    The entire method depends on LLM fidelity; only indirect validation via annotation is provided.
  • domain assumption MAGI's dependent Bayesian estimates are valid for antidepressant response.
    Self-cited, citation removed in §3.5; no evidence in this paper.
  • domain assumption Human annotation and LLM judgment correctly capture clinical accuracy of concepts.
    Used as ground truth in §3.6; the low agreement on perturbed concepts (κ=0.24–0.30) weakens this assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 25914 in / 13318 out tokens · 126129 ms · 2026-08-03T00:08:29.105407+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid." pith.science (2026). https://pith.science/paper/OU3665GE

@misc{pith2026260211391,
  author       = {Pith},
  title        = {Pith review of: A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OU3665GE}},
  note         = {Machine review of arXiv:2602.11391}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) MAP and MEASURE functions, providing an empirical basis for identifying and characterizing performance risks in conversational clinical AI across medical, linguistic, and behavioral patient variation. We applied the framework to a conversational decision aid for antidepressant selection in major depressive disorder (the AI Decision Aid). Methods: The simulator integrates three profile dimensions: (1) medical profiles constructed from All of Us electronic health records using risk-ratio gating; (2) linguistic profiles modeling a health literacy gradient and condition-specific communication; and (3) behavioral profiles representing cooperative, distracted, and adversarial engagement. We generated 500 simulated conversations and evaluated profile fidelity through human annotation and an LLM judge, then assessed downstream effects on the AI Decision Aid's concept retrieval and antidepressant recommendations. Results: The patient simulator expressed medical concepts with high fidelity (96.6% accurate across 8,210 concepts), with human inter-annotator agreement of 0.73 $\kappa$ and LLM-judge agreement against human annotators of 0.78 $\kappa$. Behavioral profiles were reliably distinguished (0.93 $\kappa$), and linguistic profiles showed moderate agreement (0.61 $\kappa$). The framework revealed monotonic degradation in AI Decision Aid performance across the health literacy gradient. Rank-1 concept retrieval increased from 47.6% for limited health literacy to 81.9% for proficient health literacy, with corresponding declines in antidepressant recommendation accuracy.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment

    cs.CY 2026-05 unverdicted novelty 6.0

    Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment whe...

Reference graph

Works this paper leans on

106 extracted references · 18 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    Abbo, Qi Zhang, Martin Zelder, and Elbert S

    Elmer D. Abbo, Qi Zhang, Martin Zelder, and Elbert S. Huang. 2008. The increasing number of clinical items addressed during the time of adult primary care visits. J. Gen. Intern. Med. 23, 12 (December 2008), 2058 –2065. https://doi.org/10.1007/s11606-008-0805-8

  2. [2]

    Sylvia, and Andrew A

    Farrokh Alemi, Mai Aljuaid, Naren Durbha, Melanie Yousefi, Hua Min, Louisa G. Sylvia, and Andrew A. Nierenberg. 2021. A surrogate measure for patient reported symptom remission in administrative data. BMC Psychiatry 21, (March 2021), 121. https://doi.org/10.1186/s12888-021-03133-1

  3. [3]

    All of Us

    All of Us Research Program Investigators, Joshua C. Denny, Joni L. Rutter, David B. Goldstein, Anthony Philippakis, Jordan W. Smoller, Gwynne Jenkins, and Eric Dishman. 2019. The “All of Us” Research Program. N. Engl. J. Med. 381, 7 (August 2019), 668–676. https://doi.org/10.1056/NEJMsr1809937

  4. [4]

    Krisztian Balog and ChengXiang Zhai. 2023. User Simulation for Evaluating Information Access Systems. In Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (SIGIR-AP ’23), November 26, 2023. Association for Computing Machinery, New York, NY, USA, 302–305. https://doi...

  5. [5]

    Zhijie Bao, Qingyun Liu, Xuanjing Huang, and Zhongyu Wei. 2025. SFMSS: Service Flow aware Medical Scenario Simulation for Conversational Data Generation. In Findings of the Association for Computational Linguistics: NAACL 2025 , April 2025. Association for Computational Linguistics, Albuquerque, New Mexico, 4586 –4604. https://doi.org/10.18653/v1/2025.fin...

  6. [6]

    Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee, Alexander Panchenko, Martin Potthast, Francisco Rangel, Paolo Rosso, Alisa Smirnova, E fstathios Stamatatos, Benno Stein, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, and Eva Zangerle. 2...

  7. [7]

    Savita Bhat and Vasudeva Varma. 2023. Large Language Models As Annotators: A Preliminary Evaluation For Annotating Low-Resource Language Content. In Proceedings of the 4th Workshop on Evaluation and Comparison of NLP Systems , November 2023. Association for Computational Linguistics, Bali, Indonesia, 100 –107. https://doi.org/10.18653/v1/2023.eval4nlp-1.8

  8. [8]

    Anna Bodonhelyi, Christian Stegemann -Philipps, Alessandra Sonanini, Lea Herschbach, Marton Szep, Anne Herrmann-Werner, Teresa Festl-Wietek, Enkelejda Kasneci, and Friederike Holderried. 2025. Modeling Challenging Patient Interactions: LLMs for Medical Communication Training. ArXiv E-Prints (March 2025), arXiv:2503.22250. https://doi.org/10.48550/arXiv.2503.22250

  9. [9]

    CDC. 2024. National Action Plan to Improve Health Literacy. Health Literacy. Retrieved November 4, 2025 from https://www.cdc.gov/health-literacy/php/develop-plan/national-action-plan.html

  10. [10]

    Cevasco, Rachel E

    Kevin E. Cevasco, Rachel E. Morrison Brown, Rediet Woldeselassie, and Seth Kaplan. 2024. Patient Engagement with Conversational Agents in Health Applications 2016 –2022: A Systematic Review and Meta -Analysis. J. Med. Syst. 48, 1 (2024), 40. https://doi.org/10.1007/s10916-024-02059-x

  11. [11]

    Xingran Chen, Zhenke Wu, Xu Shi, Hyunghoon Cho, and Bhramar Mukherjee. 2025. Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype data and open-source software. J. Am. Med. Inform. Assoc. 32, 7 (July 2025), 1227–1240. https://doi.org/10.1093/jamia/ocaf082

  12. [12]

    Young-Min Cho, Sunny Rai, Lyle Ungar, João Sedoc, and Sharath Chandra Guntuku. 2023. An Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives. Proc. Conf. Empir. Methods Nat. Lang. Process. Conf. Empir. Methods Nat. Lang. Process. 2023, (December 2023), 11346 –11369. https://doi.org/10.18653/v1/2023....

  13. [13]

    David A Cook, Joshua Overgaard, V Shane Pankratz, Guilherme Del Fiol, and Chris A Aakre. 2025. Virtual Patients Using Large Language Models: Scalable, Contextualized Simulation of Clinician -Patient Dialogue With Feedback. J. Med. Internet Res. 27, (April 2025), e68486. https://doi.org/10.2196/68486

  14. [14]

    Jessamyn Dahmen and Diane Cook. 2019. SynSys: A Synthetic Data Generation System for Healthcare Applications. Sensors 19, 5 (January 2019), 1181. https://doi.org/10.3390/s19051181

  15. [15]

    David DeVault, Ron Artstein, Grace Benn, Teresa Dey, Ed Fast, Alesia Gainer, Kallirroi Georgila, Jon Gratch, Arno Hartholt, Margaux Lhommet, Gale Lucas, Stacy Marsella, Fabrizio Morbini, Angela Nazarian, Stefan Scherer, Giota Stratou, Apar Suri, David Tra um, Rachel Wood, Yuyu Xu, Albert Rizzo, and Louis -Philippe Morency. 2014. SimSensei kiosk: a virtual...

  16. [16]

    Kathleen Kara Fitzpatrick, Alison Darcy, and Molly Vierhile. 2017. Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial. JMIR Ment. Health 4, 2 (June 2017), e7785. https://doi.org/10.2196/mental.7785

  17. [17]

    Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Ding, Zhilun Zhou, Fengli Xu, and Yong Li. 2024. Large language models empowered agent -based modeling and simulation: a survey and perspectives. Humanit. Soc. Sci. Commun. 11, 1 (September 2024), 1259. https://doi.org/10.1057/s41599-024-03611-3

  18. [18]

    Shengyue Guan, Haoyi Xiong, Jindong Wang, Jiang Bian, Bin Zhu, and Jian -guang Lou. 2025. Evaluating LLM - based Agents for Multi-Turn Conversations: A Survey. https://doi.org/10.48550/arXiv.2503.22458

  19. [19]

    Isabelle Guyon and André Elisseeff. 2003. An introduction to variable and feature selection. J Mach Learn Res 3, null (March 2003), 1157–1182

  20. [20]

    Hanchuan Peng, Fuhui Long, and C. Ding. 2005. Feature selection based on mutual information criteria of max - dependency, max -relevance, and min -redundancy. IEEE Trans. Pattern Anal. Mach. Intell. 27, 8 (August 2005), 1226–1238. https://doi.org/10.1109/TPAMI.2005.159

  21. [21]

    Friederike Holderried, Christian Stegemann -Philipps, Anne Herrmann -Werner, Teresa Festl -Wietek, Martin Holderried, Carsten Eickhoff, and Moritz Mahling. 2024. A Language Model –Powered Simulated Patient With Automated Feedback for History Taking: Prospecti ve Study. JMIR Med. Educ. 10, (August 2024), e59213. https://doi.org/10.2196/59213

  22. [22]

    Nelson, Charles Stromeyer Iv, Darlene King, Jina Suh, Li Zhou, and John Torous

    Yining Hua, Winna Xia, David Bates, George Luke Hartstein, Hyungjin Tom Kim, Michael Li, Benjamin W. Nelson, Charles Stromeyer Iv, Darlene King, Jina Suh, Li Zhou, and John Torous. 2025. Standardizing and Scaffolding Health Care AI -Chatbot Evaluation: Sys tematic Review. JMIR AI 4, 1 (November 2025), e69006. https://doi.org/10.2196/69006

  23. [23]

    Institute of Medicine (US) Committee on Health Literacy. 2004. Health Literacy: A Prescription to End Confusion. National Academies Press (US), Washington (DC). Retrieved February 8, 2026 from http://www.ncbi.nlm.nih.gov/books/NBK216032/

  24. [24]

    Mark Kalinich, James Luccarelli, Frank Moss, and John Torous. 2025. Leveraging simulation to provide a practical framework for assessing the novel scope of risk of LLMs in healthcare. 2025.11.10.25339903. https://doi.org/10.1101/2025.11.10.25339903

  25. [25]

    J. P. Kincaid, Jr Fishburne, Richard L. Rogers, and Brad S. Chissom. 1975. Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel. (February 1975). Retrieved February 10, 2026 from https://apps.dtic.mil/sti/html/tr/ADA006655/

  26. [26]

    Petrucka

    Abukari Kwame and Pammla M. Petrucka. 2021. A literature -based study of patient -centered care and communication in nurse-patient interactions: barriers, facilitators, and the way forward. BMC Nurs. 20, (September 2021), 158. https://doi.org/10.1186/s12912-021-00684-2

  27. [27]

    Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024. MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024. Association for Computational Linguistics, Miami,...

  28. [28]

    Daeun Kyung, Hyunseung Chung, Seongsu Bae, Jiho Kim, Jae Ho Sohn, Taerim Kim, Soo Kyung Kim, and Edward Choi. 2025. PatientSim: A Persona -Driven Simulator for Realistic Doctor -Patient Interactions. https://doi.org/10.48550/arXiv.2505.17818

  29. [29]

    Sahiti Labhishetty. 2023. Models and evaluation of user simulation in information retrieval. Thesis. University of Illinois at Urbana-Champaign. Retrieved November 29, 2025 from https://hdl.handle.net/2142/120103

  30. [30]

    Liliana Laranjo, Adam G Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y S Lau, and Enrico Coiera. 2018. Conversational agents in healthcare: a systematic review. J. Am. Med. Inform. Assoc. JAMIA 25, 9 (July 2018), 1248 –1258. https://doi.org/10.1093/jamia/ocy072

  31. [31]

    Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Nikolay Malkin, and Moksh Jain. 2025. LEARNING DIVERSE ATTACKS ON LARGE LANGUAGE MODELS FOR ROBUST RED-TEAMING AND SAFETY TUNING. (2025)

  32. [32]

    Abigail E Lewis, Nicole Weiskopf, Zachary B Abrams, Randi Foraker, Albert M Lai, Philip R O Payne, and Aditi Gupta. 2023. Electronic health record data quality assessment and tools: a systematic review. J. Am. Med. Inform. Assoc. 30, 10 (October 2023), 1730–1740. https://doi.org/10.1093/jamia/ocad120

  33. [33]

    Yusheng Liao, Yutong Meng, Yuhao Wang, Hongcheng Liu, Yanfeng Wang, and Yu Wang. 2024. Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator. https://doi.org/10.48550/arXiv.2403.08495

  34. [34]

    Ernest Lim, Yajie Vera He, Jared Joselowitz, Kate Preston, Mohita Chowdhury, Louis Williams, Aisling Higham, Katrina Mason, Mariane Melo, Tom Lawton, Yan Jia, and Ibrahim Habli. 2025. MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conte Xtual clinical conversational evaluation. https://doi.org/10.48550/arXiv.2508.19163

  35. [35]

    Lei Liu, Xiaoyan Yang, Junchi Lei, Yue Shen, Jian Wang, Peng Wei, Zhixuan Chu, Zhan Qin, and Kui Ren. 2024. A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions. https://doi.org/10.48550/arXiv.2406.03712

  36. [36]

    Edward Loper and Steven Bird. 2002. NLTK: The Natural Language Toolkit. In Proceedings of the ACL -02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics, July 2002. Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, 63 –70. https://doi.org/10.3115/1118108.1118117

  37. [37]

    Ming-Jie Luo, Shaowei Bi, Jianyu Pang, Lixue Liu, Ching-Kit Tsui, Yunxi Lai, Wenben Chen, Yahan Yang, Kezheng Xu, Lanqin Zhao, Ling Jin, Duoru Lin, Xiaohang Wu, Jingjing Chen, Rongxin Chen, Zhenzhen Liu, Yuxian Zou, Yangfan Yang, Yiqing Li, and Haotian Li n. 2025. A large language model digital patient system enhances ophthalmology history taking skills. ...

  38. [38]

    Xiang Luo, Zhiwen Tang, Jin Wang, and Xuejie Zhang. 2024. DuetSim: Building User Simulator with Dual Large Language Models for Task -Oriented Dialogues. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC -COLING 2024) , May 2024. ELRA and ICCL, Torino, Italia, 5414–5424. Retrieve...

  39. [39]

    Subhankar Maity and Manob Jyoti Saikia. 2025. Large Language Models in Healthcare and Medical Applications: A Review. Bioengineering 12, 6 (June 2025), 631. https://doi.org/10.3390/bioengineering12060631

  40. [40]

    Birger Moëll and Fredrik Sand Aronsson. 2025. Harm Reduction Strategies for Thoughtful Use of Large Language Models in the Medical Domain: Perspectives for Patients and Clinicians. J. Med. Internet Res. 27, 1 (July 2025), e75849. https://doi.org/10.2196/75849

  41. [41]

    Miquel Montaner, Beatriz López, and Josep Lluís de la Rosa. 2026. EVALUATION OF RECOMMENDER SYSTEMS THROUGH SIMULATED USERS. March 18, 2026. 303 –308. Retrieved March 18, 2026 from https://www.scitepress.org/Link.aspx?doi=10.5220/0002622703030308

  42. [42]

    Graham, Carrie D

    Tom Nadarzynski, Nicky Knights, Deborah Husbands, Cynthia A. Graham, Carrie D. Llewellyn, Tom Buchanan, Ian Montgomery, and Damien Ridge. 2024. Achieving health equity through conversational AI: A roadmap for design and implementation of inclusive chatbot s in healthcare. PLOS Digit. Health 3, 5 (May 2024), e0000492. https://doi.org/10.1371/journal.pdig.0000492

  43. [43]

    Don Nutbeam. 2000. Health literacy as a public health goal: a challenge for contemporary health education and communication strategies into the 21st century. Health Promot. Int. 15, 3 (September 2000), 259 –267. https://doi.org/10.1093/heapro/15.3.259

  44. [44]

    Paasche -Orlow and Michael S

    Michael K. Paasche -Orlow and Michael S. Wolf. 2007. The causal pathways linking health literacy to health outcomes. Am. J. Health Behav. 31 Suppl 1, (2007), S19-26. https://doi.org/10.5555/ajhb.2007.31.supp.S19

  45. [45]

    Maja Pavlovic and Massimo Poesio. 2024. The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation. In Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives) @ LREC -COLING 2024, May 2024. ELRA and ICCL, Torino, Italia, 100 –110. Retrieved February 11, 2026 from https://acla...

  46. [46]

    Pennebaker, Matthias R

    James W. Pennebaker, Matthias R. Mehl, and Kate G. Niederhoffer. 2003. Psychological Aspects of Natural Language Use: Our Words, Our Selves. Annu. Rev. Psychol. 54, Volume 54, 2003 (February 2003), 547 –577. https://doi.org/10.1146/annurev.psych.54.101601.145041

  47. [47]

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red Teaming Language Models with Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022. Association for Computational Linguistics, Abu Dhabi, United Arab Emira...

  48. [48]

    Paloma Rabaey, Stefan Heytens, and Thomas Demeester. 2025. SimSUM: Simulated Benchmark with Structured and Unstructured Medical Records. https://doi.org/10.48550/arXiv.2409.08936

  49. [49]

    Debra Roter and Judith A. Hall. 2006. Doctors Talking with Patients/Patients Talking with Doctors. (2006), 1–256

  50. [50]

    Stephanie Rude, Eva -Maria Gortner, and James Pennebaker. 2004. Language use of depressed and depression - vulnerable college students. Cogn. Emot. 18, 8 (December 2004), 1121 –1133. https://doi.org/10.1080/02699930441000030

  51. [51]

    Ruvini Sanjeewa, Ravi Iyer, Pragalathan Apputhurai, Nilmini Wickramasinghe, and Denny Meyer. 2024. Empathic Conversational Agent Platform Designs and Their Evaluation in the Context of Mental Health: Systematic Review. JMIR Ment. Health 11, 1 (September 2024), e58974. https://doi.org/10.2196/58974

  52. [52]

    Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. 2006. A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. Knowl. Eng. Rev. 21, 2 (June 2006), 97 –

  53. [53]

    Jim Zheng, and Kirk Roberts

    Yuqi Si, Jingcheng Du, Zhao Li, Xiaoqian Jiang, Timothy Miller, Fei Wang, W. Jim Zheng, and Kirk Roberts. 2021. Deep representation learning of patient data from Electronic Health Records (EHR): A systematic review. J. Biomed. Inform. 115, (March 2021), 103671. https://doi.org/10.1016/j.jbi.2020.103671

  54. [54]

    Scott Simpson and Anna McDowell. 2019. The Clinical Interview: Skills for More Effective Patient Encounters . Routledge, New York. https://doi.org/10.4324/9780429437243

  55. [55]

    Kristine Sørensen, Stephan Van den Broucke, James Fullam, Gerardine Doyle, Jürgen Pelikan, Zofia Slonska, Helmut Brand, and (HLS -EU) Consortium Health Literacy Project European. 2012. Health literacy and public health: a systematic review and integration of definitions and models. BMC Public Health 12, (January 2012), 80. https://doi.org/10.1186/1471-2458-12-80

  56. [56]

    Street, Gregory Makoul, Neeraj K

    Richard L. Street, Gregory Makoul, Neeraj K. Arora, and Ronald M. Epstein. 2009. How does communication heal? Pathways linking clinician –patient communication to health outcomes. Patient Educ. Couns. 74, 3 (March 2009), 295–301. https://doi.org/10.1016/j.pec.2008.11.015

  57. [57]

    Rehan Syed, Rebekah Eden, Tendai Makasi, Ignatius Chukwudi, Azumah Mamudu, Mostafa Kamalpour, Dakshi Kapugama Geeganage, Sareh Sadeghianasl, Sander J. J. Leemans, Kanika Goel, Robert Andrews, Moe Thandar Wynn, Arthur Ter Hofstede, and Trina Myers. 2023. Digital Health Data Quality Issues: Systematic Review. J. Med. Internet Res. 25, (March 2023), e42615. ...

  58. [58]

    Elham Tabassi. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST (January 2023). Retrieved October 28, 2025 from https://www.nist.gov/publications/artificial -intelligence-risk-management- framework-ai-rmf-10

  59. [59]

    Tausczik and James W

    Yla R. Tausczik and James W. Pennebaker. 2010. The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods. J. Lang. Soc. Psychol. 29, 1 (March 2010), 24 –54. https://doi.org/10.1177/0261927X09351676

  60. [60]

    Tara Templin, Sophia Fort, Prasanna Padmanabham, Pratyush Seshadri, Ram Rimal, Junier Oliva, Kristin Hassmiller Lich, Sean Sylvia, and Nasa Sinnott -Armstrong. 2025. Framework for bias evaluation in large language models in healthcare settings. Npj Digit. Med. 8, 1 (July 2025), 414. https://doi.org/10.1038/s41746-025-01786-w

  61. [61]

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nat. Med. 29, 8 (August 2023), 1930 –1940. https://doi.org/10.1038/s41591-023-02448-8

  62. [62]

    Sara Mahdavi, Christop her Semturs, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S

    Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, Elahe Vedadi, Nenad Tomasev, Shekoofeh Azizi, Karan Singhal, Le Hou, Albert Webson, Kavita Kulkarni, S. Sara Mahdavi, Christop her Semturs, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S. Corrado, Yossi Matias, Alan Karth...

  63. [63]

    Jason Walonoski, Mark Kramer, Joseph Nichols, Andre Quina, Chris Moesel, Dylan Hall, Carlton Duffett, Kudakwashe Dube, Thomas Gallagher, and Scott McLachlan. 2018. Synthea: An approach, method, and software mechanism for generating synthetic patients and t he synthetic electronic health care record. J. Am. Med. Inform. Assoc. 25, 3 (March 2018), 230–238. ...

  64. [64]

    Chiu, Jiayin Zhi, Shaun M

    Ruiyi Wang, Stephanie Milani, Jamie C. Chiu, Jiayin Zhi, Shaun M. Eack, Travis Labrum, Samuel M Murphy, Nev Jones, Kate V Hardy, Hong Shen, Fei Fang, and Zhiyu Chen. 2024. PATIENT -ψ: Using Large Language Models to Simulate Patients for Training Mental Hea lth Professionals. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Pr...

  65. [65]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain -of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22), November 28, 2022. Curran Associates Inc., Red Hook, NY,...

  66. [66]

    Janusz Wojtusiak. 2016. Towards Intelligent Patient Data Generator. Machine Learning and Inference Laboratory, George Mason University, Fairfax, VA. Retrieved March 18, 2026 from https://www.mli.gmu.edu/papers/2016/16 - 9.pdf

  67. [67]

    Wolf and Stacy Cooper Bailey

    Michael S. Wolf and Stacy Cooper Bailey. 2009. The Role of Health Literacy in Patient Safety. Role Health Lit. Patient Saf. (March 2009). Retrieved February 11, 2026 from https://psnet.ahrq.gov/perspective/role-health-literacy- patient-safety

  68. [68]

    Pfeffer, Jason Fries, and Nigam H

    Michael Wornow, Yizhe Xu, Rahul Thapa, Birju Patel, Ethan Steinberg, Scott Fleming, Michael A. Pfeffer, Jason Fries, and Nigam H. Shah. 2023. The shaky foundations of large language models and foundation models for electronic health records. NPJ Digit. Med. 6, (July 2023), 135. https://doi.org/10.1038/s41746-023-00879-8

  69. [69]

    Wynia and Chandra Y

    Matthew K. Wynia and Chandra Y. Osborn. 2010. Health Literacy and Communication Quality in Health Care Organizations. J. Health Commun. 15, Suppl 2 (2010), 102–115. https://doi.org/10.1080/10810730.2010.499981

  70. [70]

    Chao Yan, Ziqi Zhang, Steve Nyemba, and Zhuohang Li. 2024. Generating Synthetic Electronic Health Record Data Using Generative Adversarial Networks: Tutorial. JMIR AI 3, (April 2024), e52615. https://doi.org/10.2196/52615

  71. [71]

    Chi, Jilin Chen, and Alex Beutel

    Sirui Yao, Yoni Halpern, Nithum Thain, Xuezhi Wang, Kang Lee, Flavien Prost, Ed H. Chi, Jilin Chen, and Alex Beutel. 2021. Measuring Recommender System Effects with Simulated Users. https://doi.org/10.48550/arXiv.2101.04526

  72. [72]

    Ravi, Peter Brune, Fanyu Kong, Dave Anderson, George Lee, Arie Meir, Farhana Bandukwala, Elli Kanal, Sercan Ö Arık, and Tomas Pfister

    Jinsung Yoon, Michel Mizrahi, Nahid Farhady Ghalaty, Thomas Jarvinen, Ashwin S. Ravi, Peter Brune, Fanyu Kong, Dave Anderson, George Lee, Arie Meir, Farhana Bandukwala, Elli Kanal, Sercan Ö Arık, and Tomas Pfister. 2023. EHR-Safe: generating high-fidelity and privacy-preserving synthetic electronic health records. Npj Digit. Med. 6, 1 (August 2023), 141. ...

  73. [73]

    Danielle S

    Huizi Yu, Jiayan Zhou, Lingyao Li, Shan Chen, Jack Gallifant, Anye Shi, Xiang Li, Jingxian He, Wenyue Hua, Mingyu Jin, Guang Chen, Yang Zhou, Zhao Li, Trisha Gupte, Ming-Li Chen, Zahra Azizi, Yongfeng Zhang, Yanqiu Xing, Themistocles L. Danielle S. Bitter man, Themistocles L. Assimes, Xin Ma, Lin Lu, and Lizhou Fan. 2025. Simulated patient systems are int...

  74. [74]

    Taedong Yun, Eric Yang, Mustafa Safdari, Jong Ha Lee, Vaishnavi Vinod Kumar, S. Sara Mahdavi, Jonathan Amar, Derek Peyton, Reut Aharony, Andreas Michaelides PhD, Logan Douglas Schneider, Isaac Galatzer -Levy, Yugang Jia, John Canny, Arthur Gretton, and Maj a Mataric. 2025. Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for ...

  75. [75]

    Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou. 2023. LLMaAA: Making Large Language Models as Active Annotators. https://doi.org/10.48550/arXiv.2310.19596

  76. [76]

    Yinan Zhang, Xueqing Liu, and ChengXiang Zhai. 2017. Information Retrieval Evaluation as Search Simulation: A General Formal Framework for IR Evaluation. In Proceedings of the ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR ’17), October 01, 2017. Association for Computing Machinery, New York, NY, USA, 193–200. https://doi.org...

  77. [77]

    Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang, Hao Zhu, Ximing Lu, Frank Xu, Bill Yuchen Lin, Yejin Choi, Niloofar Mireshghallah, Ronan Le Bras, and Maarten Sap. 2025. HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions. https://doi.org/10.48550/arXiv.2409.16427

  78. [78]

    malexandersalazar/xlm -roberta-base-cls-depression · Hugging Face

    2023. malexandersalazar/xlm -roberta-base-cls-depression · Hugging Face. Retrieved February 11, 2026 from https://huggingface.co/malexandersalazar/xlm-roberta-base-cls-depression

  79. [79]

    Top 10 Most Common Antidepressants Dispensed in the U.S. Retrieved November 11, 2025 from https://www.definitivehc.com/resources/healthcare-insights/top-antidepressants-by-prescription-volume 9 APPENDICES 9.1 Patient Simulator We present the detailed system architecture and the core system prompt, which follows a Chain -of-thought driven reasoning mechani...

  80. [81]

    These indices are crucial for traceability and annotation, allowing evaluators or downstream models to link simulated responses back to precise clinical data points

    Medical Profile : The simulator uses a structured medical history represented by hierarchical indices (e.g., [2.3], [3.1]), where each code corresponds to a specific clinical fact or condition. These indices are crucial for traceability and annotation, allowing evaluators or downstream models to link simulated responses back to precise clinical data points

Showing first 80 references.