REVIEW 2 major objections 1 minor 33 references
PWP patient simulators conditioned on HEXACO traits match human actors in realism while preventing oversharing.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 21:13 UTC pith:FCWEAJ2K
load-bearing objection PWP adds HEXACO parametrization to patient simulators for better control over diversity and disclosure, backed by clinician ratings that beat baselines, but the realism claim rests on indirect judgments without real-patient trait data. the 2 major comments →
Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By grounding patient simulation in explicit HEXACO parametrization over a latent patient state, PWP enables fine-grained control over style, cooperativeness, and disclosure, resulting in responses judged nearly as realistic as human actors by clinicians while exhibiting wider behavioral variation and reduced oversharing compared to baselines.
What carries the argument
HEXACO personality parametrization over a latent patient state that controls conversational style and selective information disclosure
Load-bearing premise
The HEXACO personality model, when mapped to a latent patient state, produces conversational behaviors that accurately reflect the variability and selective disclosure patterns of real patients in clinical settings.
What would settle it
Compare disclosure rates and behavioral variability between PWP-simulated patients with specific HEXACO scores and actual patients with the same measured personality traits in matched clinical scenarios.
If this is right
- Clinicians rate PWP nearly as realistic as recorded human actors.
- PWP is flagged as "too informative" far less often than prior simulators.
- Configured HEXACO traits are recoverable by clinicians and an autorater.
- Personas span a substantially wider behavioral footprint than the closest baseline.
Where Pith is reading between the lines
- This framework could support large-scale benchmarking of clinical AI without recruiting real patients.
- Selective disclosure control might help simulate patients who withhold information until prompted, a common real-world pattern.
- Recoverability of traits suggests the model can be tuned for specific personality profiles in training scenarios.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PatientsWithPersonality (PWP), a framework for LLM-based patient simulation that parametrizes responses over a latent state using the six-dimensional HEXACO personality model to achieve controlled diversity, cooperativeness, and selective disclosure. It claims that clinician evaluations rate PWP nearly as realistic as recorded human actors and ahead of prior simulators, with configured HEXACO traits recoverable by clinicians and an autorater, a substantially wider behavioral footprint than baselines, and reduced oversharing.
Significance. If the evaluation results hold under fuller scrutiny, PWP could offer a practical, steerable simulator for scaling LLM benchmarking in clinical applications, addressing common limitations in realism and controllability of existing patient simulators. The grounding in an established personality inventory provides a principled mechanism for diversity that is a clear methodological strength.
major comments (2)
- [Evaluation] Evaluation section: the abstract reports positive clinician and autorater results (near-equivalence to human actors, trait recovery, wider footprint, less oversharing), but the manuscript provides no sample sizes, statistical tests, exclusion criteria, or inter-rater reliability metrics, preventing assessment of whether the comparative claims are robust.
- [Methods] Methods/Evaluation: the central claim that HEXACO conditioning produces behaviors matching real-patient variability and selective disclosure rests on clinician judgments and autorater recovery alone; no direct comparison is made to transcripts from actual patients who completed HEXACO inventories, leaving the mapping from trait axes to clinical conversational patterns unvalidated.
minor comments (1)
- [Abstract] Abstract: the phrase 'substantially wider behavioral footprint' is not quantified; a concrete metric (e.g., entropy over response categories or coverage of disclosure levels) should be stated.
Simulated Author's Rebuttal
We thank the referee for their constructive comments. We address each major comment below.
read point-by-point responses
-
Referee: [Evaluation] Evaluation section: the abstract reports positive clinician and autorater results (near-equivalence to human actors, trait recovery, wider footprint, less oversharing), but the manuscript provides no sample sizes, statistical tests, exclusion criteria, or inter-rater reliability metrics, preventing assessment of whether the comparative claims are robust.
Authors: We agree that these details are necessary to assess robustness and were omitted from the submitted manuscript. In the revised version we will report the exact sample sizes for clinician and autorater evaluations, the statistical tests performed (including p-values), any exclusion criteria, and inter-rater reliability metrics such as intraclass correlation coefficients. revision: yes
-
Referee: [Methods] Methods/Evaluation: the central claim that HEXACO conditioning produces behaviors matching real-patient variability and selective disclosure rests on clinician judgments and autorater recovery alone; no direct comparison is made to transcripts from actual patients who completed HEXACO inventories, leaving the mapping from trait axes to clinical conversational patterns unvalidated.
Authors: The manuscript's primary claims concern clinician-rated realism (near human actors), recoverability of the configured HEXACO traits, a wider behavioral range than baselines, and reduced oversharing. These are evaluated via expert clinician judgment and autorater analysis rather than a direct claim of equivalence to real-patient variability from HEXACO-inventoried transcripts. Clinician evaluation is the established standard for assessing simulation fidelity. We will add an explicit limitations paragraph discussing the evaluation design and noting that paired real-patient data would be a valuable direction for future work. revision: partial
Circularity Check
No circularity; claims rest on external clinician and autorater judgments
full rationale
The paper's central claims derive from clinician evaluations of realism, trait recoverability, behavioral range, and oversharing rates, plus an autorater, all applied to outputs generated from an external HEXACO model. These steps do not reduce by construction to the simulation parameters or any self-citation chain; the evaluations are independent measurements against human judges. No self-definitional mappings, fitted inputs renamed as predictions, or load-bearing self-citations appear in the provided derivation. The framework is therefore self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
free parameters (1)
- HEXACO trait weights and mappings
axioms (1)
- domain assumption HEXACO model captures relevant variability in patient behavior for clinical interactions
Cite this review
Pith. "Pith review of Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure." pith.science (2026). https://pith.science/paper/FCWEAJ2K
@misc{pith2026260617441,
author = {Pith},
title = {Pith review of: Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure},
year = {2026},
howpublished = {\url{https://pith.science/paper/FCWEAJ2K}},
note = {Machine review of arXiv:2606.17441}
}
read the original abstract
Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often oversharing information unprompted, and failing to capture the wide variability of patient behavior. Here, we introduce PatientsWithPersonality (PWP), a patient simulation framework that generates realistic yet diverse virtual patient responses through explicit personality parametrization over a latent patient state. Grounded in HEXACO, a six-dimensional personality space used to quantify and parameterize human behavioral traits, our approach enables fine-grained control over conversational style, cooperativeness, and information disclosure within a unified framework. In a clinician evaluation, PWP is judged nearly as realistic as recorded human actors and clearly ahead of prior simulators, while being flagged as "too informative" far less often. Conditioning on HEXACO axes yields personas whose configured traits are recoverable by both clinicians and an autorater, span a substantially wider behavioral footprint than the closest baseline, and prevent oversharing. Altogether, our framework paves the way for more accurate and informative LLM benchmarking through our realistic and steerable patient simulator.
Figures
Reference graph
Works this paper leans on
-
[1]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Senevi- ratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrad...
work page 2023
-
[2]
Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G. Nestor, Ali Soroush, Pierre A. Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, and Yifan Peng. Evaluating large language models on medical evidence summarization.npj Digital Medicine, 6(1):158, 2023
work page 2023
-
[3]
Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, and Daniel Rueckert. Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature Medicine, 30(9):2613–2622, 2024
work page 2024
-
[4]
Siun Kim and Hyung-Jin Yoon. Questioning Our Questions: How Well Do Medical QA Benchmarks Evaluate Clinical Capabilities of Language Models? In Dina Demner-Fushman, Sophia Ananiadou, Makoto Miwa, and Junichi Tsujii, editors,Proceedings of the 24th Workshop on Biomedical Language Processing, pages 274–296, Viena, Austria, 2025. Association for Computationa...
work page 2025
-
[5]
Eric Wu, Kevin Wu, Jason Hom, Paul H. Yi, Angela Zhang, Alejandro Lozano, Jeff Nirschl, Jeff Tangney, Kevin Byram, Braydon Dymm, Narender Annapureddy, Eric Topol, David Ouyang, and James Zou. MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences,
- [6]
-
[7]
Eun Jeong Gong, Chang Seok Bang, Jae Jun Lee, and Gwang Ho Baik. Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks. Journal of Medical Internet Research, 27(1):e84120, 2025
work page 2025
-
[8]
AgentClinic: a multimodal benchmark for tool-using clinical AI agents.npj Digital Medicine, 2026
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Ji Woong Kim, Eduardo Pontes Reis, Jeffrey Jopling, and Michael Moor. AgentClinic: a multimodal benchmark for tool-using clinical AI agents.npj Digital Medicine, 2026
work page 2026
-
[9]
Shreya Johri, Jaehwan Jeong, Benjamin A. Tran, Daniel I. Schlessinger, Shannon Wongvibulsin, Leandra A. Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M. Van Allen, David Kim, Roxana Daneshjou, and Pranav Rajpurkar. An evaluation framework for clinical use of large language models in patient interaction tasks.Nature Medicine, 31(1):77–86, 2025
work page 2025
-
[10]
Andrew M. Bean, Rebecca Elizabeth Payne, Guy Parsons, Hannah Rose Kirk, Juan Ciro, Rafael Mosquera-Gómez, Sara Hincapié M, Aruna S. Ekanayaka, Lionel Tarassenko, Luc Rocher, and 10 Adam Mahdi. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.Nature Medicine, 32(2):609–615, 2026
work page 2026
-
[11]
Friederike Holderried, Christian Stegemann-Philipps, Lea Herschbach, Julia-Astrid Moldt, Andrew Nevins, Jan Griewatz, Martin Holderried, Anne Herrmann-Werner, Teresa Festl-Wietek, and Moritz Mahling. A Generative Pretrained Transformer (GPT)-Powered Chatbot as a Simulated Patient to Practice History Taking: Prospective, Mixed Methods Study.JMIR medical ed...
work page 2024
-
[12]
David A Cook, Joshua Overgaard, V Shane Pankratz, Guilherme Del Fiol, and Chris A Aakre. Virtual Patients Using Large Language Models: Scalable, Contextualized Simula- tion of Clinician-Patient Dialogue With Feedback.Journal of Medical Internet Research, 27: e68486, 2025
work page 2025
-
[13]
Multi-Stage Patient Role-Playing Framework for Realistic Clinical Interactions, 2026
Shijie Jiang, Zefan Zhang, Kehua Zhu, Tian Bai, and Ruihong Zhao. Multi-Stage Patient Role-Playing Framework for Realistic Clinical Interactions, 2026. arXiv:2601.10951 [cs]
-
[14]
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors,Proceedings of the 31st International Conference on Comput...
work page 2025
-
[15]
LLMs Can Simulate Standardized Patients via Agent Coevolution
Zhuoyun Du, Lujie Zheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, and Haochao Ying. LLMs Can Simulate Standardized Patients via Agent Coevolution. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (...
work page 2025
-
[16]
Ping Wu, Yu Han, Jing Zhang, Yunqi Li, Mengna Jiang, Xinyu Lu, Haibin Zhang, Danyang Xu, Hao Ming, Lihong Wang, and Qingping Wen. From simulation to pedagogy: structured AI standardized patients for clinical communication training validated through multi-model and randomized evaluation, 2026. medRxiv 2026.04.26.26351793
work page 2026
-
[17]
Au- tomatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator,
Yusheng Liao, Yutong Meng, Yuhao Wang, Hongcheng Liu, Yanfeng Wang, and Yu Wang. Au- tomatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator,
- [18]
-
[19]
Taedong Yun, Eric Yang, Mustafa Safdari, Jong Ha Lee, Vaishnavi Vinod Kumar, S. Sara Mah- davi, Jonathan Amar, Derek Peyton, Reut Aharony, Andreas Michaelides PhD, Logan Douglas Schneider, Isaac Galatzer-Levy, Yugang Jia, John Canny, Arthur Gretton, and Maja Mataric. Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realis...
work page 2025
-
[20]
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Lei Zhou, Qianqian Xie, and Benyou Wang. Human or LLM as Standardized Patients? A Comparative Study for Medical Education. 2026. arXiv:2511.14783 [cs.CL]
-
[21]
PatientSim: A Persona-Driven Simulator for Realis- tic Doctor-Patient Interactions
Daeun Kyung, Hyunseung Chung, Seongsu Bae, Jiho Kim, Jae Ho Sohn, Taerim Kim, Soo Kyung Kim, and Edward Choi. PatientSim: A Persona-Driven Simulator for Realis- tic Doctor-Patient Interactions. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2026
work page 2026
-
[22]
Michael C. Ashton and Kibeom Lee. Empirical, theoretical, and practical advantages of the HEXACO model of personality structure.Personality and Social Psychology Review: An Official Journal of the Society for Personality and Social Psychology, Inc, 11(2):150–166, 2007
work page 2007
-
[23]
Sonia Roccas, Lilach Sagiv, Shalom H. Schwartz, and Ariel Knafo. The Big Five Personality Factors and Personal Values.Personality and Social Psychology Bulletin, 28(6):789–801, 2002
work page 2002
-
[24]
Jan Luca Pletzer, Isabel Thielmann, and Ingo Zettler. Who is healthier? A meta-analysis of the relations between the HEXACO personality domains and health outcomes.European Journal of Personality, 38(2):342–364, 2024. 11
work page 2024
-
[25]
Shuo Wang, Renhao Li, Xi Chen, Yulin Yuan, Min Yang, and Derek F. Wong. Exploring the Impact of Personality Traits on LLM Bias and Toxicity. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 4125–4143, Suzhou, China, 2025....
work page 2025
-
[26]
Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, and Meliha Yetisgen. Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation.Scientific Data, 10(1):586, 2023
work page 2023
-
[27]
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...
work page 2022
-
[28]
Large Language Models are Human-Level Prompt Engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large Language Models are Human-Level Prompt Engineers. InThe Eleventh International Conference on Learning Representations, 2023
work page 2023
-
[29]
Optimization before Evaluation: Evaluation with Unoptimized Prompts Can be Misleading
Nicholas Sadjoli, Tim Siefken, Atin Ghosh, Yifan Mai, and Daniel Dahlmeier. Optimization before Evaluation: Evaluation with Unoptimized Prompts Can be Misleading. In Georg Rehm and Yunyao Li, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), pages 619–638, Vienna, Austria, 2025. Ass...
work page 2025
-
[30]
J. P. Kincaid, Jr. Fishburne, Rogers Robert P., Chissom Richard L., and Brad S. Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel:. Technical report, Defense Technical Information Center, Fort Belvoir, V A, 1975
work page 1975
-
[31]
Dan Friedman and Adji Bousso Dieng. The Vendi Score: A Diversity Evaluation Metric for Machine Learning.Transactions on Machine Learning Research, 2023
work page 2023
-
[32]
Faiha Fareez, Tishya Parikh, Christopher Wavell, Saba Shahab, Meghan Chevalier, Scott Good, Isabella De Blasi, Rafik Rhouma, Christopher McMahon, Jean-Paul Lam, Thomas Lo, and Christopher W. Smith. A dataset of simulated patient-physician medical interviews with a focus on respiratory cases.Scientific Data, 9(1):313, 2022
work page 2022
-
[33]
Tom Zehle, Moritz Schlager, Timo Heiß, and Matthias Feurer. CAPO: Cost-Aware Prompt Optimization. InProceedings of the Fourth International Conference on Automated Machine Learning, pages 18/1–45. PMLR, 2025. 12 A Structured Case Description Fields Table 3 enumerates the named fields of the structured case description x used byPWP, together with their ass...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.