Pith. sign in

REVIEW 2 major objections 1 minor 33 references

PWP patient simulators conditioned on HEXACO traits match human actors in realism while preventing oversharing.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 21:13 UTC pith:FCWEAJ2K

load-bearing objection PWP adds HEXACO parametrization to patient simulators for better control over diversity and disclosure, backed by clinician ratings that beat baselines, but the realism claim rests on indirect judgments without real-patient trait data. the 2 major comments →

arxiv 2606.17441 v1 pith:FCWEAJ2K submitted 2026-05-13 cs.HC cs.AIcs.CY

Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

classification cs.HC cs.AIcs.CY
keywords patient simulationHEXACO personality modelvirtual patientsLLM benchmarkingclinical interactionspersonality parametrizationselective disclosure
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces PatientsWithPersonality (PWP), a framework that uses the HEXACO personality model to parametrize virtual patients in a latent state for controlled conversational behavior. It aims to provide realistic and diverse patient simulations for testing clinical LLMs, addressing issues like lack of realism and uncontrolled information disclosure in existing methods. Clinician evaluations show PWP approaches the realism of recorded human actors and outperforms prior simulators by being less likely to overshare. The approach allows recovery of configured traits and covers a broader range of behaviors.

Core claim

By grounding patient simulation in explicit HEXACO parametrization over a latent patient state, PWP enables fine-grained control over style, cooperativeness, and disclosure, resulting in responses judged nearly as realistic as human actors by clinicians while exhibiting wider behavioral variation and reduced oversharing compared to baselines.

What carries the argument

HEXACO personality parametrization over a latent patient state that controls conversational style and selective information disclosure

Load-bearing premise

The HEXACO personality model, when mapped to a latent patient state, produces conversational behaviors that accurately reflect the variability and selective disclosure patterns of real patients in clinical settings.

What would settle it

Compare disclosure rates and behavioral variability between PWP-simulated patients with specific HEXACO scores and actual patients with the same measured personality traits in matched clinical scenarios.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Clinicians rate PWP nearly as realistic as recorded human actors.
  • PWP is flagged as "too informative" far less often than prior simulators.
  • Configured HEXACO traits are recoverable by clinicians and an autorater.
  • Personas span a substantially wider behavioral footprint than the closest baseline.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • This framework could support large-scale benchmarking of clinical AI without recruiting real patients.
  • Selective disclosure control might help simulate patients who withhold information until prompted, a common real-world pattern.
  • Recoverability of traits suggests the model can be tuned for specific personality profiles in training scenarios.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces PatientsWithPersonality (PWP), a framework for LLM-based patient simulation that parametrizes responses over a latent state using the six-dimensional HEXACO personality model to achieve controlled diversity, cooperativeness, and selective disclosure. It claims that clinician evaluations rate PWP nearly as realistic as recorded human actors and ahead of prior simulators, with configured HEXACO traits recoverable by clinicians and an autorater, a substantially wider behavioral footprint than baselines, and reduced oversharing.

Significance. If the evaluation results hold under fuller scrutiny, PWP could offer a practical, steerable simulator for scaling LLM benchmarking in clinical applications, addressing common limitations in realism and controllability of existing patient simulators. The grounding in an established personality inventory provides a principled mechanism for diversity that is a clear methodological strength.

major comments (2)
  1. [Evaluation] Evaluation section: the abstract reports positive clinician and autorater results (near-equivalence to human actors, trait recovery, wider footprint, less oversharing), but the manuscript provides no sample sizes, statistical tests, exclusion criteria, or inter-rater reliability metrics, preventing assessment of whether the comparative claims are robust.
  2. [Methods] Methods/Evaluation: the central claim that HEXACO conditioning produces behaviors matching real-patient variability and selective disclosure rests on clinician judgments and autorater recovery alone; no direct comparison is made to transcripts from actual patients who completed HEXACO inventories, leaving the mapping from trait axes to clinical conversational patterns unvalidated.
minor comments (1)
  1. [Abstract] Abstract: the phrase 'substantially wider behavioral footprint' is not quantified; a concrete metric (e.g., entropy over response categories or coverage of disclosure levels) should be stated.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments. We address each major comment below.

read point-by-point responses
  1. Referee: [Evaluation] Evaluation section: the abstract reports positive clinician and autorater results (near-equivalence to human actors, trait recovery, wider footprint, less oversharing), but the manuscript provides no sample sizes, statistical tests, exclusion criteria, or inter-rater reliability metrics, preventing assessment of whether the comparative claims are robust.

    Authors: We agree that these details are necessary to assess robustness and were omitted from the submitted manuscript. In the revised version we will report the exact sample sizes for clinician and autorater evaluations, the statistical tests performed (including p-values), any exclusion criteria, and inter-rater reliability metrics such as intraclass correlation coefficients. revision: yes

  2. Referee: [Methods] Methods/Evaluation: the central claim that HEXACO conditioning produces behaviors matching real-patient variability and selective disclosure rests on clinician judgments and autorater recovery alone; no direct comparison is made to transcripts from actual patients who completed HEXACO inventories, leaving the mapping from trait axes to clinical conversational patterns unvalidated.

    Authors: The manuscript's primary claims concern clinician-rated realism (near human actors), recoverability of the configured HEXACO traits, a wider behavioral range than baselines, and reduced oversharing. These are evaluated via expert clinician judgment and autorater analysis rather than a direct claim of equivalence to real-patient variability from HEXACO-inventoried transcripts. Clinician evaluation is the established standard for assessing simulation fidelity. We will add an explicit limitations paragraph discussing the evaluation design and noting that paired real-patient data would be a valuable direction for future work. revision: partial

Circularity Check

0 steps flagged

No circularity; claims rest on external clinician and autorater judgments

full rationale

The paper's central claims derive from clinician evaluations of realism, trait recoverability, behavioral range, and oversharing rates, plus an autorater, all applied to outputs generated from an external HEXACO model. These steps do not reduce by construction to the simulation parameters or any self-citation chain; the evaluations are independent measurements against human judges. No self-definitional mappings, fitted inputs renamed as predictions, or load-bearing self-citations appear in the provided derivation. The framework is therefore self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

1 free parameters · 1 axioms · 0 invented entities

The approach depends on the applicability of the HEXACO model to clinical patient behavior and the assumption that trait parametrization translates into realistic conversational outputs without additional unstated mappings.

free parameters (1)
  • HEXACO trait weights and mappings
    Personality dimensions are used to control responses, implying parameters that shape style, cooperativeness, and disclosure levels.
axioms (1)
  • domain assumption HEXACO model captures relevant variability in patient behavior for clinical interactions
    Framework is explicitly grounded in HEXACO without additional justification in the abstract.

pith-pipeline@v0.9.1-grok · 5798 in / 1248 out tokens · 34264 ms · 2026-06-30T21:13:06.472982+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure." pith.science (2026). https://pith.science/paper/FCWEAJ2K

@misc{pith2026260617441,
  author       = {Pith},
  title        = {Pith review of: Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FCWEAJ2K}},
  note         = {Machine review of arXiv:2606.17441}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often oversharing information unprompted, and failing to capture the wide variability of patient behavior. Here, we introduce PatientsWithPersonality (PWP), a patient simulation framework that generates realistic yet diverse virtual patient responses through explicit personality parametrization over a latent patient state. Grounded in HEXACO, a six-dimensional personality space used to quantify and parameterize human behavioral traits, our approach enables fine-grained control over conversational style, cooperativeness, and information disclosure within a unified framework. In a clinician evaluation, PWP is judged nearly as realistic as recorded human actors and clearly ahead of prior simulators, while being flagged as "too informative" far less often. Conditioning on HEXACO axes yields personas whose configured traits are recoverable by both clinicians and an autorater, span a substantially wider behavioral footprint than the closest baseline, and prevent oversharing. Altogether, our framework paves the way for more accurate and informative LLM benchmarking through our realistic and steerable patient simulator.

Figures

Figures reproduced from arXiv: 2606.17441 by Avinatan Hassidim, Conrad Ketzer, Dale R. Webster, Daniel Rueckert, Eva Wende, Franziska Hartl, Friederike Jungmann, Mike Schaekermann, Moritz Schlager, Paula Ro{\ss}m\"uller, Paul Hager, Philipp Raffler, Samuel Schmidgall, Yossi Matias, Yun Liu.

Figure 1
Figure 1. Figure 1: Our PatientsWithPersonality framework resolves failure modes of current patient simula￾tors by selectively controlling information disclo￾sure and by allowing to simulate diverse characters grounded in the HEXACO personality space. However, the lack of realism and behavioral fi￾delity of the simulated patients often distorts benchmark performance. This has direct con￾sequences for benchmark validity: in a … view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the PWP framework. During initialization, the personality parametrization and structured case description are transformed into a latent patient state by a meta-LLM. During response generation, the requested case fields are extracted from each incoming question and used to update the disclosure grid. The conversational LLM then generates the patient answer based on the latent role, the currently… view at source ↗
Figure 3
Figure 3. Figure 3: Left: Clinicians judge PatientsWithPersonality conversations as real at similar rates as the original recorded encounter and well above the rates of other simulators. Right: Other simulators are flagged as “too informative” about twice as often as PatientsWithPersonality. 6 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Per-conversation differences (sim − real) from the recorded patient track. Values closer to zero indicate closer agreement with the recorded transcripts. PatientsWithPersonality is closest to the original recorded encounter across most metrics and especially on lexical diversity and token count, driving its realism scores. Finding: PWP tracks the recorded encounters most closely, with significantly tighter… view at source ↗
Figure 5
Figure 5. Figure 5: Cumulative fraction of case fields dis￾closed, averaged over conversations. PatientsWith￾Personality discloses at the same rate as both the original Human Actor and the Human Rephrase. Disclosure behavior. To measure information sharing behavior, we mark a case field as dis￾closed at the first turn at which the simulator’s ut￾terance contains content matching that field. We then calculate the cumulative fr… view at source ↗
Figure 6
Figure 6. Figure 6: Per-axis HEXACO reconstruction by the autorater versus clinician annotators on the personality task transcripts. Human and autorater agreed closely across both frameworks achieving r = 0.87. Dashed line is the identity line. To then characterize the behavioral footprint of each configuration, we embed every patient utterance with a sentence encoder and project the embeddings into a shared 2D PCA space esti… view at source ↗
Figure 7
Figure 7. Figure 7: Sentence-embedding PCA projection of patient utterances. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Reward of every prompt at the end of each step. The trajectory is colored by prompt length; [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Mean information penalty versus mean personality penalty on [PITH_FULL_IMAGE:figures/full_fig_p019_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Per-axis realism subscores from the clinician realism task. Mean ratings on a 1–5 scale [PITH_FULL_IMAGE:figures/full_fig_p022_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Inter-doctor agreement for the realism task. [PITH_FULL_IMAGE:figures/full_fig_p022_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 33 canonical work pages

  1. [1]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Senevi- ratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrad...

  2. [2]

    Nestor, Ali Soroush, Pierre A

    Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G. Nestor, Ali Soroush, Pierre A. Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, and Yifan Peng. Evaluating large language models on medical evidence summarization.npj Digital Medicine, 6(1):158, 2023

  3. [3]

    Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature Medicine, 30(9):2613–2622, 2024

    Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, and Daniel Rueckert. Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature Medicine, 30(9):2613–2622, 2024

  4. [4]

    Siun Kim and Hyung-Jin Yoon. Questioning Our Questions: How Well Do Medical QA Benchmarks Evaluate Clinical Capabilities of Language Models? In Dina Demner-Fushman, Sophia Ananiadou, Makoto Miwa, and Junichi Tsujii, editors,Proceedings of the 24th Workshop on Biomedical Language Processing, pages 274–296, Viena, Austria, 2025. Association for Computationa...

  5. [5]

    Yi, Angela Zhang, Alejandro Lozano, Jeff Nirschl, Jeff Tangney, Kevin Byram, Braydon Dymm, Narender Annapureddy, Eric Topol, David Ouyang, and James Zou

    Eric Wu, Kevin Wu, Jason Hom, Paul H. Yi, Angela Zhang, Alejandro Lozano, Jeff Nirschl, Jeff Tangney, Kevin Byram, Braydon Dymm, Narender Annapureddy, Eric Topol, David Ouyang, and James Zou. MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences,

  6. [6]

    arXiv:2603.15677 [cs]

  7. [7]

    Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks

    Eun Jeong Gong, Chang Seok Bang, Jae Jun Lee, and Gwang Ho Baik. Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks. Journal of Medical Internet Research, 27(1):e84120, 2025

  8. [8]

    AgentClinic: a multimodal benchmark for tool-using clinical AI agents.npj Digital Medicine, 2026

    Samuel Schmidgall, Rojin Ziaei, Carl Harris, Ji Woong Kim, Eduardo Pontes Reis, Jeffrey Jopling, and Michael Moor. AgentClinic: a multimodal benchmark for tool-using clinical AI agents.npj Digital Medicine, 2026

  9. [9]

    Tran, Daniel I

    Shreya Johri, Jaehwan Jeong, Benjamin A. Tran, Daniel I. Schlessinger, Shannon Wongvibulsin, Leandra A. Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M. Van Allen, David Kim, Roxana Daneshjou, and Pranav Rajpurkar. An evaluation framework for clinical use of large language models in patient interaction tasks.Nature Medicine, 31(1):77–86, 2025

  10. [10]

    Bean, Rebecca Elizabeth Payne, Guy Parsons, Hannah Rose Kirk, Juan Ciro, Rafael Mosquera-Gómez, Sara Hincapié M, Aruna S

    Andrew M. Bean, Rebecca Elizabeth Payne, Guy Parsons, Hannah Rose Kirk, Juan Ciro, Rafael Mosquera-Gómez, Sara Hincapié M, Aruna S. Ekanayaka, Lionel Tarassenko, Luc Rocher, and 10 Adam Mahdi. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.Nature Medicine, 32(2):609–615, 2026

  11. [11]

    Friederike Holderried, Christian Stegemann-Philipps, Lea Herschbach, Julia-Astrid Moldt, Andrew Nevins, Jan Griewatz, Martin Holderried, Anne Herrmann-Werner, Teresa Festl-Wietek, and Moritz Mahling. A Generative Pretrained Transformer (GPT)-Powered Chatbot as a Simulated Patient to Practice History Taking: Prospective, Mixed Methods Study.JMIR medical ed...

  12. [12]

    David A Cook, Joshua Overgaard, V Shane Pankratz, Guilherme Del Fiol, and Chris A Aakre. Virtual Patients Using Large Language Models: Scalable, Contextualized Simula- tion of Clinician-Patient Dialogue With Feedback.Journal of Medical Internet Research, 27: e68486, 2025

  13. [13]

    Multi-Stage Patient Role-Playing Framework for Realistic Clinical Interactions, 2026

    Shijie Jiang, Zefan Zhang, Kehua Zhu, Tian Bai, and Ruihong Zhao. Multi-Stage Patient Role-Playing Framework for Realistic Clinical Interactions, 2026. arXiv:2601.10951 [cs]

  14. [14]

    AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator

    Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors,Proceedings of the 31st International Conference on Comput...

  15. [15]

    LLMs Can Simulate Standardized Patients via Agent Coevolution

    Zhuoyun Du, Lujie Zheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, and Haochao Ying. LLMs Can Simulate Standardized Patients via Agent Coevolution. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (...

  16. [16]

    From simulation to pedagogy: structured AI standardized patients for clinical communication training validated through multi-model and randomized evaluation, 2026

    Ping Wu, Yu Han, Jing Zhang, Yunqi Li, Mengna Jiang, Xinyu Lu, Haibin Zhang, Danyang Xu, Hao Ming, Lihong Wang, and Qingping Wen. From simulation to pedagogy: structured AI standardized patients for clinical communication training validated through multi-model and randomized evaluation, 2026. medRxiv 2026.04.26.26351793

  17. [17]

    Au- tomatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator,

    Yusheng Liao, Yutong Meng, Yuhao Wang, Hongcheng Liu, Yanfeng Wang, and Yu Wang. Au- tomatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator,

  18. [18]

    arXiv:2403.08495 [cs.CL]

  19. [19]

    Taedong Yun, Eric Yang, Mustafa Safdari, Jong Ha Lee, Vaishnavi Vinod Kumar, S. Sara Mah- davi, Jonathan Amar, Derek Peyton, Reut Aharony, Andreas Michaelides PhD, Logan Douglas Schneider, Isaac Galatzer-Levy, Yugang Jia, John Canny, Arthur Gretton, and Maja Mataric. Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realis...

  20. [20]

    Human or LLM as Standardized Patients? A Comparative Study for Medical Education

    Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Lei Zhou, Qianqian Xie, and Benyou Wang. Human or LLM as Standardized Patients? A Comparative Study for Medical Education. 2026. arXiv:2511.14783 [cs.CL]

  21. [21]

    PatientSim: A Persona-Driven Simulator for Realis- tic Doctor-Patient Interactions

    Daeun Kyung, Hyunseung Chung, Seongsu Bae, Jiho Kim, Jae Ho Sohn, Taerim Kim, Soo Kyung Kim, and Edward Choi. PatientSim: A Persona-Driven Simulator for Realis- tic Doctor-Patient Interactions. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2026

  22. [22]

    Ashton and Kibeom Lee

    Michael C. Ashton and Kibeom Lee. Empirical, theoretical, and practical advantages of the HEXACO model of personality structure.Personality and Social Psychology Review: An Official Journal of the Society for Personality and Social Psychology, Inc, 11(2):150–166, 2007

  23. [23]

    Schwartz, and Ariel Knafo

    Sonia Roccas, Lilach Sagiv, Shalom H. Schwartz, and Ariel Knafo. The Big Five Personality Factors and Personal Values.Personality and Social Psychology Bulletin, 28(6):789–801, 2002

  24. [24]

    Who is healthier? A meta-analysis of the relations between the HEXACO personality domains and health outcomes.European Journal of Personality, 38(2):342–364, 2024

    Jan Luca Pletzer, Isabel Thielmann, and Ingo Zettler. Who is healthier? A meta-analysis of the relations between the HEXACO personality domains and health outcomes.European Journal of Personality, 38(2):342–364, 2024. 11

  25. [25]

    Shuo Wang, Renhao Li, Xi Chen, Yulin Yuan, Min Yang, and Derek F. Wong. Exploring the Impact of Personality Traits on LLM Bias and Toxicity. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 4125–4143, Suzhou, China, 2025....

  26. [26]

    Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation.Scientific Data, 10(1):586, 2023

    Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, and Meliha Yetisgen. Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation.Scientific Data, 10(1):586, 2023

  27. [27]

    Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

    Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...

  28. [28]

    Large Language Models are Human-Level Prompt Engineers

    Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large Language Models are Human-Level Prompt Engineers. InThe Eleventh International Conference on Learning Representations, 2023

  29. [29]

    Optimization before Evaluation: Evaluation with Unoptimized Prompts Can be Misleading

    Nicholas Sadjoli, Tim Siefken, Atin Ghosh, Yifan Mai, and Daniel Dahlmeier. Optimization before Evaluation: Evaluation with Unoptimized Prompts Can be Misleading. In Georg Rehm and Yunyao Li, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), pages 619–638, Vienna, Austria, 2025. Ass...

  30. [30]

    J. P. Kincaid, Jr. Fishburne, Rogers Robert P., Chissom Richard L., and Brad S. Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel:. Technical report, Defense Technical Information Center, Fort Belvoir, V A, 1975

  31. [31]

    The Vendi Score: A Diversity Evaluation Metric for Machine Learning.Transactions on Machine Learning Research, 2023

    Dan Friedman and Adji Bousso Dieng. The Vendi Score: A Diversity Evaluation Metric for Machine Learning.Transactions on Machine Learning Research, 2023

  32. [32]

    Faiha Fareez, Tishya Parikh, Christopher Wavell, Saba Shahab, Meghan Chevalier, Scott Good, Isabella De Blasi, Rafik Rhouma, Christopher McMahon, Jean-Paul Lam, Thomas Lo, and Christopher W. Smith. A dataset of simulated patient-physician medical interviews with a focus on respiratory cases.Scientific Data, 9(1):313, 2022

  33. [33]

    <chiefcomplaint>

    Tom Zehle, Moritz Schlager, Timo Heiß, and Matthias Feurer. CAPO: Cost-Aware Prompt Optimization. InProceedings of the Fourth International Conference on Automated Machine Learning, pages 18/1–45. PMLR, 2025. 12 A Structured Case Description Fields Table 3 enumerates the named fields of the structured case description x used byPWP, together with their ass...