Pith. sign in

REVIEW 5 major objections 5 minor 62 references

Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Persona-prompted GPT-4 matches high-education patients on discharge-summary questions, but not female or low-visit patients.

desk verdict A useful pilot on LLM persona simulation for health communication, but the headline 88% claim is not supported by the undefined alignment metric. read the letter →

arxiv 2501.06964 v1 pith:AIFCKOIF submitted 2025-01-12 cs.AI cs.HC

classification cs.AIcs.HC
keywords largelanguagemodelsrole-playingpersonasimulationdischargesummariespatientcomprehensionhealthcommunicationGPT-4alignmentevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a large language model can stand in for real patients when testing whether discharge summaries are understandable. The authors prompt GPT-4 with personas such as "someone who has never received a college degree" or "someone whose gender is female," ask it the same ten comprehension questions given to 96 human participants, and compare answer patterns. They find that education-primed personas align with human answers at an average of 77.5%, with high-education information questions reaching 88%, while female, low-doctor-visit, and low-ER-visit personas align at 31–47%, near random chance. The point is to decide whether LLM role-play can cheaply replace user studies for patient-centered health communication, and the answer is: for some groups yes, for others no, and adding demographic context can make things worse.

What carries the argument

The intervention is in-context impersonation: the authors prepend "If you were a {persona}" to a fixed task instruction, and the model answers multiple-choice questions by letter only. Personas are built from four demographic attributes (education, gender, doctor-visit frequency, ER-visit frequency) split into high and low. The evaluation machinery is the reported alignment rate, which compares the persona-prompted LLM's letter choice with the dominant answer of the corresponding human group on the same questions, reported separately for eight information-based questions and two perception-based questions, with a random-guess baseline for comparison.

What would settle it

Re-score the survey using full answer distributions rather than modal agreement: for each of the ten questions, record the LLM's letter distribution under the "high education" prompt, the "low education" prompt, and a no-persona prompt, and compare them with the human group distributions. The claimed persona effect is falsified if the LLM's distribution barely moves across personas on questions where human groups differ, or if the no-persona prompt matches a group as well as that group's persona prompt does; the easiest check is Q2 and Q3, where most humans answer "A. Yes" and the model might do the same regardless of persona.

Watch

Extended reading notes

Core claim

The authors set out to establish that LLM-driven personas can reproduce how real patients answer questions about discharge summaries, and to map where that reproduction fails. Their central quantitative finding is that the model's alignment with human answers depends strongly on the persona and the task type: education-based personas average 77.5% alignment, high-education information questions hit 88%, and male information questions hit 97%, while female, low-doctor-visit, and low-ER-visit personas fall to 31–47%. The paper also finds that the model oversimplifies: it concentrates answers on one or two options, never chooses "I don't know," and overestimates comprehension of long summaries. The authors read these results as evidence that in-context impersonation is promising for self-similar groups but not ready for diverse patient populations, and that adding patient-specific context can reduce rather than improve accuracy.

Load-bearing premise

The whole evaluation assumes that agreeing with the most popular answer in a human group demonstrates that the model is simulating that group's perspective; if the agreement comes from the model's generic preference for common answers, the persona scores do not measure what they claim to.

Editorial extensions

If this is right

  • Education-primed LLM personas could serve as a low-cost first-pass screening tool for how well-educated patients understand discharge instructions.
  • Automated patient communication systems should not yet use LLM personas to represent female, low doctor-visit, or low ER-visit populations, because alignment there is near random.
  • Including demographic details beyond education in a prompt can hurt performance, so a simpler query-response format may be safer for generating patient-facing health text.
  • Because the model overestimates comprehension of long summaries and never expresses uncertainty, any LLM-generated discharge material needs human review before release.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The alignment metric may be rewarding base-rate matching: on questions where almost all humans choose "A. Yes," a model that prefers "A" scores high for every persona; comparing full answer distributions would separate persona-specific simulation from default answer tendencies.
  • A direct test of the mechanism would be to run the same prompts with no persona and with several personas; if the no-persona distribution matches high-education humans as well as the high-education persona does, the persona prefix is not doing the work.
  • The failure to ever answer "I don't know" is a plausible cause of both the inflated male and high-education scores and the depressed female and low-visit scores; adding a calibration step that permits uncertainty may change which groups look well simulated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper evaluates whether GPT-4 can simulate the behavior of patient groups when answering questions about ICU discharge summaries. The authors recruited 96 Prolific participants, collected answers to eight information-based and two perception-based questions across four discharge summaries, and compared these with GPT-4 responses generated under personas defined by education, gender, doctor-visit frequency, and ER-visit frequency. They report alignment rates per persona (Table 2) and use them to conclude that education-primed personas achieve high alignment (88% for high-education information tasks), that performance drops for female and low-ER-visit personas, and that adding other information can reduce performance below random guessing. The paper frames this as a first step toward using LLMs as proxies for human subjects in personalized health communication.

Significance. If the quantitative claims were well supported, this would be a useful preliminary contribution to the growing literature on LLM role-playing and to debates about replacing human participants with LLM simulations. The external human-response data, IRB oversight, randomized discharge summaries, and transparent reporting of the per-persona percentages are genuine strengths, as is the honest qualitative discussion of overconfidence (Section 4.5) and of the risk of 'misportray[ing]' identity groups (Section 5). However, the paper's central quantitative claims currently rest on an undefined similarity metric, a weak random-guess baseline, and the absence of any no-persona control; until those are addressed, the alignment percentages cannot be interpreted as evidence for persona-simulation capability.

major comments (5)
  1. [§4.2, Table 2] The paper never defines the 'similarity'/'alignment rate' between the LLM response and human responses. It is not stated whether the comparison is to the modal human answer for each question, to the full human response distribution (e.g., via a distributional metric), or to a per-question pooled statistic. Without this definition, the headline numbers (0.88, 0.97, 0.31) and the aggregate 'average alignment rate of 54.97%' are uninterpretable, and the central claim that the LLM 'simulates' these personas is not falsifiable from the table.
  2. [Abstract and §4.2] The abstract claims that education-primed LLMs 'deliver accurate and actionable medical guidance 88% of the time,' but Table 2 shows 0.88 under 'High Edu.' for the information-based task, which is an alignment rate with human answers, not a measure of medical accuracy or actionability. Moreover, the abstract's claim that performance 'falls below random chance levels' is not generally true: in the information tasks every persona row is above the reported random-guess value (0.278), and in the perception tasks only the Male persona (0.25) falls below 0.267. The wording must be corrected to describe alignment with modal human responses and to give the precise exception.
  3. [§4.2] The random-guess baseline is not a meaningful control for skewed answer distributions. The reported baseline assumes uniform random choice among the answer options, but the human responses are highly skewed (e.g., Figure 4 shows a strong majority choosing 'A. Yes' for Q2-Q4). A majority-class baseline or a chance level conditional on the observed marginal answer distribution would be far higher than 0.278 for such questions. The paper also lacks a no-persona control condition: without asking the same questions without persona priming, or with a neutral persona, the high alignment of, say, the Male persona (0.97) may simply reflect the LLM's default tendency to choose common answers, a possibility the paper itself acknowledges in Section 5 when it notes that LLM responses are 'inherently identical or leptokurtic.'
  4. [§4.1 and §4.2] No confidence intervals, significance tests, or sampling details are provided for the LLM responses. Section 4.1 states only that GPT4_0613 was used via the Azure OpenAI API; the number of repeated samples per prompt, temperature, and number of questions per condition are not given. Section 4.2 uses the phrase 'significantly higher alignment rate' without any statistical test, and the differences between personas (e.g., High Edu. 0.88 vs. Low Edu. 0.72) are single point estimates from 96 human respondents split into multiple subgroups. These omissions make it impossible to judge whether any of the reported differences are robust.
  5. [Abstract and §7] The claim that 'a straightforward query-response model could outperform a more tailored approach' is not supported by any experiment reported in the paper. The evaluation only compares persona-primed prompts to random guessing; there is no non-persona or 'straightforward' baseline in Table 2 or elsewhere. If the authors have such data, it should be reported; otherwise this conclusion should be removed or reframed as a hypothesis.
minor comments (5)
  1. [Section 5] The terms 'LMM' and 'LMMs' appear several times (e.g., 'failure modes of LMM'); these should be 'LLM'/'LLMs'.
  2. [Appendix C.1] Q5 is labeled 'Do you know your diagnosis?' but the same label is used for Q3; the question text appears to be about other prescriptions, so the label is likely a copy-paste error.
  3. [Appendix C.1] Q8's option A reads 'A void fruit' and should read 'Avoid fruit'.
  4. [Table 2 and §4.2] The terms 'similarity' and 'alignment rate' are used interchangeably; they should be unified and formally defined.
  5. [References] The sentence 'Significant gender biases in LLM-generated content have been analyzed by Wang et al. in [47]' would benefit from a brief summary of the relevant finding, since the reference is to a general preprint.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: LLM outputs are compared against independently collected human survey responses, and no fitted parameter, self-citation chain, or by-construction identity defines the reported alignment rates.

full rationale

The paper's central evaluation is an external comparison: human participants recruited through Prolific answered questions about discharge summaries, and GPT-4 was prompted with persona descriptions and asked the same questions. The reported 'alignment rate' is a similarity between LLM-generated answers and these pre-existing human answers. There is no fitted parameter derived from the human data that is then used to define the target, no training or fine-tuning on the human responses, and no equation in which the prediction reduces to its own input. The random-guess baseline is an independent reference point, and the paper's own discussion of misalignment (e.g., LLMs never choosing 'I don't know') is based on observed distribution differences rather than on a self-referential construction. The abstract's phrase 'accurate and actionable medical guidance 88% of the time' overstates what the 0.88 alignment figure means, and the alignment metric itself is not formally defined in the paper; these are reporting and validity concerns, not circularity. There is also no load-bearing self-citation: the cited related work on role-playing and persona simulation is background context rather than a premise that forces the paper's results. Therefore, the derivation chain is self-contained with respect to the external human benchmark, and no circular step meets the evidentiary standard required for a positive finding.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper's claims rest on the external human survey as ground truth, on the validity of the comprehension questions, and on the stability of GPT-4 sampling. No fitted parameters or invented entities are introduced, but several methodological choices (sample count, metric definition, baseline computation) are unstated.

free parameters (1)
  • LLM samples per prompt = not reported
    Figures 3-5 show LLM response distributions, which require multiple generations per prompt, but the sample count and temperature are not stated. This choice affects the stability of the reported alignment rates.
assumptions (5)
  • domain assumption Human self-reported demographic categories on Prolific are accurate labels for ground-truth groups.
    Section 4.1 uses Prolific recruitment and self-reported education, gender, doctor visit, and ER visit frequency to define the human groups the LLM is compared against.
  • domain assumption The ten survey questions validly measure comprehension of discharge summaries.
    Section 3.2 states the questions assess understanding and recall, but no validation or pilot results are given.
  • domain assumption Multiple LLM generations approximate a persona's response distribution.
    Figures 3-5 plot LLM answer distributions, which requires sampling; the paper assumes these samples characterize the persona despite noting LLM responses are 'identical or leptokurtic' in Section 5.
  • standard math The random guess baseline is computed correctly and serves as a meaningful chance level.
    Table 2 reports 0.278 and 0.267 for information and perception questions; the computation is not shown, and the option counts vary by question.
  • domain assumption GPT-4's behavior at inference time is stable enough for comparison across personas.
    The study uses one model snapshot (GPT4_0613) but does not control for API nondeterminism or drift.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives." pith.science (2026). https://pith.science/paper/AIFCKOIF

@misc{pith2026250106964,
  author       = {Pith},
  title        = {Pith review of: Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AIFCKOIF}},
  note         = {Machine review of arXiv:2501.06964}
}
read the original abstract

Large Language Models (LLMs) have demonstrated impressive capabilities in role-playing scenarios, particularly in simulating domain-specific experts using tailored prompts. This ability enables LLMs to adopt the persona of individuals with specific backgrounds, offering a cost-effective and efficient alternative to traditional, resource-intensive user studies. By mimicking human behavior, LLMs can anticipate responses based on concrete demographic or professional profiles. In this paper, we evaluate the effectiveness of LLMs in simulating individuals with diverse backgrounds and analyze the consistency of these simulated behaviors compared to real-world outcomes. In particular, we explore the potential of LLMs to interpret and respond to discharge summaries provided to patients leaving the Intensive Care Unit (ICU). We evaluate and compare with human responses the comprehensibility of discharge summaries among individuals with varying educational backgrounds, using this analysis to assess the strengths and limitations of LLM-driven simulations. Notably, when LLMs are primed with educational background information, they deliver accurate and actionable medical guidance 88% of the time. However, when other information is provided, performance significantly drops, falling below random chance levels. This preliminary study shows the potential benefits and pitfalls of automatically generating patient-specific health information from diverse populations. While LLMs show promise in simulating health personas, our results highlight critical gaps that must be addressed before they can be reliably used in clinical settings. Our findings suggest that a straightforward query-response model could outperform a more tailored approach in delivering health information. This is a crucial first step in understanding how LLMs can be optimized for personalized health communication while maintaining accuracy.

Figures

Figures reproduced from arXiv: 2501.06964 by the authors.

Figure 1
Figure 1. Four discharge summaries on different categories. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Effectiveness of the LLM’s Ability to Simulate Personas Across Different Discharge Summary Categories [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Results distribution for perception-based questions for DS1, DS2, DS3 and DS4. “A” to “E” represents [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Examples of Answers distributions of Selected Questions. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Examples of Answers distributions of Selected Questions. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 35 canonical work pages

  1. [1]

    Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society . 298–306

  2. [2]

    Gati Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2022. Using Large Language Models to Simulate Multiple Humans. arXiv:2208.10264 (2022)

  3. [3]

    Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

    Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...

  4. [4]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis (2023)

  5. [5]

    Simran Arora, Avanika Narayan, Mayee F Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, and Christopher Re

  6. [6]

    Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, ...

  7. [7]

    Francoise Baylis. 1996. Women and health research: working for change. The Journal of Clinical Ethics 7, 3 (1996), 229–242

  8. [8]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In ACM FAccT

Show all 62 references
  1. [9]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mo- hammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A Suite for Analyzing Lar...

  2. [10]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  3. [11]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  4. [12]

    Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017)

  5. [13]

    Akhil Chintalapati, Shaurya Agarwal, Aparajita Senapati, Mohammed Misbah, RS Nithyashree, and Sunil Kumar

  6. [14]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  7. [15]

    In 2023 International Conference on Next Generation Electronics (NEleX)

    Textual Alchemy: Transmuting BERT Summaries with GAN Elegance. In 2023 International Conference on Next Generation Electronics (NEleX). IEEE, 1–7. , Vol. 1, No. 1, Article . Publication date: January 2025. 16 Xinyao Ma, Rui Zhu, Zihao Wang, Jingwei Xiong, Qingyu Chen, Haixu Ta...

  8. [16]

    Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz. 2023. Inducing anxiety in large language models increases exploration and bias. arXiv:2304.11111 (2023)

  9. [17]

    Zhao, Yanping Huang, Andrew M

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa De- hghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vince...

  10. [18]

    Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M Johnson. 2024. EvaluLLM: LLM assisted evaluation of generative outputs. In Companion Proceedings of the 29th International Conference on Intelligent User Interfaces. 30–32

  11. [19]

    Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv:2304.05335 (2023)

  12. [20]

    Seungju Han, Beomsu Kim, Jin Yong Yoo, Seokjun Seo, Sangbum Kim, Enkhbayar Erdenee, and Buru Chang. 2022. Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a Few Utterances. In NAACL-HLT

  13. [21]

    Stacie E Geller, Abby Koch, Beth Pellettieri, and Molly Carnes. 2011. Inclusion, analysis, and reporting of sex and race/ethnicity in clinical trials: have we made progress? Journal of women’s health 20, 3 (2011), 315–320

  14. [22]

    Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan. 2023. Causal Reasoning and Large Language Models: Opening a New Frontier for Causality. arXiv:2305.00050 (2023)

  15. [23]

    Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv:1909.05858 (2019)

  16. [24]

    Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong

  17. [25]

    Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. 2024. Understanding users’ dissatisfaction with chatgpt responses: Types, resolving tactics, and the effect of knowledge level. In Proceedings of the 29th International Conference on Intelligent User Interfaces . 385–404

  18. [26]

    Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill. 2022. Can language models learn from explanations in context?. In EMNLP. ACL

  19. [27]

    Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing S...

  20. [28]

    Gleb Kumichev, Pavel Blinov, Yulia Kuzkina, Vasily Goncharov, Galina Zubkova, Nikolai Zenovkin, Aleksei Goncharov, and Andrey Savchenko. 2024. MedSyn: LLM-Based Synthetic Medical Text Generation Framework. In Joint European Conference on Machine Learning and Knowledge Discover...

  21. [29]

    Juan Antonio Lossio-Ventura, Rachel Weger, Angela Y Lee, Emily P Guinee, Joyce Chung, Lauren Atlas, Eleni Linos, and Francisco Pereira. 2024. A comparison of ChatGPT and fine-tuned Open Pre-Trained Transformers (OPT) against widely used sentiment analysis tools: sentiment anal...

  22. [30]

    Carolyn M Mazure and Daniel P Jones. 2015. Twenty years and still counting: including women as participants and studying sex and gender in biomedical research. BMC Women’s Health 15 (2015), 1–16

  23. [31]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In ACL

  24. [32]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023). https://doi.org/10.48550/arXiv.2303.08774 arXiv:2303.08774

  25. [33]

    Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen. 2023. Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering. arXiv:2303.13534 (2023)

  26. [34]

    Sachit Menon and Carl Vondrick. 2023. Visual Classification via Description from Large Language Models. In ICLR

  27. [35]

    Max Pellert, Clemens M Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. 2023. AI Psychometrics: Using psychometric inventories to obtain psychological profiles of large language models. (2023)

  28. [36]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML

  29. [37]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sand- hini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Le...

  30. [38]

    Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2023. In-Context Impersonation Reveals Large Language Models’ Strengths and Biases. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Sy...

  31. [39]

    Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, D...

  32. [40]

    Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In CHI

  33. [41]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role-Play with Large Language Models. ArXiv:2305.16367 (2023)

  34. [42]

    Logan IV, Eric Wallace, and Sameer Singh

    Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In EMNLP

  35. [43]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexan- dra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît...

  36. [44]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume , Vol. 1, No. 1, Article . Publication date: January ...

  37. [45]

    Carl van Walraven and Ella Rokosh. 1999. What is necessary for high-quality discharge summaries? American Journal of Medical Quality 14, 4 (1999), 160–169

  38. [46]

    Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Q...

  39. [47]

    kelly is a warm person, joseph is a role model

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. " kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters. arXiv preprint arXiv:2310.09219 (2023)

  40. [48]

    Angelina Wang, Jamie Morgenstern, and John P Dickerson. 2024. Large language models cannot replace human participants because they cannot portray identity groups. arXiv preprint arXiv:2402.01908 (2024)

  41. [49]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned Language Models are Zero-Shot Learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April ...

  42. [50]

    Samer Abdulateef Waheeb, Naseer Ahmed Khan, Bolin Chen, and Xuequn Shang. 2020. Machine learning based sentiment text classification for evaluating treatment quality of discharge summary. Information 11, 5 (2020), 281

  43. [51]

    Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua. 2023. Fundamental Limitations of Alignment in Large Language Models. arXiv:2304.11082 (2023)

  44. [52]

    Discharge Me!

    Haotian Wu, Paul Boulenger, Antonin Faure, Berta Céspedes, Farouk Boukil, Nastasia Morel, Zeming Chen, and Antoine Bosselut. 2024. EPFL-MAKE at “Discharge Me!”: An LLM System for Automatically Generating Discharge Summaries of Clinical Electronic Health Record. In Proceedings ...

  45. [53]

    Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023. Large Language Models are Diverse Role-Players for Summarization Evaluation. In Natural Language Processing and Chinese Computing - 12th National CCF Conference, NLPCC 2023, Foshan, China, October 12-15, 20...

  46. [54]

    Chi, Quoc V Le, and Denny Zhou

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. In NeurIPS

  47. [55]

    Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar. 2022. Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification. arXiv:2211.11158 (2022)

  48. [56]

    Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023. How well do Large Language Models perform in Arithmetic tasks? arXiv:2304.02015 (2023)

  49. [57]

    Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer...

  50. [58]

    Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. 2018. Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly. TPAMI 41, 9 (2018)

  51. [62]

    Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large Language Models Are Human-Level Prompt Engineers. In NeurIPS Workshops. Appendix A Prompt Example with Lower Education Level Persona You are someone who has neve...

  52. [2022]

    CoRR abs/2201.08239 (2022)

    LaMDA: Language Models for Dialog Applications. CoRR abs/2201.08239 (2022). arXiv:2201.08239 https: //arxiv.org/abs/2201.08239

  53. [2023]

    Ask Me Anything: A simple strategy for prompting language models. In ICLR

  54. [2024]

    Better Zero-Shot Reasoning with Role-Play Prompting. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024 , Ke...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.