REVIEW 5 major objections 5 minor 62 references
Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Persona-prompted GPT-4 matches high-education patients on discharge-summary questions, but not female or low-visit patients.
desk verdict A useful pilot on LLM persona simulation for health communication, but the headline 88% claim is not supported by the undefined alignment metric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The intervention is in-context impersonation: the authors prepend "If you were a {persona}" to a fixed task instruction, and the model answers multiple-choice questions by letter only. Personas are built from four demographic attributes (education, gender, doctor-visit frequency, ER-visit frequency) split into high and low. The evaluation machinery is the reported alignment rate, which compares the persona-prompted LLM's letter choice with the dominant answer of the corresponding human group on the same questions, reported separately for eight information-based questions and two perception-based questions, with a random-guess baseline for comparison.
What would settle it
Re-score the survey using full answer distributions rather than modal agreement: for each of the ten questions, record the LLM's letter distribution under the "high education" prompt, the "low education" prompt, and a no-persona prompt, and compare them with the human group distributions. The claimed persona effect is falsified if the LLM's distribution barely moves across personas on questions where human groups differ, or if the no-persona prompt matches a group as well as that group's persona prompt does; the easiest check is Q2 and Q3, where most humans answer "A. Yes" and the model might do the same regardless of persona.
Extended reading notes
Core claim
The authors set out to establish that LLM-driven personas can reproduce how real patients answer questions about discharge summaries, and to map where that reproduction fails. Their central quantitative finding is that the model's alignment with human answers depends strongly on the persona and the task type: education-based personas average 77.5% alignment, high-education information questions hit 88%, and male information questions hit 97%, while female, low-doctor-visit, and low-ER-visit personas fall to 31–47%. The paper also finds that the model oversimplifies: it concentrates answers on one or two options, never chooses "I don't know," and overestimates comprehension of long summaries. The authors read these results as evidence that in-context impersonation is promising for self-similar groups but not ready for diverse patient populations, and that adding patient-specific context can reduce rather than improve accuracy.
Load-bearing premise
The whole evaluation assumes that agreeing with the most popular answer in a human group demonstrates that the model is simulating that group's perspective; if the agreement comes from the model's generic preference for common answers, the persona scores do not measure what they claim to.
Editorial extensions
If this is right
- Education-primed LLM personas could serve as a low-cost first-pass screening tool for how well-educated patients understand discharge instructions.
- Automated patient communication systems should not yet use LLM personas to represent female, low doctor-visit, or low ER-visit populations, because alignment there is near random.
- Including demographic details beyond education in a prompt can hurt performance, so a simpler query-response format may be safer for generating patient-facing health text.
- Because the model overestimates comprehension of long summaries and never expresses uncertainty, any LLM-generated discharge material needs human review before release.
Reading between the lines
- The alignment metric may be rewarding base-rate matching: on questions where almost all humans choose "A. Yes," a model that prefers "A" scores high for every persona; comparing full answer distributions would separate persona-specific simulation from default answer tendencies.
- A direct test of the mechanism would be to run the same prompts with no persona and with several personas; if the no-persona distribution matches high-education humans as well as the high-education persona does, the persona prefix is not doing the work.
- The failure to ever answer "I don't know" is a plausible cause of both the inflated male and high-education scores and the depressed female and low-visit scores; adding a calibration step that permits uncertainty may change which groups look well simulated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates whether GPT-4 can simulate the behavior of patient groups when answering questions about ICU discharge summaries. The authors recruited 96 Prolific participants, collected answers to eight information-based and two perception-based questions across four discharge summaries, and compared these with GPT-4 responses generated under personas defined by education, gender, doctor-visit frequency, and ER-visit frequency. They report alignment rates per persona (Table 2) and use them to conclude that education-primed personas achieve high alignment (88% for high-education information tasks), that performance drops for female and low-ER-visit personas, and that adding other information can reduce performance below random guessing. The paper frames this as a first step toward using LLMs as proxies for human subjects in personalized health communication.
Significance. If the quantitative claims were well supported, this would be a useful preliminary contribution to the growing literature on LLM role-playing and to debates about replacing human participants with LLM simulations. The external human-response data, IRB oversight, randomized discharge summaries, and transparent reporting of the per-persona percentages are genuine strengths, as is the honest qualitative discussion of overconfidence (Section 4.5) and of the risk of 'misportray[ing]' identity groups (Section 5). However, the paper's central quantitative claims currently rest on an undefined similarity metric, a weak random-guess baseline, and the absence of any no-persona control; until those are addressed, the alignment percentages cannot be interpreted as evidence for persona-simulation capability.
major comments (5)
- [§4.2, Table 2] The paper never defines the 'similarity'/'alignment rate' between the LLM response and human responses. It is not stated whether the comparison is to the modal human answer for each question, to the full human response distribution (e.g., via a distributional metric), or to a per-question pooled statistic. Without this definition, the headline numbers (0.88, 0.97, 0.31) and the aggregate 'average alignment rate of 54.97%' are uninterpretable, and the central claim that the LLM 'simulates' these personas is not falsifiable from the table.
- [Abstract and §4.2] The abstract claims that education-primed LLMs 'deliver accurate and actionable medical guidance 88% of the time,' but Table 2 shows 0.88 under 'High Edu.' for the information-based task, which is an alignment rate with human answers, not a measure of medical accuracy or actionability. Moreover, the abstract's claim that performance 'falls below random chance levels' is not generally true: in the information tasks every persona row is above the reported random-guess value (0.278), and in the perception tasks only the Male persona (0.25) falls below 0.267. The wording must be corrected to describe alignment with modal human responses and to give the precise exception.
- [§4.2] The random-guess baseline is not a meaningful control for skewed answer distributions. The reported baseline assumes uniform random choice among the answer options, but the human responses are highly skewed (e.g., Figure 4 shows a strong majority choosing 'A. Yes' for Q2-Q4). A majority-class baseline or a chance level conditional on the observed marginal answer distribution would be far higher than 0.278 for such questions. The paper also lacks a no-persona control condition: without asking the same questions without persona priming, or with a neutral persona, the high alignment of, say, the Male persona (0.97) may simply reflect the LLM's default tendency to choose common answers, a possibility the paper itself acknowledges in Section 5 when it notes that LLM responses are 'inherently identical or leptokurtic.'
- [§4.1 and §4.2] No confidence intervals, significance tests, or sampling details are provided for the LLM responses. Section 4.1 states only that GPT4_0613 was used via the Azure OpenAI API; the number of repeated samples per prompt, temperature, and number of questions per condition are not given. Section 4.2 uses the phrase 'significantly higher alignment rate' without any statistical test, and the differences between personas (e.g., High Edu. 0.88 vs. Low Edu. 0.72) are single point estimates from 96 human respondents split into multiple subgroups. These omissions make it impossible to judge whether any of the reported differences are robust.
- [Abstract and §7] The claim that 'a straightforward query-response model could outperform a more tailored approach' is not supported by any experiment reported in the paper. The evaluation only compares persona-primed prompts to random guessing; there is no non-persona or 'straightforward' baseline in Table 2 or elsewhere. If the authors have such data, it should be reported; otherwise this conclusion should be removed or reframed as a hypothesis.
minor comments (5)
- [Section 5] The terms 'LMM' and 'LMMs' appear several times (e.g., 'failure modes of LMM'); these should be 'LLM'/'LLMs'.
- [Appendix C.1] Q5 is labeled 'Do you know your diagnosis?' but the same label is used for Q3; the question text appears to be about other prescriptions, so the label is likely a copy-paste error.
- [Appendix C.1] Q8's option A reads 'A void fruit' and should read 'Avoid fruit'.
- [Table 2 and §4.2] The terms 'similarity' and 'alignment rate' are used interchangeably; they should be unified and formally defined.
- [References] The sentence 'Significant gender biases in LLM-generated content have been analyzed by Wang et al. in [47]' would benefit from a brief summary of the relevant finding, since the reference is to a general preprint.
Circularity Check
No significant circularity: LLM outputs are compared against independently collected human survey responses, and no fitted parameter, self-citation chain, or by-construction identity defines the reported alignment rates.
full rationale
The paper's central evaluation is an external comparison: human participants recruited through Prolific answered questions about discharge summaries, and GPT-4 was prompted with persona descriptions and asked the same questions. The reported 'alignment rate' is a similarity between LLM-generated answers and these pre-existing human answers. There is no fitted parameter derived from the human data that is then used to define the target, no training or fine-tuning on the human responses, and no equation in which the prediction reduces to its own input. The random-guess baseline is an independent reference point, and the paper's own discussion of misalignment (e.g., LLMs never choosing 'I don't know') is based on observed distribution differences rather than on a self-referential construction. The abstract's phrase 'accurate and actionable medical guidance 88% of the time' overstates what the 0.88 alignment figure means, and the alignment metric itself is not formally defined in the paper; these are reporting and validity concerns, not circularity. There is also no load-bearing self-citation: the cited related work on role-playing and persona simulation is background context rather than a premise that forces the paper's results. Therefore, the derivation chain is self-contained with respect to the external human benchmark, and no circular step meets the evidentiary standard required for a positive finding.
Assumptions & free parameters
free parameters (1)
- LLM samples per prompt =
not reported
assumptions (5)
- domain assumption Human self-reported demographic categories on Prolific are accurate labels for ground-truth groups.
- domain assumption The ten survey questions validly measure comprehension of discharge summaries.
- domain assumption Multiple LLM generations approximate a persona's response distribution.
- standard math The random guess baseline is computed correctly and serves as a meaningful chance level.
- domain assumption GPT-4's behavior at inference time is stable enough for comparison across personas.
Cite this review
Pith. "Pith review of Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives." pith.science (2026). https://pith.science/paper/AIFCKOIF
@misc{pith2026250106964,
author = {Pith},
title = {Pith review of: Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/AIFCKOIF}},
note = {Machine review of arXiv:2501.06964}
}
read the original abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in role-playing scenarios, particularly in simulating domain-specific experts using tailored prompts. This ability enables LLMs to adopt the persona of individuals with specific backgrounds, offering a cost-effective and efficient alternative to traditional, resource-intensive user studies. By mimicking human behavior, LLMs can anticipate responses based on concrete demographic or professional profiles. In this paper, we evaluate the effectiveness of LLMs in simulating individuals with diverse backgrounds and analyze the consistency of these simulated behaviors compared to real-world outcomes. In particular, we explore the potential of LLMs to interpret and respond to discharge summaries provided to patients leaving the Intensive Care Unit (ICU). We evaluate and compare with human responses the comprehensibility of discharge summaries among individuals with varying educational backgrounds, using this analysis to assess the strengths and limitations of LLM-driven simulations. Notably, when LLMs are primed with educational background information, they deliver accurate and actionable medical guidance 88% of the time. However, when other information is provided, performance significantly drops, falling below random chance levels. This preliminary study shows the potential benefits and pitfalls of automatically generating patient-specific health information from diverse populations. While LLMs show promise in simulating health personas, our results highlight critical gaps that must be addressed before they can be reliably used in clinical settings. Our findings suggest that a straightforward query-response model could outperform a more tailored approach in delivering health information. This is a crucial first step in understanding how LLMs can be optimized for personalized health communication while maintaining accuracy.
Figures
Reference graph
Works this paper leans on
-
[1]
Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society . 298–306
2021
-
[2]
Gati Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2022. Using Large Language Models to Simulate Multiple Humans. arXiv:2208.10264 (2022)
arXiv 2022
-
[3]
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...
-
[4]
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis (2023)
work page 2023
-
[5]
Simran Arora, Avanika Narayan, Mayee F Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, and Christopher Re
-
[6]
Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, ...
-
[7]
Francoise Baylis. 1996. Women and health research: working for change. The Journal of Clinical Ethics 7, 3 (1996), 229–242
work page 1996
-
[8]
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In ACM FAccT
work page 2021
Show all 62 references
-
[9]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mo- hammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A Suite for Analyzing Lar...
2023
-
[10]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901
2020
-
[11]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[12]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017)
2017
-
[13]
Akhil Chintalapati, Shaurya Agarwal, Aparajita Senapati, Mohammed Misbah, RS Nithyashree, and Sunil Kumar
- [14]
-
[15]
In 2023 International Conference on Next Generation Electronics (NEleX)
Textual Alchemy: Transmuting BERT Summaries with GAN Elegance. In 2023 International Conference on Next Generation Electronics (NEleX). IEEE, 1–7. , Vol. 1, No. 1, Article . Publication date: January 2025. 16 Xinyao Ma, Rui Zhu, Zihao Wang, Jingwei Xiong, Qingyu Chen, Haixu Ta...
2023
-
[16]
Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz. 2023. Inducing anxiety in large language models increases exploration and bias. arXiv:2304.11111 (2023)
2023 arXiv
-
[17]
Zhao, Yanping Huang, Andrew M
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa De- hghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vince...
-
[18]
Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M Johnson. 2024. EvaluLLM: LLM assisted evaluation of generative outputs. In Companion Proceedings of the 29th International Conference on Intelligent User Interfaces. 30–32
2024
-
[19]
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv:2304.05335 (2023)
2023 arXiv
-
[20]
Seungju Han, Beomsu Kim, Jin Yong Yoo, Seokjun Seo, Sangbum Kim, Enkhbayar Erdenee, and Buru Chang. 2022. Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a Few Utterances. In NAACL-HLT
2022
-
[21]
Stacie E Geller, Abby Koch, Beth Pellettieri, and Molly Carnes. 2011. Inclusion, analysis, and reporting of sex and race/ethnicity in clinical trials: have we made progress? Journal of women’s health 20, 3 (2011), 315–320
2011
-
[22]
Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan. 2023. Causal Reasoning and Large Language Models: Opening a New Frontier for Causality. arXiv:2305.00050 (2023)
2023 arXiv
-
[23]
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv:1909.05858 (2019)
2019 arXiv
-
[24]
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong
-
[25]
Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. 2024. Understanding users’ dissatisfaction with chatgpt responses: Types, resolving tactics, and the effect of knowledge level. In Proceedings of the 29th International Conference on Intelligent User Interfaces . 385–404
2024
-
[26]
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill. 2022. Can language models learn from explanations in context?. In EMNLP. ACL
2022
-
[27]
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing S...
2023
-
[28]
Gleb Kumichev, Pavel Blinov, Yulia Kuzkina, Vasily Goncharov, Galina Zubkova, Nikolai Zenovkin, Aleksei Goncharov, and Andrey Savchenko. 2024. MedSyn: LLM-Based Synthetic Medical Text Generation Framework. In Joint European Conference on Machine Learning and Knowledge Discover...
2024
-
[29]
Juan Antonio Lossio-Ventura, Rachel Weger, Angela Y Lee, Emily P Guinee, Joyce Chung, Lauren Atlas, Eleni Linos, and Francisco Pereira. 2024. A comparison of ChatGPT and fine-tuned Open Pre-Trained Transformers (OPT) against widely used sentiment analysis tools: sentiment anal...
2024
-
[30]
Carolyn M Mazure and Daniel P Jones. 2015. Twenty years and still counting: including women as participants and studying sex and gender in biomedical research. BMC Women’s Health 15 (2015), 1–16
2015
-
[31]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In ACL
2022
- [32]
-
[33]
Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen. 2023. Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering. arXiv:2303.13534 (2023)
2023 arXiv
-
[34]
Sachit Menon and Carl Vondrick. 2023. Visual Classification via Description from Large Language Models. In ICLR
2023
-
[35]
Max Pellert, Clemens M Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. 2023. AI Psychometrics: Using psychometric inventories to obtain psychological profiles of large language models. (2023)
2023
-
[36]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML
2021
-
[37]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sand- hini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Le...
2022
-
[38]
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2023. In-Context Impersonation Reveals Large Language Models’ Strengths and Biases. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Sy...
2023
-
[39]
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, D...
2022
-
[40]
Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In CHI
2021
-
[41]
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role-Play with Large Language Models. ArXiv:2305.16367 (2023)
2023 arXiv
-
[42]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In EMNLP
2020
- [43]
- [44]
-
[45]
Carl van Walraven and Ella Rokosh. 1999. What is necessary for high-quality discharge summaries? American Journal of Medical Quality 14, 4 (1999), 160–169
1999
-
[46]
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Q...
-
[47]
kelly is a warm person, joseph is a role model
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. " kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters. arXiv preprint arXiv:2310.09219 (2023)
2023 arXiv
-
[48]
Angelina Wang, Jamie Morgenstern, and John P Dickerson. 2024. Large language models cannot replace human participants because they cannot portray identity groups. arXiv preprint arXiv:2402.01908 (2024)
2024 arXiv
-
[49]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned Language Models are Zero-Shot Learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April ...
2022
-
[50]
Samer Abdulateef Waheeb, Naseer Ahmed Khan, Bolin Chen, and Xuequn Shang. 2020. Machine learning based sentiment text classification for evaluating treatment quality of discharge summary. Information 11, 5 (2020), 281
2020
-
[51]
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua. 2023. Fundamental Limitations of Alignment in Large Language Models. arXiv:2304.11082 (2023)
2023 arXiv
-
[52]
Discharge Me!
Haotian Wu, Paul Boulenger, Antonin Faure, Berta Céspedes, Farouk Boukil, Nastasia Morel, Zeming Chen, and Antoine Bosselut. 2024. EPFL-MAKE at “Discharge Me!”: An LLM System for Automatically Generating Discharge Summaries of Clinical Electronic Health Record. In Proceedings ...
2024
-
[53]
Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023. Large Language Models are Diverse Role-Players for Summarization Evaluation. In Natural Language Processing and Chinese Computing - 12th National CCF Conference, NLPCC 2023, Foshan, China, October 12-15, 20...
2023 doi
-
[54]
Chi, Quoc V Le, and Denny Zhou
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. In NeurIPS
2022
-
[55]
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar. 2022. Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification. arXiv:2211.11158 (2022)
2022 arXiv
-
[56]
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023. How well do Large Language Models perform in Arithmetic tasks? arXiv:2304.02015 (2023)
2023 arXiv
-
[57]
Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer...
-
[58]
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. 2018. Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly. TPAMI 41, 9 (2018)
2018
-
[62]
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large Language Models Are Human-Level Prompt Engineers. In NeurIPS Workshops. Appendix A Prompt Example with Lower Education Level Persona You are someone who has neve...
2022
-
[2022]
CoRR abs/2201.08239 (2022)
LaMDA: Language Models for Dialog Applications. CoRR abs/2201.08239 (2022). arXiv:2201.08239 https: //arxiv.org/abs/2201.08239
2022 arXiv
-
[2023]
Ask Me Anything: A simple strategy for prompting language models. In ICLR
-
[2024]
Better Zero-Shot Reasoning with Role-Play Prompting. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024 , Ke...
2024 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.