REVIEW 3 major objections 5 minor 38 references
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Asking chatbots to argue emotionally makes their wording test as more rational than asking for rational arguments.
desk verdict Useful cross-model mapping of emotional vs rational prompting, but the headline LIWC claims need significance testing before they can be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is an experimental contrast between three prompt setups, baseline, "use emotional arguments," and "use rational arguments," applied 60 times to each of twelve models. Responses are scored with LIWC-22, a word-count tool that returns the percentage of words in psycholinguistic categories; the paper uses cognitive-process categories as indicators of rational persuasion and affect categories as indicators of emotional persuasion. A separate human-annotation step has four judges code each response for presence of the six principles of social influence, commitment and consistency, reciprocity, scarcity, liking and sympathy, authority, and social proof, with at least three agreeing required. The machinery converts a qualitative prompt instruction into quantitative linguistic indicators and principle frequencies, which is what allows the paper to compare setups and model families.
What would settle it
On the same 60-response-per-model data, compute bootstrap 95% confidence intervals for the difference between emotional and rational setups on the insight and affect categories; if the intervals include zero for most of the twelve models, the claim that emotional prompting consistently outperforms rational prompting on cognitive complexity would be falsified.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a paradox: emotional persuasion makes LLM language look more rational, not less. Relative to the rational-prompt condition, the emotional condition produced higher mean values in almost every cognitive-processing category in LIWC-22, which the authors interpret as more elaborate, confident, and insightful argumentation, alongside more polarized all-or-none thinking. The baseline condition, by contrast, had the lowest cognitive-complexity scores and a muted rational style with elevated anger and sadness in some models, suggesting that models default to restrained rational persuasion with subtle negative affect. A human annotation of six principles of social influence adds that emotional prompts predominantly trigger commitment and consistency, liking and sympathy, and social proof, while rational prompts predominantly trigger authority and social proof; the chi-square comparisons are significant except for reciprocity. The authors also report model-family differences, with some models adapting their social-influence strategy to the prompt and others not.
Load-bearing premise
The central comparison assumes that the reported differences in average LIWC-22 percentages between emotional, rational, and baseline setups are real signals rather than sampling noise, but the paper gives no confidence intervals or significance tests for those differences.
Editorial extensions
If this is right
- Emotional prompting is a practical lever for more cognitively complex output: users or systems that want insight, causation, and certainty in an answer may get more of those markers by asking for emotional arguments than by asking for rational arguments.
- Baseline behavior is not neutral: when no persuasion style is specified, models default to a cautious rational style with a subtle negative tint, so unsolicited persuasive text still carries emotional weight.
- The social-influence tactics models use depend on the requested persuasion style: emotional prompts recruit commitment-and-consistency, liking, and social proof, while rational prompts recruit authority and social proof.
- Model families differ in how flexibly they adapt, so a mitigation or amplification strategy that works for one model cannot be assumed to transfer to another.
- Asking for rational arguments does not purge manipulative elements from the response, because rational-prompt outputs still lean on authority and social proof.
Reading between the lines
- The paper's paradox that emotional prompts increase measured rationality suggests that "rational" and "emotional" are not clean opposites in LLM text; a user study could test whether emotional-prompted messages are in fact more persuasive to human readers, which the linguistic analysis alone does not establish.
- The central comparison could be tested directly by re-analyzing the same collected responses with bootstrap confidence intervals or mixed-effects models; if the emotional-versus-rational gaps on insight or affect fall within response-to-response noise for most models, the central conclusion would lose support.
- The observation that one model produced no emotional responses invites a prompt-sensitivity metric that quantifies how much each model's linguistic fingerprint moves with the persuasion instruction, potentially linking that flexibility to model size, family, or alignment.
- The six principles cover a narrow slice of influence tactics; extending the annotation to other documented strategies such as framing, foot-in-the-door, or personalization would show whether the emotional-versus-rational split generalizes beyond this taxonomy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how twelve large language models from four families (OpenAI, Mixtral, Meta, Anthropic) respond to three prompting setups: a baseline prompt, an emotional-persuasion prompt, and a rational-persuasion prompt. Using LIWC-22, the authors compare linguistic indicators of rational and emotional persuasion across setups, and using human annotation, they examine which of Cialdini's six social-influence principles appear in the generated responses. The paper's central claims are that emotional prompting consistently increases LIWC-measured cognitive-complexity indicators (RQ1), that baseline prompting tends toward rational language with subtle negative affect (RQ2), and that emotional versus rational prompting evokes different social-influence principles (RQ3). The authors position the work within human-centered AI and discuss implications for mitigating risks of LLM-driven misinformation.
Significance. If the empirical claims were statistically well supported, the paper would make a useful contribution to the growing literature on persuasive LLMs by mapping linguistic differences across a broad set of current models and linking them to established psychological frameworks. Strengths include the wide model coverage, the standardized prompt template, the 60 repeated trials per condition, the explicit three-of-four-judge annotation criterion, and the use of chi-square tests for the social-influence analysis. However, the headline finding for RQ1 currently rests on raw means without any quantification of sampling variability, which is a substantial gap for a study whose conclusions are primarily about comparative differences.
major comments (3)
- [§4.1, Figures 2–3] The claim that "the emotional setup consistently outperforms both rational and baseline setups across nearly every linguistic indicator of rational persuasion" is supported only by raw LIWC-22 mean values; no confidence intervals, standard deviations, significance tests, or effect sizes are reported for any of these comparisons. Because the methods state that 60 repeated trials were run per model and setup, the reported mean differences are subject to sampling variability and cannot, as presented, be distinguished from noise; this is load-bearing since the same claim is restated in Section 5 as the answer to RQ1.
- [§4.2, §3.3] The annotation analysis reports chi-square tests for differences in use of Cialdini's principles between emotional and rational setups, but no inter-rater reliability statistic (e.g., Cohen's kappa or Krippendorff's alpha) is provided for the four judges, despite the explicit "at least three judges" decision rule; without evidence of agreement quality, the principle-frequency comparisons are difficult to interpret. In addition, the six chi-square tests are not corrected for multiple comparisons, so the claim that all principles except reciprocity differ significantly should be treated with caution.
- [§3.4, Table 2] The paper assumes that LIWC cognitive-process categories (insight, causation, certainty, differentiation, and so on) serve as linguistic indicators of rational persuasion, but this labeling decision is not validated; emotional language can also contain these words, and the 'paradox' reported in Section 5—that emotional prompts increase rational indicators—may partly reflect the selected indicator vocabulary rather than a genuine enhancement of rational argumentation. A validation discussion or at least a specific justification for mapping these LIWC categories to rational persuasion is needed before the conclusion can be accepted.
minor comments (5)
- [§3.3] The phrase "Claude Sonet" is a typo for Claude 3 Sonnet, and the sentence explaining why no emotional responses were available for annotation should specify whether this was a model refusal, an API error, or an annotation-filtering outcome.
- [§4.1] Several typographical and spacing issues appear in the results text, including "mostcertitude words," "discrepency," and "rised" in the introduction; these should be corrected before publication.
- [Figures 2–3] The line plots would be much more informative with error bars or shaded confidence bands; at minimum, the caption or text should state how many responses contribute to each plotted mean.
- [§4.2] The text mentions "OpenAI models (3.5 Turbo, 4, 4 Turbo, 4o, o1)" but o1 is not listed among the twelve models in §3.1; this discrepancy should be resolved.
- [Abstract and §5] The abstract and introduction frame the work as asking "whether and how we can mitigate the risks" of LLM-driven misinformation, but no mitigation intervention is tested; the conclusions only mention future safeguards, so the framing should be aligned with the actual scope.
Circularity Check
No significant circularity: the LIWC comparisons and chi-square tests are empirical observations, not reductions to the paper's own inputs.
full rationale
This paper is an empirical measurement study. The central claim in Section 4.1 that the emotional setup yields higher LIWC cognitive-process and affective scores than rational or baseline setups is a comparison of observed output statistics across fixed prompt templates. Nothing is fitted to the data and then renamed as a prediction, and no result is defined in terms of the quantity it is said to explain. The prompt definitions in Table 1, 'Use emotional arguments' versus 'Use rational arguments', do not analytically imply the LIWC category scores in Figures 2 and 3; the link is established only by running twelve models and measuring the generated text. The reuse in Section 3.2 of a prompt schema from the authors' prior work, Mieleszczenko-Kowszewicz et al. (2024), is a data and prompt provenance statement, not a self-citation that carries logical weight for the outcome. The labeling of LIWC categories as 'rational' or 'emotional' indicators in Table 2 is a coding decision based on the LIWC psycholinguistic dictionary, not a derivation that makes the empirical result true by definition. The chi-square test in Section 4.2 is an inferential check on the annotation data and is independent of the prompt wording. Therefore, no circular step reduces the paper's conclusions to its inputs. The absence of confidence intervals or significance tests for the LIWC differences is a statistical-support concern, not a circularity concern.
Assumptions & free parameters
assumptions (4)
- domain assumption LIWC-22 category percentages are valid indicators of rational persuasion (cognitive processes) and emotional persuasion (affect).
- domain assumption The 'independent judges rating method' with a majority of at least 3 of 4 judges reliably identifies presence of Cialdini principles.
- ad hoc to paper The prompt template variables (gender, trait level, belief) are the only systematic differences across setups besides the persuasion instruction.
- domain assumption The 10% random subset of 517 responses is representative for annotation.
Cite this review
Pith. "Pith review of Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models." pith.science (2026). https://pith.science/paper/CKY252DA
@misc{pith2026250209687,
author = {Pith},
title = {Pith review of: Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/CKY252DA}},
note = {Machine review of arXiv:2502.09687}
}
read the original abstract
Be careful what you ask for, you just might get it. This saying fits with the way large language models (LLMs) are trained, which, instead of being rewarded for correctness, are increasingly rewarded for pleasing the recipient. So, they are increasingly effective at persuading us that their answers are valuable. But what tricks do they use in this persuasion? In this study, we examine what are the psycholinguistic features of the responses used by twelve different language models. By grouping response content according to rational or emotional prompts and exploring social influence principles employed by LLMs, we ask whether and how we can mitigate the risks of LLM-driven mass misinformation. We position this study within the broader discourse on human-centred AI, emphasizing the need for interdisciplinary approaches to mitigate cognitive and societal risks posed by persuasive AI responses.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
I ntroducing C laude 3.5 S onnet --- anthropic.com
Anthropic (2024a). I ntroducing C laude 3.5 S onnet --- anthropic.com. https://www.anthropic.com/news/claude-3-5-sonnet. [Accessed 07-09-2024]
work page 2024
-
[2]
T he C laude 3 M odel F amily: O pus, S onnet, H aiku
Anthropic (2024b). T he C laude 3 M odel F amily: O pus, S onnet, H aiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf. [Accessed 07-09-2024]
work page 2024
-
[3]
L., Rapp, A., Meyer, T., and Mullins, R
Baker, T. L., Rapp, A., Meyer, T., and Mullins, R. (2014). The role of brand communications on front line service employee beliefs, behaviors, and performance. Journal of the academy of marketing science , 42:642--657
work page 2014
-
[4]
Binyamin, G. (2020). Do leader expectations shape employee service performance? enhancing self-expectations and internalization in employee role identity. Journal of Management & Organization , 26(4):536--554
work page 2020
-
[5]
L., Ashokkumar, A., Seraj, S., and Pennebaker, J
Boyd, R. L., Ashokkumar, A., Seraj, S., and Pennebaker, J. W. (2022). The development and psychometric properties of liwc-22. Austin, TX: University of Texas at Austin , 10
work page 2022
-
[6]
Brader, T. (2005). Striking a responsive chord: How political ads motivate and persuade voters by appealing to emotions. American Journal of Political Science , 49(2):388--405
work page 2005
-
[7]
Breum, S. M., Egdal, D. V., Mortensen, V. G., M ller, A. G., and Aiello, L. M. (2024). The persuasive power of large language models. In Proceedings of the International AAAI Conference on Web and Social Media , volume 18, pages 152--163
work page 2024
-
[8]
D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems , 33:1877--1901
2020
Show all 38 references
-
[9]
Carrasco-Farre, C. (2024). Large language models are as persuasive as humans, but how? about the cognitive effort and moral-emotional language of llm arguments. arXiv preprint arXiv:2404.09329
2024 arXiv
-
[10]
Cialdini, Robert, B. (2021). Influence, new and expanded: the psychology of persuasion. City/Country. New York
2021
-
[11]
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., ..., and Zhao, Z. (2024). The llama 3 herd of models
2024
-
[12]
R., Hochwarter, W
Ferris, G. R., Hochwarter, W. A., Douglas, C., Blass, F. R., Kolodinsky, R. W., and Treadway, D. C. (2002). Social influence processes in organizations and human resources systems. In Research in personnel and human resources management , pages 65--127. Emerald Group Publishin...
2002
-
[13]
Fogg, B. J. (1998). Captology: the study of computers as persuasive technologies. In CHI 98 Conference Summary on Human Factors in Computing Systems , page 385
1998
-
[14]
and Amichai-Hamburger, Y
Fox, S. and Amichai-Hamburger, Y. (2001). The power of emotional appeals in promoting organizational change programs. Academy of Management Perspectives , 15(4):84--94
2001
-
[15]
A., Chao, J., Grossman, S., Stamos, A., and Tomz, M
Goldstein, J. A., Chao, J., Grossman, S., Stamos, A., and Tomz, M. (2024). How persuasive is ai-generated propaganda? PNAS nexus , 3(2):pgae034
2024
-
[16]
Heath, R., Brandt, D., and Nairn, A. (2006). Brand relationships: Strengthened by emotion, weakened by attention. Journal of advertising research , 46(4):410--419
2006
-
[17]
Jin, D., Mehri, S., Hazarika, D., Padmakumar, A., Lee, S., Liu, Y., and Namazifar, M. (2023). Data-efficient alignment of large language models with human feedback through natural language. arXiv preprint arXiv:2311.14543
2023 arXiv
-
[18]
Jones, C. R. and Bergen, B. K. (2024). Lies, damned lies, and distributional language statistics: Persuasion and deception with large language models. arXiv preprint arXiv:2412.17128
2024 arXiv
-
[19]
X., Park, J
Karinshak, E., Liu, S. X., Park, J. S., and Hancock, J. T. (2023). Working with ai to persuade: Examining a large language model's ability to generate pro-vaccination messages. Proceedings of the ACM on Human-Computer Interaction , 7(CSCW1):1--29
2023
-
[20]
Kelman, H. C. (2017). Further thoughts on the processes of compliance, identification, and internalization. In Social power and political influence , pages 125--171. Routledge
2017
-
[21]
https://www.langchain.com
Langchain (2025). https://www.langchain.com. [Accessed 02-02-2025]
2025
-
[22]
Mazzotta, I., De Rosis, F., and Carofiglio, V. (2007). Portia: A user-adapted persuasion system in the healthy-eating domain. IEEE Intelligent systems , 22(6):42--51
2007
-
[23]
McDonald, N., Schoenebeck, S., and Forte, A. (2019). Reliability and inter-rater reliability in qualitative research: Norms and guidelines for cscw and hci practice. Proceedings of the ACM on human-computer interaction , 3(CSCW):1--23
2019
-
[24]
d., and Poggi, I
Miceli, M., Rosis, F. d., and Poggi, I. (2006). Emotional and non-emotional persuasion. Applied Artificial Intelligence , 20(10):849--879
2006
-
[25]
Mieleszczenko-Kowszewicz, W., P udowski, D., Ko odziejczyk, F., \'S wistak, J., Sienkiewicz, J., and Biecek, P. (2024). The dark patterns of personalized persuasion in large language models: Exposing persuasive linguistic features for big five personality traits in llms respon...
2024 arXiv
-
[26]
C heaper, B etter, F aster, S tronger
MistralAI (2024). C heaper, B etter, F aster, S tronger. https://mistral.ai/news/mixtral-8x22b/. [Accessed 07-09-2024]
2024
-
[27]
https://platform.openai.com/docs/models
OpenAI (2024). https://platform.openai.com/docs/models. [Accessed 07-09-2024]
2024
-
[28]
and Joffe, H
O’Connor, C. and Joffe, H. (2020). Intercoder reliability in qualitative research: debates and practical guidelines. International journal of qualitative methods , 19:1609406919899220
2020
-
[29]
B., Augenstein, I., and Assent, I
Pauli, A. B., Augenstein, I., and Assent, I. (2024). Measuring and benchmarking large language models' capabilities to generate persuasive language. arXiv preprint arXiv:2406.17753
2024 arXiv
-
[30]
and Hogg, M
Raftopoulou, E. and Hogg, M. K. (2010). The political role of government-sponsored social marketing campaigns. European Journal of Marketing , 44(7/8):1206--1227
2010
-
[31]
J., and Mackie, D
Rosselli, F., Skelly, J. J., and Mackie, D. M. (1995). Processing rational and emotional messages: The cognitive and affective mediation of persuasion. Journal of experimental social psychology , 31(2):163--190
1995
-
[32]
Simons, H. W. (1976). Persuasion: Understanding, practice, and analysis. (No Title)
1976
-
[33]
Tausczik, Y. R. and Pennebaker, J. W. (2010). The psychological meaning of words: Liwc and computerized text analysis methods. Journal of language and social psychology , 29(1):24--54
2010
-
[34]
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. (2022). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682
2022 arXiv
-
[35]
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al. (2021). Ethical and social risks of harm from language models. arxiv. arXiv preprint arXiv:2112.04359 , 10
2021 arXiv
-
[36]
Wilczy \'n ski, P., Mieleszczenko-Kowszewicz, W., and Biecek, P. (2024). Resistance against manipulative ai: key factors and possible actions. arXiv preprint arXiv:2404.14230
2024 arXiv
-
[37]
Wilson, E. V. (2003). Perceived effectiveness of interpersonal persuasion strategies in computer-mediated communication. Computers in Human Behavior , 19(5):537--552
2003
-
[38]
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024). How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.