Pith. sign in

REVIEW 3 major objections 5 minor 38 references

Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Asking chatbots to argue emotionally makes their wording test as more rational than asking for rational arguments.

desk verdict Useful cross-model mapping of emotional vs rational prompting, but the headline LIWC claims need significance testing before they can be believed. read the letter →

arxiv 2502.09687 v1 pith:CKY252DA submitted 2025-02-13 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords largelanguagemodelspersuasionemotionalrationalLIWC-22socialinfluenceprinciplescognitivecomplexitymisinformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the wording of a persuasion prompt changes the psycholinguistic fingerprints of LLM responses in a counterintuitive way. Across twelve models from four families, responses prompted with "use emotional arguments" score higher than responses prompted with "use rational arguments" on nearly every LIWC-22 cognitive-processing indicator the authors selected, including insight, causation, discrepancy, tentativeness, certitude, differentiation, and memory, along with stronger affective indicators. The paper also tries to establish that the baseline prompt, which names no persuasion style, is the least cognitively complex and carries a subtle negative emotional slant, and that emotional versus rational prompts recruit different social-influence principles. If true, these findings show that a few words of instruction control which persuasive mechanisms a model deploys, giving both manipulators and defenders a concrete lever over AI-generated persuasion.

What carries the argument

The carrying mechanism is an experimental contrast between three prompt setups, baseline, "use emotional arguments," and "use rational arguments," applied 60 times to each of twelve models. Responses are scored with LIWC-22, a word-count tool that returns the percentage of words in psycholinguistic categories; the paper uses cognitive-process categories as indicators of rational persuasion and affect categories as indicators of emotional persuasion. A separate human-annotation step has four judges code each response for presence of the six principles of social influence, commitment and consistency, reciprocity, scarcity, liking and sympathy, authority, and social proof, with at least three agreeing required. The machinery converts a qualitative prompt instruction into quantitative linguistic indicators and principle frequencies, which is what allows the paper to compare setups and model families.

What would settle it

On the same 60-response-per-model data, compute bootstrap 95% confidence intervals for the difference between emotional and rational setups on the insight and affect categories; if the intervals include zero for most of the twelve models, the claim that emotional prompting consistently outperforms rational prompting on cognitive complexity would be falsified.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a paradox: emotional persuasion makes LLM language look more rational, not less. Relative to the rational-prompt condition, the emotional condition produced higher mean values in almost every cognitive-processing category in LIWC-22, which the authors interpret as more elaborate, confident, and insightful argumentation, alongside more polarized all-or-none thinking. The baseline condition, by contrast, had the lowest cognitive-complexity scores and a muted rational style with elevated anger and sadness in some models, suggesting that models default to restrained rational persuasion with subtle negative affect. A human annotation of six principles of social influence adds that emotional prompts predominantly trigger commitment and consistency, liking and sympathy, and social proof, while rational prompts predominantly trigger authority and social proof; the chi-square comparisons are significant except for reciprocity. The authors also report model-family differences, with some models adapting their social-influence strategy to the prompt and others not.

Load-bearing premise

The central comparison assumes that the reported differences in average LIWC-22 percentages between emotional, rational, and baseline setups are real signals rather than sampling noise, but the paper gives no confidence intervals or significance tests for those differences.

Editorial extensions

If this is right

  • Emotional prompting is a practical lever for more cognitively complex output: users or systems that want insight, causation, and certainty in an answer may get more of those markers by asking for emotional arguments than by asking for rational arguments.
  • Baseline behavior is not neutral: when no persuasion style is specified, models default to a cautious rational style with a subtle negative tint, so unsolicited persuasive text still carries emotional weight.
  • The social-influence tactics models use depend on the requested persuasion style: emotional prompts recruit commitment-and-consistency, liking, and social proof, while rational prompts recruit authority and social proof.
  • Model families differ in how flexibly they adapt, so a mitigation or amplification strategy that works for one model cannot be assumed to transfer to another.
  • Asking for rational arguments does not purge manipulative elements from the response, because rational-prompt outputs still lean on authority and social proof.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's paradox that emotional prompts increase measured rationality suggests that "rational" and "emotional" are not clean opposites in LLM text; a user study could test whether emotional-prompted messages are in fact more persuasive to human readers, which the linguistic analysis alone does not establish.
  • The central comparison could be tested directly by re-analyzing the same collected responses with bootstrap confidence intervals or mixed-effects models; if the emotional-versus-rational gaps on insight or affect fall within response-to-response noise for most models, the central conclusion would lose support.
  • The observation that one model produced no emotional responses invites a prompt-sensitivity metric that quantifies how much each model's linguistic fingerprint moves with the persuasion instruction, potentially linking that flexibility to model size, family, or alignment.
  • The six principles cover a narrow slice of influence tactics; extending the annotation to other documented strategies such as framing, foot-in-the-door, or personalization would show whether the emotional-versus-rational split generalizes beyond this taxonomy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies how twelve large language models from four families (OpenAI, Mixtral, Meta, Anthropic) respond to three prompting setups: a baseline prompt, an emotional-persuasion prompt, and a rational-persuasion prompt. Using LIWC-22, the authors compare linguistic indicators of rational and emotional persuasion across setups, and using human annotation, they examine which of Cialdini's six social-influence principles appear in the generated responses. The paper's central claims are that emotional prompting consistently increases LIWC-measured cognitive-complexity indicators (RQ1), that baseline prompting tends toward rational language with subtle negative affect (RQ2), and that emotional versus rational prompting evokes different social-influence principles (RQ3). The authors position the work within human-centered AI and discuss implications for mitigating risks of LLM-driven misinformation.

Significance. If the empirical claims were statistically well supported, the paper would make a useful contribution to the growing literature on persuasive LLMs by mapping linguistic differences across a broad set of current models and linking them to established psychological frameworks. Strengths include the wide model coverage, the standardized prompt template, the 60 repeated trials per condition, the explicit three-of-four-judge annotation criterion, and the use of chi-square tests for the social-influence analysis. However, the headline finding for RQ1 currently rests on raw means without any quantification of sampling variability, which is a substantial gap for a study whose conclusions are primarily about comparative differences.

major comments (3)
  1. [§4.1, Figures 2–3] The claim that "the emotional setup consistently outperforms both rational and baseline setups across nearly every linguistic indicator of rational persuasion" is supported only by raw LIWC-22 mean values; no confidence intervals, standard deviations, significance tests, or effect sizes are reported for any of these comparisons. Because the methods state that 60 repeated trials were run per model and setup, the reported mean differences are subject to sampling variability and cannot, as presented, be distinguished from noise; this is load-bearing since the same claim is restated in Section 5 as the answer to RQ1.
  2. [§4.2, §3.3] The annotation analysis reports chi-square tests for differences in use of Cialdini's principles between emotional and rational setups, but no inter-rater reliability statistic (e.g., Cohen's kappa or Krippendorff's alpha) is provided for the four judges, despite the explicit "at least three judges" decision rule; without evidence of agreement quality, the principle-frequency comparisons are difficult to interpret. In addition, the six chi-square tests are not corrected for multiple comparisons, so the claim that all principles except reciprocity differ significantly should be treated with caution.
  3. [§3.4, Table 2] The paper assumes that LIWC cognitive-process categories (insight, causation, certainty, differentiation, and so on) serve as linguistic indicators of rational persuasion, but this labeling decision is not validated; emotional language can also contain these words, and the 'paradox' reported in Section 5—that emotional prompts increase rational indicators—may partly reflect the selected indicator vocabulary rather than a genuine enhancement of rational argumentation. A validation discussion or at least a specific justification for mapping these LIWC categories to rational persuasion is needed before the conclusion can be accepted.
minor comments (5)
  1. [§3.3] The phrase "Claude Sonet" is a typo for Claude 3 Sonnet, and the sentence explaining why no emotional responses were available for annotation should specify whether this was a model refusal, an API error, or an annotation-filtering outcome.
  2. [§4.1] Several typographical and spacing issues appear in the results text, including "mostcertitude words," "discrepency," and "rised" in the introduction; these should be corrected before publication.
  3. [Figures 2–3] The line plots would be much more informative with error bars or shaded confidence bands; at minimum, the caption or text should state how many responses contribute to each plotted mean.
  4. [§4.2] The text mentions "OpenAI models (3.5 Turbo, 4, 4 Turbo, 4o, o1)" but o1 is not listed among the twelve models in §3.1; this discrepancy should be resolved.
  5. [Abstract and §5] The abstract and introduction frame the work as asking "whether and how we can mitigate the risks" of LLM-driven misinformation, but no mitigation intervention is tested; the conclusions only mention future safeguards, so the framing should be aligned with the actual scope.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the LIWC comparisons and chi-square tests are empirical observations, not reductions to the paper's own inputs.

full rationale

This paper is an empirical measurement study. The central claim in Section 4.1 that the emotional setup yields higher LIWC cognitive-process and affective scores than rational or baseline setups is a comparison of observed output statistics across fixed prompt templates. Nothing is fitted to the data and then renamed as a prediction, and no result is defined in terms of the quantity it is said to explain. The prompt definitions in Table 1, 'Use emotional arguments' versus 'Use rational arguments', do not analytically imply the LIWC category scores in Figures 2 and 3; the link is established only by running twelve models and measuring the generated text. The reuse in Section 3.2 of a prompt schema from the authors' prior work, Mieleszczenko-Kowszewicz et al. (2024), is a data and prompt provenance statement, not a self-citation that carries logical weight for the outcome. The labeling of LIWC categories as 'rational' or 'emotional' indicators in Table 2 is a coding decision based on the LIWC psycholinguistic dictionary, not a derivation that makes the empirical result true by definition. The chi-square test in Section 4.2 is an inferential check on the annotation data and is independent of the prompt wording. Therefore, no circular step reduces the paper's conclusions to its inputs. The absence of confidence intervals or significance tests for the LIWC differences is a statistical-support concern, not a circularity concern.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's comparisons rest on four unstated premises: the LIWC-22 categories selected in Section 3.4 validly operationalize rational and emotional persuasion; the majority-of-judges rule in Section 3.3 yields reliable annotations without a reported kappa; the prompt structure isolates the persuasion instruction (which is true for rational vs emotional but not for baseline comparisons, since baseline omits gender, trait, and belief variables); and the 10% annotated subset is representative. No free parameters are fitted to data and no new entities are introduced.

assumptions (4)
  • domain assumption LIWC-22 category percentages are valid indicators of rational persuasion (cognitive processes) and emotional persuasion (affect).
    Section 3.4 chooses two LIWC-22 categories as indicators of rational and emotional persuasion; the central comparisons in Section 4.1 depend entirely on this mapping.
  • domain assumption The 'independent judges rating method' with a majority of at least 3 of 4 judges reliably identifies presence of Cialdini principles.
    Section 3.3 defines presence by majority of judges but reports no inter-rater reliability (e.g., Cohen's kappa).
  • ad hoc to paper The prompt template variables (gender, trait level, belief) are the only systematic differences across setups besides the persuasion instruction.
    The emotional/rational prompts include {gender}, {level}, {trait}, and {belief} while the baseline prompt omits them, confounding baseline comparisons (Table 1).
  • domain assumption The 10% random subset of 517 responses is representative for annotation.
    Section 3.3 annotates only 10% of responses; no stratification or balance check is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models." pith.science (2026). https://pith.science/paper/CKY252DA

@misc{pith2026250209687,
  author       = {Pith},
  title        = {Pith review of: Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CKY252DA}},
  note         = {Machine review of arXiv:2502.09687}
}
read the original abstract

Be careful what you ask for, you just might get it. This saying fits with the way large language models (LLMs) are trained, which, instead of being rewarded for correctness, are increasingly rewarded for pleasing the recipient. So, they are increasingly effective at persuading us that their answers are valuable. But what tricks do they use in this persuasion? In this study, we examine what are the psycholinguistic features of the responses used by twelve different language models. By grouping response content according to rational or emotional prompts and exploring social influence principles employed by LLMs, we ask whether and how we can mitigate the risks of LLM-driven mass misinformation. We position this study within the broader discourse on human-centred AI, emphasizing the need for interdisciplinary approaches to mitigate cognitive and societal risks posed by persuasive AI responses.

Figures

Figures reproduced from arXiv: 2502.09687 by the authors.

Figure 1
Figure 1. Experimental setup for evaluating two types of persuasion in large language models (LLMs). The process consists of three stages: [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The graph compares emotional, rational, and baseline setups across various emotional linguistic indicators. The lines represent the [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The graph compares emotional, rational, and baseline setups across various rational linguistic indicators. The lines represent the [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Comparison of social influence principle frequencies in LLM responses across emotional and rational setups Note: * [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Frequencies of social influence principles across models in emotional and rational setups. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 28 canonical work pages

  1. [1]

    I ntroducing C laude 3.5 S onnet --- anthropic.com

    Anthropic (2024a). I ntroducing C laude 3.5 S onnet --- anthropic.com. https://www.anthropic.com/news/claude-3-5-sonnet. [Accessed 07-09-2024]

  2. [2]

    T he C laude 3 M odel F amily: O pus, S onnet, H aiku

    Anthropic (2024b). T he C laude 3 M odel F amily: O pus, S onnet, H aiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf. [Accessed 07-09-2024]

  3. [3]

    L., Rapp, A., Meyer, T., and Mullins, R

    Baker, T. L., Rapp, A., Meyer, T., and Mullins, R. (2014). The role of brand communications on front line service employee beliefs, behaviors, and performance. Journal of the academy of marketing science , 42:642--657

  4. [4]

    Binyamin, G. (2020). Do leader expectations shape employee service performance? enhancing self-expectations and internalization in employee role identity. Journal of Management & Organization , 26(4):536--554

  5. [5]

    L., Ashokkumar, A., Seraj, S., and Pennebaker, J

    Boyd, R. L., Ashokkumar, A., Seraj, S., and Pennebaker, J. W. (2022). The development and psychometric properties of liwc-22. Austin, TX: University of Texas at Austin , 10

  6. [6]

    Brader, T. (2005). Striking a responsive chord: How political ads motivate and persuade voters by appealing to emotions. American Journal of Political Science , 49(2):388--405

  7. [7]

    M., Egdal, D

    Breum, S. M., Egdal, D. V., Mortensen, V. G., M ller, A. G., and Aiello, L. M. (2024). The persuasive power of large language models. In Proceedings of the International AAAI Conference on Web and Social Media , volume 18, pages 152--163

  8. [8]

    D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

    Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems , 33:1877--1901

Show all 38 references
  1. [9]

    Carrasco-Farre, C. (2024). Large language models are as persuasive as humans, but how? about the cognitive effort and moral-emotional language of llm arguments. arXiv preprint arXiv:2404.09329

  2. [10]

    Cialdini, Robert, B. (2021). Influence, new and expanded: the psychology of persuasion. City/Country. New York

  3. [11]

    Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., ..., and Zhao, Z. (2024). The llama 3 herd of models

  4. [12]

    R., Hochwarter, W

    Ferris, G. R., Hochwarter, W. A., Douglas, C., Blass, F. R., Kolodinsky, R. W., and Treadway, D. C. (2002). Social influence processes in organizations and human resources systems. In Research in personnel and human resources management , pages 65--127. Emerald Group Publishin...

  5. [13]

    Fogg, B. J. (1998). Captology: the study of computers as persuasive technologies. In CHI 98 Conference Summary on Human Factors in Computing Systems , page 385

  6. [14]

    and Amichai-Hamburger, Y

    Fox, S. and Amichai-Hamburger, Y. (2001). The power of emotional appeals in promoting organizational change programs. Academy of Management Perspectives , 15(4):84--94

  7. [15]

    A., Chao, J., Grossman, S., Stamos, A., and Tomz, M

    Goldstein, J. A., Chao, J., Grossman, S., Stamos, A., and Tomz, M. (2024). How persuasive is ai-generated propaganda? PNAS nexus , 3(2):pgae034

  8. [16]

    Heath, R., Brandt, D., and Nairn, A. (2006). Brand relationships: Strengthened by emotion, weakened by attention. Journal of advertising research , 46(4):410--419

  9. [17]

    Jin, D., Mehri, S., Hazarika, D., Padmakumar, A., Lee, S., Liu, Y., and Namazifar, M. (2023). Data-efficient alignment of large language models with human feedback through natural language. arXiv preprint arXiv:2311.14543

  10. [18]

    Jones, C. R. and Bergen, B. K. (2024). Lies, damned lies, and distributional language statistics: Persuasion and deception with large language models. arXiv preprint arXiv:2412.17128

  11. [19]

    X., Park, J

    Karinshak, E., Liu, S. X., Park, J. S., and Hancock, J. T. (2023). Working with ai to persuade: Examining a large language model's ability to generate pro-vaccination messages. Proceedings of the ACM on Human-Computer Interaction , 7(CSCW1):1--29

  12. [20]

    Kelman, H. C. (2017). Further thoughts on the processes of compliance, identification, and internalization. In Social power and political influence , pages 125--171. Routledge

  13. [21]

    https://www.langchain.com

    Langchain (2025). https://www.langchain.com. [Accessed 02-02-2025]

  14. [22]

    Mazzotta, I., De Rosis, F., and Carofiglio, V. (2007). Portia: A user-adapted persuasion system in the healthy-eating domain. IEEE Intelligent systems , 22(6):42--51

  15. [23]

    McDonald, N., Schoenebeck, S., and Forte, A. (2019). Reliability and inter-rater reliability in qualitative research: Norms and guidelines for cscw and hci practice. Proceedings of the ACM on human-computer interaction , 3(CSCW):1--23

  16. [24]

    d., and Poggi, I

    Miceli, M., Rosis, F. d., and Poggi, I. (2006). Emotional and non-emotional persuasion. Applied Artificial Intelligence , 20(10):849--879

  17. [25]

    Mieleszczenko-Kowszewicz, W., P udowski, D., Ko odziejczyk, F., \'S wistak, J., Sienkiewicz, J., and Biecek, P. (2024). The dark patterns of personalized persuasion in large language models: Exposing persuasive linguistic features for big five personality traits in llms respon...

  18. [26]

    C heaper, B etter, F aster, S tronger

    MistralAI (2024). C heaper, B etter, F aster, S tronger. https://mistral.ai/news/mixtral-8x22b/. [Accessed 07-09-2024]

  19. [27]

    https://platform.openai.com/docs/models

    OpenAI (2024). https://platform.openai.com/docs/models. [Accessed 07-09-2024]

  20. [28]

    and Joffe, H

    O’Connor, C. and Joffe, H. (2020). Intercoder reliability in qualitative research: debates and practical guidelines. International journal of qualitative methods , 19:1609406919899220

  21. [29]

    B., Augenstein, I., and Assent, I

    Pauli, A. B., Augenstein, I., and Assent, I. (2024). Measuring and benchmarking large language models' capabilities to generate persuasive language. arXiv preprint arXiv:2406.17753

  22. [30]

    and Hogg, M

    Raftopoulou, E. and Hogg, M. K. (2010). The political role of government-sponsored social marketing campaigns. European Journal of Marketing , 44(7/8):1206--1227

  23. [31]

    J., and Mackie, D

    Rosselli, F., Skelly, J. J., and Mackie, D. M. (1995). Processing rational and emotional messages: The cognitive and affective mediation of persuasion. Journal of experimental social psychology , 31(2):163--190

  24. [32]

    Simons, H. W. (1976). Persuasion: Understanding, practice, and analysis. (No Title)

  25. [33]

    Tausczik, Y. R. and Pennebaker, J. W. (2010). The psychological meaning of words: Liwc and computerized text analysis methods. Journal of language and social psychology , 29(1):24--54

  26. [34]

    Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. (2022). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682

  27. [35]

    Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al. (2021). Ethical and social risks of harm from language models. arxiv. arXiv preprint arXiv:2112.04359 , 10

  28. [36]

    Wilczy \'n ski, P., Mieleszczenko-Kowszewicz, W., and Biecek, P. (2024). Resistance against manipulative ai: key factors and possible actions. arXiv preprint arXiv:2404.14230

  29. [37]

    Wilson, E. V. (2003). Perceived effectiveness of interpersonal persuasion strategies in computer-mediated communication. Computers in Human Behavior , 19(5):537--552

  30. [38]

    Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024). How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.