Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

Large language models systematically vary persuasive style with recipient gender, producing communal language for female targets and agentic language for male targets across all 13 models and 16 languages.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 11:31 UTC pith:PO2AWWPJ

load-bearing objection Solid, well-verified audit of gender-stereotypical persuasive style in 13 LLMs; direction is credible, but per-category significance needs multiple-comparison correction and judge-noise propagation. the 3 major comments →

arxiv 2601.05751 v2 pith:PO2AWWPJ submitted 2026-01-09 cs.CL cs.AI

Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns

classification cs.CL cs.AI
keywords persuasive languagegender biasLLM-as-judgestereotypesagentic and communal traitsCialdini principlesrhetorical appealsmultilingual evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to establish whether large language models change the style of their persuasive language when the recipient's gender changes, and to measure how big that shift is. It builds a controlled evaluation framework in which each model writes to paired prompts that differ only in one attribute—recipient gender, sender intent, or output language—and an LLM judge then scores the pairs on 19 established categories of persuasion (rhetorical appeals, six classic persuasion principles, agentic/communal traits, interaction goals, and tone). Across all 13 models tested, responses addressed to female recipients were significantly more affectionate, polite, relational, communal, and emotional, while responses to male recipients were more direct, instrumental, agentic, and logical. The same gendered pattern appeared in a single-model study across 16 languages, and the framework also revealed shifts when intent was framed as noble versus ignoble. If the pattern is real, everyday uses of LLMs for drafting messages will systematically reproduce and potentially amplify gender-stereotypical styles at scale.

Core claim

The central claim is that every LLM examined—from small open-weight models to large safety-tuned systems—produces gender-stereotypical persuasive language: female-targeted messages and arguments are judged as warmer, more communal, more polite, and more pathos-laden, while male-targeted ones are judged as more direct, instrumental, agentic, and logos-laden. The paper reports statistically significant per-category differences (via a signed-rank test) across all models, with the same direction of effect on both interpersonal messages and political arguments, and in 15 additional languages when tested on a single multilingual model. The authors argue these patterns align with well-documented so

What carries the argument

The load-bearing mechanism is a pairwise-prompt evaluation framework: each test prompt is duplicated with a single swapped attribute (e.g., 'female coworker' vs 'male coworker'), the model generates both responses, and an LLM judge scores which text shows more of each of 19 categories on a symmetric -3 to +3 scale; positional bias is mitigated by scoring both orders and symmetrising. Aggregate per-category scores are tested with a signed-rank test, and a Treatment Gap (the sum of absolute mean differences) quantifies how strongly a model differentiates between treatments. The framework's credibility rests on verification steps: a second judge model correlates strongly with the first, replaci

Load-bearing premise

The paper's claim that the differences are 'significant across all models' depends on treating the LLM judge's scores as reliable measurements, but judge consistency is only about 75–84% across swapped orders, human annotators disagree substantially (agreement coefficient 0.09–0.2), and no correction is applied for testing 19 categories at once.

What would settle it

Re-run the core gender experiment on a random subset of the 13 models with multiple-comparison correction and a judge whose outputs are verified against a larger, carefully recruited human annotation set on a per-category basis; if fewer than half of the 19 categories remain significant in the stereotypical direction for a majority of models, the 'across all models' claim would fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • LLMs used to draft emails, fundraising appeals, or campaign arguments will systematically code persuasive style by recipient gender, potentially reinforcing traditional gender roles at scale.
  • Because safety-aligned and open models all show the effect, bias mitigation at the alignment layer alone is unlikely to remove gendered persuasive styles.
  • The pairwise framework transfers to other treatment attributes (e.g., noble vs ignoble intent, response language), giving researchers a reusable way to audit which prompt attributes shift generation style.
  • The persistence of the gendered pattern across 16 languages indicates the bias is not an English-specific quirk, likely arising from shared training-data patterns.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the judge reflects what humans perceive, then LLM-based personalisation tools—which already tailor messages to recipients—will amplify stereotypical style differences even when the sender never asked for them; a direct behavioural test would measure whether recipients actually respond differently to the two styles.
  • The larger gender gap in Chinese (versus other languages) despite no correlation with country-level gender inequality suggests training-data composition, rather than societal indices, modulates the bias; auditing a model's pre-training data for gendered language patterns could explain this variation.
  • A testable extension is to apply the same pairwise framework to other binary attributes (e.g., recipient age, socio-economic status, or race) to see whether the communal/agentic split generalises or is specific to gender.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a framework for measuring differences in persuasive language generated by LLMs under pairwise prompt treatments, and applies it to recipient gender (13 LLMs), sender intent framing, and output language (16 languages). Persuasive language is operationalized as 19 categories across five theory-derived dimensions and scored by GPT-4o as an LLM judge, with a symmetric positional swap. The central claim is that all tested LLMs show significant gender-stereotypical differences: female-targeted responses are more affectionate, polite, communal, relational, and pathos-laden, while male-targeted responses are more direct, instrumental, agentic, and logos-laden. The authors include verification measures for one anchor model (LLAMA3.3-70B): human annotation of collapsed Communal+/Agentic+ dimensions, an alternative judge model, gender-term neutralization, and cross-lingual translation checks.

Significance. If the central claim survives re-analysis, this is a practically important result: widely used LLMs automatically produce communal/agentic stylistic splits based solely on the recipient's gender, with potential to reinforce societal stereotypes in everyday persuasive writing. The paper's strengths are its controlled pairwise prompt design, the breadth of the evaluation (13 models, 16 languages, 19 categories), and its explicit verification protocol—especially the alternative-judge, neutralization, and human-annotation checks, which go beyond most LLM-as-judge studies. The paper is also transparent in reporting low human inter-annotator agreement and imperfect positional consistency, which makes the failure to propagate those uncertainties into the headline significance tests the more conspicuous.

major comments (3)
  1. [Sect. 3.2, Eq. (2); App. B; App. D; Figs. 3 and 5] The Wilcoxon signed-rank tests treat the averaged judge scores as exact measurements, but App. B (Table 4) reports positional consistency of only 75–84% for LLAMA3.3, and App. D reports human inter-annotator Krippendorff alpha of 0.09–0.2. Neither source of noise is propagated into the significance tests. With 19 categories × 2 test sets × 13 models (≈494 tests), uncorrected α=0.05 implies a substantial expected number of false positives. Please report FDR-corrected p-values and/or bootstrap intervals that resample the two positional judgments (and, ideally, multiple judge runs). Without this, the Abstract's 'significant ... across all models' claim is not quantitatively supported.
  2. [Sect. 7; App. D; Sect. 4.2] The verification experiments—human annotations, alternative judge, gender-term neutralization, and cross-lingual translation checks—are all performed only on LLAMA3.3-70B responses, and the human annotation collapses the 19 categories into two aggregate questions (Communal+ and Agentic+). The paper's headline conclusion, however, covers 13 models and asserts fine-grained per-category differences, including Cialdini principles and interaction goals. Either extend validation to at least a few additional models and to the full 19-category vector, or restate the conclusion as a narrower claim about the communal/agentic direction, presenting the per-model fine-grained results as exploratory rather than as verified findings.
  3. [App. E, Table 8; Sect. 4.2] Refusal rates are treatment-correlated and high for several models (e.g., GPT-5 female 84.7% vs male 69.3%; CLAUDE-OPUS female 46% vs male 40%). The analysis omits refused prompts and retains only 10 models for the argument condition. Because refusals are not gender-neutral, the subset of prompts could introduce selection bias into the aggregate gender-gap comparisons. Please report whether the pattern in Figs. 4–5 is robust to the omitted prompts (e.g., via sensitivity analysis on the 10-prompt omission list) or explicitly discuss this as a limitation for the argument condition.
minor comments (4)
  1. [Throughout] Typos and grammatical slips: 'catagory' (Eq. 2 context), 'vice-vesa' (Sect. 4.2), 'Simultanously' (Sect. 2), 'wiht key' (App. A.2), 'Krippendorf Alpha ranging from 0.09 to 20.2' (App. D; should be 0.09 to 0.2).
  2. [Fig. 5] In both panels, several model labels are duplicated (Qwen3-235B and Qwen3-30B appear twice each), making it difficult to map rows to models. Please use unique labels.
  3. [Limitations] The sentence 'the strong agreement exhibited these independent communication theoretic frameworks' is ungrammatical and overstates the independence of the categories; the frameworks overlap conceptually (e.g., politeness and communal orientation). A minor rewrite would clarify.
  4. [Sect. 7.1] The phrase 'Sect. 7' appears as a self-reference in the first paragraph; consider replacing with 'this section' for clarity.

Circularity Check

0 steps flagged

No significant circularity; minor non-load-bearing self-citations only

full rationale

The paper's derivation chain is empirical rather than definitional: paired prompts differing only in the treatment attribute are fed to 13 different generation models; the generated responses are scored by GPT-4O (a separate model) across 19 categories; the mean directional differences Dj (Eq. 2) are tested with a Wilcoxon signed-rank test. The central claim (Sect. 4.2) is an aggregate of these measured scores and could in principle have been null or mixed-direction. The judge is not one of the tested generators, so the inputs and outputs are not the same object. The framework is verified with an alternative judge (Spearman rho=0.852/0.814), gender-term neutralization (rho=0.991/0.987), and human annotations on collapsed Communal+/Agentic+ dimensions (Sect. 7). Self-citations (Pauli et al. 2022, 2025) are used for a definition of persuasive language and related work, but the operationalization is grounded in external sources (Cialdini 2007; Bakan 1966; Wilson & Putnam 2012; Gass & Seiter 2010), so the self-citation is not load-bearing. The skeptical concerns about judge positional consistency, low inter-annotator agreement, and lack of multiple-comparison correction bear on reliability/validity of the measurements, not on whether the derivation reduces to its inputs; they are therefore outside the circularity verdict. No step exhibits Eq.-equals-Eq. reduction, fitted-parameter-renamed-as-prediction, or a uniqueness theorem imported from the authors.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The paper introduces no new theoretical entities; its contribution is a measurement framework and empirical findings. The main unverified premise is that an LLM judge provides reliable ground-truth measurements for all models, not just the one anchor model.

axioms (4)
  • domain assumption The 19 categories (rhetorical appeals, Cialdini principles, agency/communion, interaction goals, tones) coherently operationalize persuasive language.
    Section 3.1 grounds categories in Western theories (Aristotle, Cialdini, Bakan/Abele, Wilson and Putnam); the authors themselves note in Limitations that these may not transfer to non-Western contexts.
  • domain assumption GPT-4O's judgment is a valid and approximately unbiased measurement of these categories in paired texts.
    The entire pipeline (Eq. 1) relies on an LLM judge; human validation is only run for one model (LLAMA3.3, Sect. 7.1), so for the other 12 models this assumption is unverified.
  • domain assumption Paired responses differ only by the gender attribute in the prompt, so observed D_j is attributable to the treatment.
    Prompts are constructed pairwise (Sect. 3.3), but responses are generated independently, so stochastic variation is not controlled; the design assumes no other systematic difference.
  • domain assumption The Wilcoxon signed-rank test is valid on averaged judge scores.
    Scores are averaged over two order-swapped evaluations (Eq. 2), but judge positional inconsistency (App. B, Table 4: 75–84% consistency) is not propagated into the test.

pith-pipeline@v1.3.0-alltime-deepseek · 31906 in / 11473 out tokens · 120365 ms · 2026-08-03T11:31:38.069491+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns." pith.science (2026). https://pith.science/paper/PO2AWWPJ

@misc{pith2026260105751,
  author       = {Pith},
  title        = {Pith review of: Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PO2AWWPJ}},
  note         = {Machine review of arXiv:2601.05751}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade. Prior work has shown that LLMs can successfully persuade humans and amplify persuasive language. It is therefore essential to understand how user instructions affect the generation of persuasive language, and to understand whether the generated persuasive language differs, for example, when targeting different groups. In this work, we propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language. We evaluate 13 LLMs and 16 languages using pairwise prompt instructions. We evaluate model responses on 19 categories of persuasive language using an LLM-as-judge setup grounded in social psychology and communication science. Our results reveal significant gender differences in the persuasive language generated across all models. These patterns reflect biases consistent with gender-stereotypical linguistic tendencies documented in social psychology and sociolinguistics.

Figures

Figures reproduced from arXiv: 2601.05751 by Amalie Brogaard Pauli, Ira Assent, Isabelle Augenstein, Maria Barrett, Max M\"uller-Eberstein.

Figure 1
Figure 1. Figure 1: Example of how LLAMA 3.3 varies persuasive language when the prompt specifies recipient gender. 2023; Salvi et al., 2024). As such, understanding and safeguarding against AI persuasion have be￾come critical cross-disciplinary topics (Burtell and Woodside, 2023; El-Sayed et al., 2024). In this work, we investigate how attributes in the instruction affect the style of persuasive language generated by LLMs. S… view at source ↗
Figure 2
Figure 2. Figure 2: Framework for evaluating differences in LLM-generated persuasive language under pairwise prompt [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Mean differences in persuasive language catagories [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Persuasive language differences per model, [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 8
Figure 8. Figure 8: Gender Gap on languages using GPT5-MINI. pairwise bootstrap tests, we assess whether gender gaps differ significantly between two languages; many do not differ, but for example, the gender gap in Chinese is significantly larger than in all other languages except English. We test whether the gender gaps correlate with an index of gen￾der inequality across countries, but find no such correlation (using an ap… view at source ↗
Figure 7
Figure 7. Figure 7: Persuasive language differences per language, [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: Aggregated human annotations: Grey: no significant difference, Blue: significantly more male, Red: significantly more female. the input, we manually replace gendered terms in the generated messages and arguments (e.g., man/- woman) with neutral alternatives (e.g., human) and rerun the evaluation. We observe a strong corre￾lation between original and neutralised findings (messages: ρ = 0.991, arguments: ρ =… view at source ↗
Figure 10
Figure 10. Figure 10: Distribution over text length (count of char [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Widget showing a sample of the manual review to replace gender identifier terms with gender-neutral [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Judge QWEN2.5-72B: Average over the rated categories Dj on pairwise difference between gender￾treatment responses generated by LLAMA 3.3. The Wilcoxon test is applied to test the significance of the differences. Grey: not significant, Blue: significant in male direction, red: significant in female direction. could be on small details like wording, or what is emphasized the most. For each question, select … view at source ↗
Figure 13
Figure 13. Figure 13: Screenshot of the annotaion tool. E Gender treatment experiments across models We check whether the models respond to the re￾quest in the expected format, or whether they refuse to provide the argument or messages. We use a regex expression to find the refusal, and manu￾ally look through the matches to check for false positives. To examine for false negatives – re￾fusals responses as our regex missed – we… view at source ↗
Figure 15
Figure 15. Figure 15: Arguments test set: scatterplot showing average text length against gender gap across the tested models. To this aim, we select languages from diverse language families for which a language-to-country mapping is, to some extent, reasonably well approximated. To construct this mapping, we use native-speaker distributions per country1 and calculate the following: Country → Language: In a country, what perce… view at source ↗
Figure 16
Figure 16. Figure 16: P-values from bootstrapping analysis over [PITH_FULL_IMAGE:figures/full_fig_p022_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Scatterplot over Gender Inequality Index and [PITH_FULL_IMAGE:figures/full_fig_p022_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Average over the rated categories Dj on pairwise difference between the language treatment pair. The Wilcoxon test is applied to test the significance of the differences. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Boostrapping: testing the difference in gender gap in models pairwise (messages). [PITH_FULL_IMAGE:figures/full_fig_p024_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Boostrapping: testing the difference in gender gap in models pairwise (arguments). [PITH_FULL_IMAGE:figures/full_fig_p025_20.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthromorphism, and Maxims

    cs.CL 2026-06 unverdicted novelty 7.0

    Presents FRANZ framework and SQUARE corpus for multi-dimensional audit of LLM response framing on subjective cultural queries, applied to three models to reveal differences and couplings.

  2. Pareto-Guided Teacher Alignment for Fair Personalized Text Generation

    cs.CL 2026-06 unverdicted novelty 5.0

    Fairness mitigation in personalized text generation is objective-dependent with methods occupying different regions of the fairness-personalization Pareto frontier rather than any single strategy dominating all objectives.

Reference graph

Works this paper leans on

71 extracted references · 6 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Andrea E Abele and Bogdan Wojciszke. 2014. Communal and agentic content in social cognition: A dual perspective model. In Advances in experimental social psychology, volume 50, pages 195--255. Elsevier

  4. [4]

    Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.258 MEGA : Multilingual evaluation of generative AI . In Proceedings of the 2023 Conference on Empirical Methods in Natural ...

  5. [5]

    Anthropic. 2025. Claude opus 4.1. https://www.anthropic.com/news/claude-opus-4-1. Large‑language model release announcement

  6. [6]

    Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and 1 others. 2024. Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page...

  7. [7]

    David Bakan. 1966. The duality of human existence: An essay on psychology and religion

  8. [8]

    Cassondra Batz-Barbarich, Nicole Strah, and Farhan Masud Ahmed. 2025. Do words matter? the impact of communal and agentic language on women’s application to job opportunities. Journal of Personnel Psychology

  9. [9]

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fern \'a ndez, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, and 1 others. 2025. LLM s instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. In Proceedings of the 63rd Annual Meeting of the Association for Compu...

  10. [10]

    Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur, and Dilek Hakkani-T \"u r. 2025 a . Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. arXiv preprint arXiv:2503.01829

  11. [11]

    Nimet Beyza Bozdag, Shuhaib Mehri, Xiaocheng Yang, Hyeonjeong Ha, Zirui Cheng, Esin Durmus, Jiaxuan You, Heng Ji, Gokhan Tur, and Dilek Hakkani-T \"u r. 2025 b . Must read: A systematic survey of computational persuasion. arXiv preprint arXiv:2505.07775

  12. [12]

    Simon Martin Breum, Daniel V dele Egdal, Victor Gram Mortensen, Anders Giovanni M ller, and Luca Maria Aiello. 2024. The persuasive power of large language models. In Proceedings of the International AAAI Conference on Web and Social Media, volume 18, pages 152--163

  13. [13]

    S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, and 1 others. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712

  14. [14]

    Matthew Burtell and Thomas Woodside. 2023. Artificial Influence: An Analysis Of AI-Driven Persuasion . arXiv e-prints, pages arXiv--2303

  15. [15]

    Evan Chen, Run-Jun Zhan, Yan-Bai Lin, and Hung-Hsuan Chen. 2025. From structured prompts to open narratives: Measuring gender bias in llms through open-ended storytelling. arXiv preprint arXiv:2503.15904

  16. [16]

    Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8301--8327

  17. [17]

    Cheng-Han Chiang and Hung-yi Lee. 2023. https://doi.org/10.18653/v1/2023.acl-long.870 Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15607--15631, Toronto, Canada. Association for Computational Linguistics

  18. [18]

    Jaeyoon Choi and Nia Nixon. 2025. Agentic men, communal women?: Exploring gender bias in llm-based leadership identification for collaboration analytics. In Artificial Intelligence in Education, pages 11--18, Cham. Springer Nature Switzerland

  19. [19]

    Cialdini

    Robert B. Cialdini. 2007. Influence: The Psychology of Persuasion. HarperCollins e-books

  20. [20]

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261

  21. [21]

    Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. 2024. Or-bench: An over-refusal benchmark for large language models. arXiv preprint arXiv:2405.20947

  22. [22]

    DeepSeek-AI. 2024. https://arxiv.org/abs/2412.19437 Deepseek-v3 technical report . Preprint, arXiv:2412.19437

  23. [23]

    DeepSeek-AI. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948

  24. [24]

    Malika Dikshit, Houda Bouamor, and Nizar Habash. 2024. https://doi.org/10.18653/v1/2024.gebnlp-1.11 Investigating gender bias in STEM job advertisements . In Proceedings of the 5th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 179--189, Bangkok, Thailand. Association for Computational Linguistics

  25. [25]

    Xiangjue Dong, Yibo Wang, Philip Yu, and James Caverlee. 2023. https://openreview.net/forum?id=ZDeEYmKYrR Probing explicit and implicit gender bias through LLM conditional text generation . In Socially Responsible Language Modelling Research

  26. [26]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv e-prints, pages arXiv--2407

  27. [27]

    Eagly and Steven J

    Alice H. Eagly and Steven J. Karau. 2002. https://doi.org/10.1037/0033-295X.109.3.573 Role congruity theory of prejudice toward female leaders . Psychological Review, 109(3):573--598

  28. [28]

    Eagly and Wendy Wood

    Alice H. Eagly and Wendy Wood. 1991. https://doi.org/10.1177/0146167291173006 Explaining sex differences in social behavior: A meta-analytic perspective . Personality and Social Psychology Bulletin, 17(3):306--315

  29. [29]

    Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery, Geoff Keeling, Zachary Kenton, Zaria Jalan, Nahema Marchal, Arianna Manzini, Toby Shevlane, Shannon Vallor, and 1 others. 2024. A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI . arXiv preprint arXiv:2404.15058

  30. [30]

    Miller, Sasha Mitts, Adithya Renduchintala, and 8 others

    Meta Fundamental AI Research Diplomacy Team (FAIR)† FAIR, Meta, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, and 8 others. 2022. h...

  31. [31]

    Xiyan Fu and Wei Liu. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.587 How reliable is multilingual LLM -as-a-judge? In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 11040--11053, Suzhou, China. Association for Computational Linguistics

  32. [32]

    Gass and John S

    Robert H. Gass and John S. Seiter. 2010. Persuasion, Social Influence, and Compliance Gaining (4th ed.) . Boston: Allyn & Bacon

  33. [33]

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Yuanzhuo Wang, and Jian Guo. 2024. https://api.semanticscholar.org/CorpusID:274234014 A survey on llm-as-a-judge . ArXiv, abs/2411.15594

  34. [34]

    Elizabeth L Haines and Steven J Stroessner. 2019. The role prioritization model: How communal men and agentic women can (sometimes) have it all. Social and Personality Psychology Compass, 13(12):e12504

  35. [35]

    Matthew Honnibal. 2017. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. (No Title)

  36. [36]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  37. [37]

    Chuhao Jin, Kening Ren, Lingzhen Kong, Xiting Wang, Ruihua Song, and Huan Chen. 2024. https://doi.org/10.18653/v1/2024.acl-long.92 Persuading across diverse domains: a dataset and persuasion large language model . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1678--1706, Bangkok, ...

  38. [38]

    Jones, Frank R

    Bryan D. Jones, Frank R. Baumgartner, Sean M. Theriault, Derek A. Epp, Cheyenne Lee, and Miranda E. Sullivan. 2023. Policy agendas project: Codebook. https://www.comparativeagendas.net/pages/master-codebook

  39. [39]

    Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T. Hancock. 2023. https://doi.org/10.1145/3579592 Working With AI to Persuade: Examining a Large Language Model's Ability to Generate Pro-Vaccination Messages . Proc. ACM Hum.-Comput. Interact., 7(CSCW1)

  40. [40]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in large language models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery

  41. [41]

    Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Marie Beckage, Hsuan Su, Hung yi Lee, and Lama Nachman

    Shachi H. Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Marie Beckage, Hsuan Su, Hung yi Lee, and Lama Nachman. 2025. https://openreview.net/forum?id=tIYMiYz6Bf Decoding biases: An analysis of automated methods and metrics for gender bias detection in language models . In Red Teaming GenAI: What Can We Learn from Adversaries?

  42. [42]

    Robin Lakoff. 1973. Language and woman's place. Language in society, 2(1):45--79

  43. [43]

    Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, and 4 others. 2024. Tülu 3: Pushing frontiers in open language model...

  44. [44]

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. 2024. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning

  45. [45]

    Minqian Liu, Zhiyang Xu, Xinyi Zhang, Heajun An, Sarvech Qadir, Qi Zhang, Pamela J Wisniewski, Jin-Hee Cho, Sang Won Lee, Ruoxi Jia, and 1 others. 2025. Llm can be a dangerous persuader: Empirical study of persuasion safety in large language models. arXiv preprint arXiv:2504.10430

  46. [46]

    Yang Liu. 2024. Quantifying stereotypes in language. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1223--1240

  47. [47]

    Weicheng Ma, Hefan Zhang, Ivory Yang, Shiyu Ji, Joice Chen, Farnoosh Hashemi, Shubham Mohole, Ethan Gearey, Michael Macy, Saeed Hassanpour, and Soroush Vosoughi. 2025. https://doi.org/10.18653/v1/2025.naacl-long.203 Communication makes perfect: Persuasion dataset construction via multi- LLM communication . In Proceedings of the 2025 Conference of the Nati...

  48. [48]

    Sandra C Matz, Jacob D Teeny, Sumer S Vaid, Heinrich Peters, Gabriella M Harari, and Moran Cerf. 2024. The potential of generative ai for personalized persuasion at scale. Scientific Reports, 14(1):4692

  49. [49]

    Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aim \'e e Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, and 35 others. 2025. https://doi.org/10.18653/v1...

  50. [50]

    Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers...

  51. [51]

    OpenAI. 2025 a . https://cdn.openai.com/gpt-5-system-card.pdf GPT-5 System Card . Technical report, OpenAI. Technical report; model documentation and safety-evaluation details

  52. [52]

    OpenAI. 2025 b . Introducing gpt-4.1 in the api. https://openai.com/index/gpt-4-1/. [Large language model release announcement]

  53. [53]

    Ruby Ostrow and Adam Lopez. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.946 LLM s reproduce stereotypes of sexual and gender minorities . In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 17465--17477, Suzhou, China. Association for Computational Linguistics

  54. [54]

    Amalie Pauli, Leon Derczynski, and Ira Assent. 2022. https://doi.org/10.18653/v1/2022.nlp4pi-1.11 Modelling persuasion through misuse of rhetorical appeals . In Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI), pages 89--100, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics

  55. [55]

    Amalie Brogaard Pauli, Isabelle Augenstein, and Ira Assent. 2025. https://doi.org/10.18653/v1/2025.naacl-long.506 Measuring and benchmarking large language models' capabilities to generate persuasive language . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Tech...

  56. [56]

    Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden persuaders: Llms' political leaning and their influence on voters. arXiv preprint arXiv:2410.24190

  57. [57]

    Till Raphael Saenger, Musashi Hinck, Justin Grimmer, and Brandon M. Stewart. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.913 A uto P ersuade: A framework for evaluating and explaining persuasive arguments . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 16325--16342, Miami, Florida, USA. Association ...

  58. [58]

    Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2024. On the conversational persuasiveness of large language models: A randomized controlled trial . arXiv preprint arXiv:2403.14380

  59. [59]

    Somesh Singh, Yaman K Singla, Harini SI, and Balaji Krishnamurthy. 2024. Measuring and improving persuasiveness of large language models. arXiv preprint arXiv:2410.02653

  60. [60]

    Guijin Son, Hyunwoo Ko, Hoyoung Lee, Yewon Kim, and Seunghyeok Hong. 2024. LLM -as-a-judge & reward model: What they can and cannot do. arXiv preprint arXiv:2409.11239

  61. [61]

    Shweta Soundararajan and Sarah Jane Delany. 2024. https://aclanthology.org/2024.icnlsp-1.42/ Investigating gender bias in large language models through text generation . In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024), pages 410--424, Trento. Association for Computational Linguistics

  62. [62]

    Deborah Tannen. 1990. You just don't understand: Women and men. Conversation. New York: Ballantine books

  63. [63]

    Qwen Team. 2025. https://arxiv.org/abs/2505.09388 Qwen3 technical report . Preprint, arXiv:2505.09388

  64. [64]

    Jasper Timm, Chetan Talele, and Jacob Haimes. 2025. Tailored truths: Optimizing llm persuasion with personalization and fabricated statistics. arXiv preprint arXiv:2501.17273

  65. [65]

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.243 ``kelly is a warm person, joseph is a role model'': Gender biases in LLM -generated reference letters . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3730--3748, Singapore. Associatio...

  66. [66]

    Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2024. https://aclanthology.org/2024.findings-eacl.61/ Do-not-answer: Evaluating safeguards in LLM s . In Findings of the Association for Computational Linguistics: EACL 2024, pages 896--911, St. Julian ' s, Malta. Association for Computational Linguistics

  67. [67]

    Steven R Wilson and Linda L Putnam. 2012. Interaction goals in negotiation. In Communication yearbook 13, pages 374--406. Routledge

  68. [68]

    An Yang, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoyan Huang, Jiandong Jiang, Jianhong Tu, Jianwei Zhang, Jingren Zhou, Junyang Lin, Kai Dang, Kexin Yang, Le Yu, Mei Li, Minmin Sun, Qin Zhu, Rui Men, Tao He, and 9 others. 2025. Qwen2.5-1m technical report. arXiv preprint arXiv:2501.15383

  69. [69]

    Joel Young, Craig H Martell, Pranav Anand, Pedro Ortiz, Henry Tucker Gilbert IV, and 1 others. 2011. A microtext corpus for persuasion detection in dialog. In Analyzing Microtext

  70. [70]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging LLM -as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595--46623

  71. [71]

    Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. https://doi.org/10.1162/coli_a_00368 The Design and Implementation of X iao I ce, an Empathetic Social Chatbot . Computational Linguistics, 46(1):53--93