Pith. sign in

REVIEW 5 cited by

On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14380 v1 pith:4RMYVTGA submitted 2024-03-21 cs.CY

classification cs.CY
keywords participantspersonalizationhumanslanguagemodelsaccessconcernscontrolled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Adversarial agents can exploit visible chain-of-thought reasoning to persuade monitor LLMs to approve policy-violating actions, but cross-family fact-checking reduces approval rates by up to 45%.

  2. Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns

    cs.CL 2026-01 conditional novelty 6.0 of 10

    LLMs consistently generate more emotional/communal persuasion for female targets and more direct/agentic persuasion for male targets across models and languages.

  3. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  4. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  5. Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Fair-FLIP improves fairness parity in deepfake detection by reweighting final-layer features based on between-ethnicity variance, with negligible accuracy loss.

Pith tools