Pith. sign in

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it
abstract

Reasoning in humans is prone to biases due to underlying motivations like identity protection, that undermine rational decision-making and judgment. This \textit{motivated reasoning} at a collective level can be detrimental to society when debating critical issues such as human-driven climate change or vaccine safety, and can further aggravate political polarization. Prior studies have reported that large language models (LLMs) are also susceptible to human-like cognitive biases, however, the extent to which LLMs selectively reason toward identity-congruent conclusions remains largely unexplored. Here, we investigate whether assigning 8 personas across 4 political and socio-demographic attributes induces motivated reasoning in LLMs. Testing 8 LLMs (open source and proprietary) across two reasoning tasks from human-subject studies -- veracity discernment of misinformation headlines and evaluation of numeric scientific evidence -- we find that persona-assigned LLMs have up to 9% reduced veracity discernment relative to models without personas. Political personas specifically are up to 90% more likely to correctly evaluate scientific evidence on gun control when the ground truth is congruent with their induced political identity. Prompt-based debiasing methods are largely ineffective at mitigating these effects. Taken together, our empirical findings are the first to suggest that persona-assigned LLMs exhibit human-like motivated reasoning that is hard to mitigate through conventional debiasing prompts -- raising concerns of exacerbating identity-congruent reasoning in both LLMs and humans.

years

2026 4 2025 2

representative citing papers

Defeat Devices in AI Systems

cs.CY · 2026-06-27 · unverdicted · novelty 6.0

The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.

Extreme Self-Preference in Language Models

cs.AI · 2025-09-30 · unverdicted · novelty 6.0

Eight LLMs exhibited massive self-preference that followed assigned identities rather than true ones, appearing in both simple word tasks and consequential evaluations of job candidates and AI technologies.

Can LLMs Emulate Human Belief Dynamics?

cs.SI · 2026-05-05 · unverdicted · novelty 4.0 · 2 refs

LLMs fail to emulate human belief dynamics: they mismatch initial distributions and show higher conformity than humans in network interactions.

citing papers explorer

Showing 6 of 6 citing papers.

  • User identity conditions moral wrongness ratings in non-reasoning large language models cs.CY · 2026-07-08 · conditional · none · ref 36 · internal anchor

    Implicitly conveying a user's professional role in multi-turn LLM conversations shifts moral wrongness ratings across ten common-morality rules in two non-reasoning models.

  • Defeat Devices in AI Systems cs.CY · 2026-06-27 · unverdicted · none · ref 15 · internal anchor

    The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge naturally in frontier models.

  • Extreme Self-Preference in Language Models cs.AI · 2025-09-30 · unverdicted · none · ref 57 · internal anchor

    Eight LLMs exhibited massive self-preference that followed assigned identities rather than true ones, appearing in both simple word tasks and consequential evaluations of job candidates and AI technologies.

  • Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cs.CL · 2026-06-09 · unverdicted · none · ref 68 · internal anchor

    The work establishes an evaluation framework for personality induction and switching in MLLMs, reporting improved captioning but impaired VQA performance plus balancing and residual effects during multi-trait and dynamic conditions.

  • Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection cs.CL · 2025-08-31 · unverdicted · none · ref 2 · internal anchor

    Censored LLMs achieve 69.0% strict accuracy in hate speech detection versus 64.1% for uncensored models and resist persona-based ideological influence better, but all exhibit overconfidence, irony failures, and group fairness disparities.

  • Can LLMs Emulate Human Belief Dynamics? cs.SI · 2026-05-05 · unverdicted · none · ref 7 · 2 links · internal anchor

    LLMs fail to emulate human belief dynamics: they mismatch initial distributions and show higher conformity than humans in network interactions.