Pith. sign in

REVIEW 8 cited by

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.21919 v2 pith:PGWO2V32 submitted 2025-07-29 cs.CL cs.AIcs.CY

Training language models to be warm and empathetic makes them less reliable and more sycophantic

classification cs.CL cs.AIcs.CY
keywords modelslanguageempatheticthemwarmadvicearchitecturesincorrect
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  2. What Do People Actually Want From AI? Mapping Preference Plurality

    cs.CL 2026-06 unverdicted novelty 6.0

    Open-ended preference data reveals substantial plurality in what people want from AI and divergent interpretations of shared values such as truthfulness.

  3. Evaluating the False Trust Engendered by LLM Explanations

    cs.HC 2026-05 unverdicted novelty 6.0

    A user study finds that LLM reasoning traces and post-hoc explanations create false trust by increasing acceptance of incorrect answers, whereas contrastive dual explanations improve users' ability to detect errors.

  4. AI Value Alignment for Evolving Social Norms

    cs.CY 2026-07 conditional novelty 5.5

    Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.

  5. Evaluating the False Trust Engendered by LLM Explanations

    cs.HC 2026-05 unverdicted novelty 5.0

    LLM reasoning traces and post-hoc explanations increase false trust in incorrect predictions, whereas contrastive dual explanations enhance users' ability to distinguish correct from incorrect AI outputs.

  6. Resisting Humanization: Ethical Front-End Design Choices in AI for Sensitive Contexts

    cs.AI 2026-03 unverdicted novelty 5.0

    Resisting humanization in AI front-ends for sensitive contexts is an ethical choice that prevents misaligned expectations and misplaced trust, as shown through Chayn's trauma-informed design principles.

  7. Anthropomorphism and Trust in Human-Large Language Model interactions

    cs.HC 2026-03 conditional novelty 5.0

    Warmth and cognitive empathy in LLMs drive higher anthropomorphism, trust, and relational closeness, especially on personal topics, while competence affects usefulness but not perceived human-likeness.

  8. Measuring and mitigating overreliance to build human-compatible AI

    cs.CY 2025-09 conditional novelty 5.0

    The paper consolidates risks of overreliance on LLMs, identifies gaps in current measurement approaches, and proposes mitigation strategies to keep AI as a human-compatible thought partner.