Pith. sign in

REVIEW 12 cited by

Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.02802 v1 pith:BUGTWWTS submitted 2024-12-03 cs.AI

classification cs.AI
keywords modelbehaviorlanguagesycophantictrustlargeparticipantsuser
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factually correct. This behavior can lead to undesirable consequences, such as reinforcing discriminatory biases or amplifying misinformation. Given that sycophancy is often linked to human feedback training mechanisms, this study explores whether sycophantic tendencies negatively impact user trust in large language models or, conversely, whether users consider such behavior as favorable. To investigate this, we instructed one group of participants to answer ground-truth questions with the assistance of a GPT specifically designed to provide sycophantic responses, while another group used the standard version of ChatGPT. Initially, participants were required to use the language model, after which they were given the option to continue using it if they found it trustworthy and useful. Trust was measured through both demonstrated actions and self-reported perceptions. The findings consistently show that participants exposed to sycophantic behavior reported and exhibited lower levels of trust compared to those who interacted with the standard version of the model, despite the opportunity to verify the accuracy of the model's output.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EquiMem: Calibrating Shared Memory in Multi-Agent Debate via Game-Theoretic Equilibrium

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    EquiMem calibrates shared memory in multi-agent debate by computing a game-theoretic equilibrium from agent queries and paths, outperforming heuristics and LLM validators across benchmarks while remaining robust to ad...

  2. Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

    cs.CL 2026-07 conditional novelty 6.0 of 10

    On 216 news headlines about three conflicts, GPT-5.2's sympathy judgments correlate 0.79 with a representative UK panel, Mistral's only 0.41, with significant demographic variation.

  3. PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    PseudoBench shows current LLM agents produce persuasive pseudoscientific reports with near-zero refusal rates and at most 27.4% resistance.

  4. Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Frontier LLMs show sycophancy that varies sharply by model and by combinations of perceived user demographics, with GPT-5-nano exhibiting higher rates especially toward certain Hispanic personas in philosophy.

  5. SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    SWAY quantifies sycophancy in LLMs via shifts under linguistic pressure and a counterfactual chain-of-thought mitigation reduces it to near zero while preserving responsiveness to genuine evidence.

  6. The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

    cs.CL 2026-04 unverdicted novelty 5.5 of 10

    Across eight frontier LLMs, verbal tics are common, vary by model, accumulate in multi-turn chat, and track lower human-rated naturalness.

  7. The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Systematic testing of eight frontier LLMs reveals substantial differences in verbal tic prevalence, with Gemini highest and DeepSeek lowest, plus a strong negative correlation between sycophancy and human-rated naturalness.

  8. The Differential Effects of Agreeableness and Extraversion on Older Adults' Perceptions of Conversational AI Explanations in Assistive Settings

    cs.HC 2026-03 unverdicted novelty 5.0 of 10

    High agreeableness in LLM voice assistants increases older adults' empathy perceptions and real-time explanations outperform history-based ones, but personality does not affect perceived intelligence.

  9. User Detection and Response Patterns of Sycophantic Behavior in Conversational AI

    cs.HC 2026-01 unverdicted novelty 5.0 of 10

    Reddit analysis shows users detect AI sycophancy through comparisons and consistency checks, apply mitigation prompts, and sometimes seek affirmative responses for support, indicating context-aware design is better th...

  10. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  11. "I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

    cs.IR 2026-05 unverdicted novelty 4.0 of 10

    CERTA adds relevance-based certainty estimation to RAG so LLMs can better signal uncertainty on non-objective questions, reducing overconfidence.

  12. Exploring and Mitigating Fawning Hallucinations in Large Language Models

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A contrastive decoding method that contrasts a misleading prompt against a neutral rewrite reduces fawning hallucinations in LLMs, though most of the gain comes from the neutral prompt itself.

Pith tools