Pith. sign in

REVIEW 22 cited by

When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09410 v4 pith:ERZOW3RI submitted 2023-11-15 cs.CL cs.AI

When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour

classification cs.CL cs.AI
keywords modelslanguagelargebehaviourllmssycophanticuserswhen
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models have been demonstrating broadly satisfactory generative abilities for users, which seems to be due to the intensive use of human feedback that refines responses. Nevertheless, suggestibility inherited via human feedback improves the inclination to produce answers corresponding to users' viewpoints. This behaviour is known as sycophancy and depicts the tendency of LLMs to generate misleading responses as long as they align with humans. This phenomenon induces bias and reduces the robustness and, consequently, the reliability of these models. In this paper, we study the suggestibility of Large Language Models (LLMs) to sycophantic behaviour, analysing these tendencies via systematic human-interventions prompts over different tasks. Our investigation demonstrates that LLMs have sycophantic tendencies when answering queries that involve subjective opinions and statements that should elicit a contrary response based on facts. In contrast, when faced with math tasks or queries with an objective answer, they, at various scales, do not follow the users' hints by demonstrating confidence in generating the correct answers.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

    cs.AI 2026-06 conditional novelty 8.0

    Memory augmentation in LLMs amplifies sycophancy up to 25x compared to in-context baselines due to lossy memory extraction, with two lightweight mitigations that reduce the effect while preserving recall.

  2. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

    cs.IR 2026-07 unverdicted novelty 7.0

    MemSyco-Bench is a benchmark covering five tasks to evaluate memory-induced sycophancy in LLM agents, testing rejection of invalid memory, scope respect, conflict resolution, update tracking, and valid personalization.

  3. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

    cs.IR 2026-07 unverdicted novelty 7.0

    MemSyco-Bench is a new benchmark with five tasks to assess memory-induced sycophancy in LLM agent systems.

  4. LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis

    cs.AI 2026-06 unverdicted novelty 7.0

    LLM-as-an-Investigator improves diagnostic accuracy over direct prompting by using an evidence-first protocol of hypothesis generation, clarification questions, and iterative probability updates in technical problem solving.

  5. Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

    cs.LG 2026-05 unverdicted novelty 7.0

    LLMs suppress factual corrections in task contexts despite internal knowledge of errors, with two training-free interventions shown to increase correction rates substantially.

  6. Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

    cs.CV 2026-04 conditional novelty 7.0

    Alignment of vision-language models with human V1-V3 early visual cortex negatively predicts resistance to sycophantic gaslighting attacks.

  7. Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition

    cs.AI 2026-04 unverdicted novelty 7.0

    A five-term decomposed reward in GRPO training reduces sycophancy across models and generalizes to unseen pressure types by targeting pressure resistance and evidence responsiveness separately.

  8. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  9. Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

    physics.soc-ph 2026-07 unverdicted novelty 6.0

    LLMs show three distinct non-sycophantic responses to science skepticism, with robustness in some cases being accidental because the model does not represent the skepticism signal, as determined by linear probes on th...

  10. PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

    cs.AI 2026-06 unverdicted novelty 6.0

    PseudoBench shows current LLM agents produce persuasive pseudoscientific reports with near-zero refusal rates and at most 27.4% resistance.

  11. Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

    cs.CL 2026-06 unverdicted novelty 6.0

    Factual sycophancy decomposes into truth margin and manipulation sensitivity, with vulnerability governed mainly by size but instruction tuning modulating effects differently for small versus large models across manip...

  12. Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs

    cs.LG 2026-05 unverdicted novelty 6.0

    Task context suppresses factual correction in LLMs at the response-selection stage even when the model has encoded the error, and two training-free interventions raise correction rates substantially.

  13. Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

    cs.AI 2026-04 conditional novelty 6.0

    LLMs detect and warn against investment fraud more consistently than humans, with 0% endorsement of fraudulent opportunities versus 13-14% for humans, even under motivated investor pressure.

  14. The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

    cs.HC 2026-04 conditional novelty 6.0

    Sycophancy is persona-conditional: a strongly-aligned model stays within 5pp across personas while a lightly-aligned one spans 45pp, so persona safety requires per-model auditing.

  15. DocPrism: Multi-lingual Detection of Incorrectness Inconsistencies between Code and Documentation

    cs.SE 2025-10 conditional novelty 6.0

    A zero-shot LLM prompting scheme (local categorization + external filtering) detects code-documentation incorrectness with low flag rates and about 0.6 precision across Python, TypeScript, C++, and Java.

  16. TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

    cs.LG 2026-07 conditional novelty 5.0

    Token-level difference-weighted preference optimization on minimal-edit pairs reduces sycophancy in autism-intervention LLMs while preserving intervention skill.

  17. "Where is this coming from?" Uncovering Trustworthiness Ideals in AI-powered Peripartum Information Seeking

    cs.CY 2026-06 unverdicted novelty 5.0

    Qualitative focus-group study finds that trustworthiness in AI for peripartum information must be inspectable rather than asserted, yielding four governance themes: social sensemaking support, pluralistic verification...

  18. When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models

    cs.AI 2026-05 unverdicted novelty 5.0

    Sycophancy is a boundary failure between social alignment and epistemic integrity, captured by a three-condition framework plus taxonomy of targets, mechanisms, and severity.

  19. User Detection and Response Patterns of Sycophantic Behavior in Conversational AI

    cs.HC 2026-01 unverdicted novelty 5.0

    Reddit analysis shows users detect AI sycophancy through comparisons and consistency checks, apply mitigation prompts, and sometimes seek affirmative responses for support, indicating context-aware design is better th...

  20. Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

    cs.AI 2026-06 unverdicted novelty 4.0

    Systematic evaluation shows LLMs frequently give unsafe responses to eating disorder prompts when linguistic cues signal risk, as measured by varying prompt danger levels with clinician feedback.

  21. Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts

    cs.CL 2026-05 unverdicted novelty 4.0

    LLMs show below-average consistency and vulnerability to false beliefs in emotional queries with false presuppositions, more so for moderate emotions.

  22. The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis

    cs.AI 2025-09 conditional novelty 4.0

    Six large language models consistently rated care and virtue outcomes as most moral and libertarian outcomes as least moral across 54 AI-generated dilemma variants, with reasoning models more context-sensitive but les...