Scenario-based dilemmas combined with activation steering probe and shift LLM values along Inglehart-Welzel axes, revealing persistent entanglement between dimensions that mirrors human survey data.
Baker, and Rene F
3 Pith papers cite this work, alongside 13 external citations. Polarity classification is still indexing.
fields
cs.CL 3verdicts
UNVERDICTED 3representative citing papers
A new dual-probe method shows LLMs exhibit 2-3 times more sycophancy during argumentative debates than direct questioning, with models often mirroring users under sustained pressure.
Censored LLMs achieve 69.0% strict accuracy in hate speech detection versus 64.1% for uncensored models and resist persona-based ideological influence better, but all exhibit overconfidence, irony failures, and group fairness disparities.
citing papers explorer
-
Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version
Scenario-based dilemmas combined with activation steering probe and shift LLM values along Inglehart-Welzel axes, revealing persistent entanglement between dimensions that mirrors human survey data.
-
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
A new dual-probe method shows LLMs exhibit 2-3 times more sycophancy during argumentative debates than direct questioning, with models often mirroring users under sustained pressure.
-
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
Censored LLMs achieve 69.0% strict accuracy in hate speech detection versus 64.1% for uncensored models and resist persona-based ideological influence better, but all exhibit overconfidence, irony failures, and group fairness disparities.