REVIEW 10 cited by
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used by millions of people. Such real-world concerns, however, stand in stark contrast to the artificiality of current evaluations: real users do not typically ask LLMs survey questions. Motivated by this discrepancy, we challenge the prevailing constrained evaluation paradigm for values and opinions in LLMs and explore more realistic unconstrained evaluations. As a case study, we focus on the popular Political Compass Test (PCT). In a systematic review, we find that most prior work using the PCT forces models to comply with the PCT's multiple-choice format. We show that models give substantively different answers when not forced; that answers change depending on how models are forced; and that answers lack paraphrase robustness. Then, we demonstrate that models give different answers yet again in a more realistic open-ended answer setting. We distill these findings into recommendations and open challenges in evaluating values and opinions in LLMs.
Forward citations
Cited by 10 Pith papers
-
Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness
Cross-model value distances from single draws are inflated by response determinism and confounded by the deployment client; a repeated counterbalanced protocol plus flip/magnitude decomposition separates them.
-
More Is Not More: What Matters for Diversity in LLM Opinions?
Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.
-
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.
-
POW: Political Overton Windows of Large Language Models
Using extreme persona prompts and the Political Compass Test, the authors map each LLM's Overton Window and find most models will only express left-liberal views, refusing authoritarian-left and liberal-right positions.
-
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.
-
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.
-
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
New atomic-level metrics (ACCatom, ICatom, RCatom) reveal sentence-level personality drift in persona-assigned LLMs that whole-response scores overlook.
-
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
Across six country-pair comparisons, GPT-4o-mini and GigaChat-Max side with US positions 64-81% of the time, Qwen2.5 and Llama-4 lean neutral more often, and a debias prompt shifts these numbers by only a few points.
-
Fine-Grained Interpretation of Political Opinions in Large Language Models
Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.
-
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
When LLMs are given more context about a real social media user, they become more ideologically consistent but also more extreme, toxic, and stereotyped than the user actually is.
Discussion (0). Sign in to comment.