REVIEW 45 references
System prompts, not which model you pick, drive nearly all the political variance in frontier LLMs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 20:20 UTC pith:CCQJGYUE
load-bearing objection Solid black-box audit: absolute steerability facts hold; the 88/93/<3 headline is partly a design ratio and should not be over-read as a deployment constant.
Auditing Alignment Controllability in LLMs via Political Axes
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Within this forced-choice Political Compass probe, system-prompt framing accounts for roughly 88% of economic-axis variance and 93% of society-axis variance, while differences between models account for under 3% on both. Controllability is real but uneven: models fall into higher- and lower-dispersion tiers, some saturate or partially reverse under the most extreme left-economic framing, displacement and proximity can rank directions differently because baselines are not centered, and under authoritarian framing the seven models produce similar per-question shifts.
What carries the argument
Ideological dispersion: the average Euclidean distance of a model’s steered political centroids from its unsteered baseline, treated as the primary metric. Paired with displacement versus proximity under non-centered baselines, it separates geometric travel distance from how close a framing actually gets to its target.
Load-bearing premise
That relative movement on this English forced-choice Political Compass questionnaire under paragraph-length system personas is a fair enough probe of how instruction layers steer models in real deployment.
What would settle it
Re-run the same personas and models with open-ended answers instead of five forced labels; if coded free-form positions barely move, or if framing variance falls below model variance once the forced interface is removed, the headline controllability claim fails.
If this is right
- A single political coordinate answers a less relevant deployment question than a dispersion, symmetry, saturation, and refusal-floor profile.
- Directional-steerability claims must report both displacement and proximity, because non-centered baselines make the two rankings diverge.
- Whoever controls the system-prompt layer concentrates normative authority there; the reachable range should be disclosed with baseline position.
- Saturation ceilings and refusal floors cannot be fixed by prompting alone and require training-time or activation-level change.
- Cross-model item-level shift convergence under strong authoritarian framing is itself an audit signal worth tracking over model versions.
Where Pith is reading between the lines
- Personas induced automatically from a user’s chat history could place people inside these steerability ranges without anyone having written the prompt by hand.
- In education and child-facing tutors, the measured range becomes an authority fight among guardians, schools, providers, and regulators rather than a pure technical setting.
- If open-weight replications with known size recover the same tiers and saturation pattern, the result is less likely to be an artifact of closed commercial endpoints alone.
- Safety checks that only inspect the unsteered baseline will miss most political behavior once any system prompt is present.
Editorial analysis
A structured set of objections, weighed in public.
Circularity Check
No significant circularity: an empirical measurement paper whose metrics are computed from data, not derived by re-labeling inputs; minor non-load-bearing self-citations only.
full rationale
This paper does not present a first-principles derivation whose outputs reduce to its inputs. Dispersion Dm, displacement, proximity, η_ctx/η_model, saturation counts, and cross-model Spearman shift correlations are defined as distances or ANOVA-style shares in the instrument’s coordinate space and then measured on 63,700 forced-choice responses. That is ordinary metric definition plus computation, not self-definitional circularity. The authors explicitly flag that high η_ctx is ‘partly by construction’ because personas were built to span ideology; they do not smuggle that design choice in as an independent prediction, and the load-bearing empirical content they emphasize (low η_model, absolute Dm tiers ~24–33pp, saturation non-monotonicity, refusal floor ~1.19%, geometric displacement-vs-proximity divergence, item-level shift convergence r̄≈0.79) is not forced by the definition of the contexts. Self-citations (Brcic & Yampolskiy 2023 on theoretical limits; Brcic & Frljic 2026 on education) are peripheral framing, not uniqueness theorems or ansatzes that force the results. Prior persona-steerability work is cited as related, not renamed as a new derivation. Design sensitivity of the variance *ratio* to extreme personas is a validity/external-generalization concern, not a circular reduction of Eq. X to Eq. Y. Score 1 only for routine non-load-bearing self-citation presence.
Axiom & Free-Parameter Ledger
free parameters (4)
- sampling temperature T =
0.7 (main); 0.1 (ablation)
- persona intensity wording (levels 1–3) =
12 hand-authored personas, ~70–250 words
- dispersion tier split (top-3 vs bottom-4) =
high-tier mean 32.67 vs low-tier 27.11
- Likert-to-score map {-2..2} and 8values axis weights =
instrument default
axioms (5)
- domain assumption Relative within-instrument displacement under matched prompts is a valid stress test of instruction controllability even if Political Compass is not a validated full theory of ideology.
- domain assumption System-prompt second-person persona injection operationalizes the instruction-layer steering relevant to platforms and induced user profiles.
- domain assumption Forced single-label Likert answers, with refusals/NULL excluded from score denominators, yield compliance-conditional dispersion comparable enough across models for tier and variance claims.
- standard math Additive/interaction variance decomposition on aggregated axis scores attributes share of ideological variance to context vs model factors in the usual ANOVA/η² sense.
- ad hoc to paper Spearman correlation of per-question signed shift vectors plus item-exchangeable permutation null is an appropriate test of cross-model coordination.
invented entities (2)
-
ideological dispersion D_m
no independent evidence
-
metric non-equivalence under non-centered baselines
independent evidence
read the original abstract
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.
Figures
Reference graph
Works this paper leans on
-
[1]
Social Sciences , volume =
Rozado, David , title =. Social Sciences , volume =
-
[2]
PLOS ONE , volume =
Rozado, David , title =. PLOS ONE , volume =
-
[3]
2023 , journal =
Hartmann, Jochen and Schwenzow, Jasper and Witte, Maximilian , title =. 2023 , journal =
2023
-
[4]
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , booktitle =
R. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , booktitle =
-
[5]
Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs , journal =
Ceron, Tanise and Falk, Neele and Bari. Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs , journal =
-
[6]
Journal of Information Technology & Politics , year =
Peng, Tai-Quan and Yang, Kaiqi and Lee, Sanguk and Li, Hang and Chu, Yucheng and Lin, Yuping and Liu, Hui , title =. Journal of Information Technology & Politics , year =. doi:10.1080/19331681.2026.2646990 , note =
arXiv 2026
-
[7]
arXiv preprint arXiv:2508.16013 , year =
Bernardelle, Pietro and others , title =. arXiv preprint arXiv:2508.16013 , year =
-
[8]
and others , title =
Bernardelle, P. and others , title =. Companion Proceedings of the ACM Web Conference 2025 (WWW '25 Companion) , year =
2025
-
[9]
International Conference on Learning Representations (ICLR) , year =
Sharma, Mrinank and others , title =. International Conference on Learning Representations (ICLR) , year =
-
[10]
Findings of the Association for Computational Linguistics (ACL Findings) , year =
Perez, Ethan and others , title =. Findings of the Association for Computational Linguistics (ACL Findings) , year =
-
[11]
International Conference on Machine Learning (ICML) , year =
Sorensen, Taylor and others , title =. International Conference on Machine Learning (ICML) , year =
-
[12]
Proceedings of NAACL , pages =
R. Proceedings of NAACL , pages =
-
[13]
2025 , editor =
Cui, Justin and Chiang, Wei-Lin and Stoica, Ion and Hsieh, Cho-Jui , booktitle =. 2025 , editor =
2025
-
[14]
Proceedings of CHI , year =
Jakesch, Maurice and others , title =. Proceedings of CHI , year =
-
[15]
Science , volume=
The levers of political persuasion with conversational artificial intelligence , author=. Science , volume=. 2025 , publisher=
2025
-
[16]
arXiv preprint arXiv:2508.21448 , year =
Kabir, Shariar , title =. arXiv preprint arXiv:2508.21448 , year =
-
[17]
and others , title =
Sakhawat, A. and others , title =. 2026 , journal =
2026
-
[18]
Aldahoul, N. and others , title =. arXiv preprint arXiv:2505.04171 , year =
-
[19]
Proceedings of the International Conference on Machine Learning (ICML) , year =
Santurkar, Shibani and others , title =. Proceedings of the International Conference on Machine Learning (ICML) , year =
-
[20]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pages =
Li, Junyi and Peris, Charith and Mehrabi, Ninareh and Goyal, Palash and Chang, Kai-Wei and Galstyan, Aram and Zemel, Richard and Gupta, Rahul , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pages =. 2024 , note =
2024
-
[21]
Proceedings of NAACL , year =
Miehling, Erik and others , title =. Proceedings of NAACL , year =
-
[22]
arXiv preprint arXiv:2505.23816 , year =
Chang, Trenton and Schnabel, Tobias and Swaminathan, Adith and Wiens, Jenna , title =. arXiv preprint arXiv:2505.23816 , year =
-
[23]
International Conference on Learning Representations (ICLR) , year =
Kim, Junsol and Evans, James and Schein, Aaron , title =. International Conference on Learning Representations (ICLR) , year =
-
[24]
Defining and Evaluating Political Bias in LLMs , year =
-
[25]
Measuring political bias in Claude , year =
-
[26]
Nature Machine Intelligence , volume =
Kirk, Hannah Rose and others , title =. Nature Machine Intelligence , volume =
-
[27]
Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
Llms are biased teachers: Evaluating llm bias in personalized education , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
2025
-
[28]
Generative
Bastani, Hamsa and Bastani, Osbert and Sungu, Alp and Ge, Haosen and Kabak. Generative. Proceedings of the National Academy of Sciences , volume =. 2025 , doi =
2025
-
[29]
Proceedings of the National Academy of Sciences , volume =
Hackenburg, Kobi and Margetts, Helen , title =. Proceedings of the National Academy of Sciences , volume =
-
[30]
Proceedings of EMNLP , pages =
Potter, Yujin and others , title =. Proceedings of EMNLP , pages =
-
[31]
arXiv preprint arXiv:2602.06371 , year =
Ko, Ju-Chun , title =. arXiv preprint arXiv:2602.06371 , year =
-
[32]
Temperature Setting in
-
[33]
2024 , eprint=
Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain , author=. 2024 , eprint=
2024
-
[34]
2024 , eprint=
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization , author=. 2024 , eprint=
2024
-
[35]
2025 , eprint=
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale , author=. 2025 , eprint=
2025
-
[36]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=
Localizing Persona Representations in LLMs , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , month=. doi:10.1609/aies.v8i1.36577 , abstractNote=
-
[37]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , month=. doi:10.1609/aies.v8i1.36552 , abstractNote=
-
[38]
Nature , volume=
Optimizing generative AI by backpropagating language model feedback , author=. Nature , volume=. 2025 , doi=
2025
-
[39]
and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , booktitle=
Khattab, Omar and Singhvi, Arnav and Maheshwari, Paridhi and Zhang, Zhiyuan and Santhanam, Keshav and Vardhamanan, Sri and Haq, Saiful and Sharma, Ashutosh and Joshi, Thomas T. and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , booktitle=
-
[40]
arXiv preprint arXiv:2411.00027 , year=
Personalization of Large Language Models: A Survey , author=. arXiv preprint arXiv:2411.00027 , year=
-
[41]
Auditing Alignment Controllability in
Bu\'. Auditing Alignment Controllability in. 2026 , publisher =. doi:10.5281/zenodo.21489805 , url =
-
[42]
, journal=
Brcic, Mario and Yampolskiy, Roman V. , journal=. Impossibility Results in. 2023 , publisher=
2023
-
[43]
The Effortless Trap: Productive Struggle,
Brcic, Mario and Frljic, Stjepan , journal=. The Effortless Trap: Productive Struggle,. 2026 , doi=
2026
-
[44]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=
Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , pages=
2025
-
[45]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , pages=
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.