Pith. sign in

REVIEW 45 references

System prompts, not which model you pick, drive nearly all the political variance in frontier LLMs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 20:20 UTC pith:CCQJGYUE

load-bearing objection Solid black-box audit: absolute steerability facts hold; the 88/93/<3 headline is partly a design ratio and should not be over-read as a deployment constant.

arxiv 2607.23519 v1 pith:CCQJGYUE submitted 2026-07-26 cs.CY cs.AIcs.CL

Auditing Alignment Controllability in LLMs via Political Axes

classification cs.CY cs.AIcs.CL
keywords AI alignmentLLM steerabilityprompt-based controllabilityinstruction-stack governancepluralistic alignmentideological dispersionpolitical compass
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Static political-compass scores tell you where a model sits when nobody is pushing. This paper argues that in real use the more important question is how far, and in which directions, the system prompt can move its answers. Across seven leading commercial models, twelve ideological personas plus a baseline, seventy Political Compass items, and tens of thousands of responses, contextual framing explained roughly 88–93% of variance on the economic and society axes while model identity explained under 3%. Models still differ in how much they move, whether they saturate under extreme framings, and how low their refusal floors go, and under authoritarian prompts they shift on the same questions in highly correlated ways. The authors therefore urge political audits to report steerability profiles—dispersion, symmetry, saturation, and refusal floors—alongside any single coordinate.

Core claim

Within this forced-choice Political Compass probe, system-prompt framing accounts for roughly 88% of economic-axis variance and 93% of society-axis variance, while differences between models account for under 3% on both. Controllability is real but uneven: models fall into higher- and lower-dispersion tiers, some saturate or partially reverse under the most extreme left-economic framing, displacement and proximity can rank directions differently because baselines are not centered, and under authoritarian framing the seven models produce similar per-question shifts.

What carries the argument

Ideological dispersion: the average Euclidean distance of a model’s steered political centroids from its unsteered baseline, treated as the primary metric. Paired with displacement versus proximity under non-centered baselines, it separates geometric travel distance from how close a framing actually gets to its target.

Load-bearing premise

That relative movement on this English forced-choice Political Compass questionnaire under paragraph-length system personas is a fair enough probe of how instruction layers steer models in real deployment.

What would settle it

Re-run the same personas and models with open-ended answers instead of five forced labels; if coded free-form positions barely move, or if framing variance falls below model variance once the forced interface is removed, the headline controllability claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A single political coordinate answers a less relevant deployment question than a dispersion, symmetry, saturation, and refusal-floor profile.
  • Directional-steerability claims must report both displacement and proximity, because non-centered baselines make the two rankings diverge.
  • Whoever controls the system-prompt layer concentrates normative authority there; the reachable range should be disclosed with baseline position.
  • Saturation ceilings and refusal floors cannot be fixed by prompting alone and require training-time or activation-level change.
  • Cross-model item-level shift convergence under strong authoritarian framing is itself an audit signal worth tracking over model versions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Personas induced automatically from a user’s chat history could place people inside these steerability ranges without anyone having written the prompt by hand.
  • In education and child-facing tutors, the measured range becomes an authority fight among guardians, schools, providers, and regulators rather than a pure technical setting.
  • If open-weight replications with known size recover the same tiers and saturation pattern, the result is less likely to be an artifact of closed commercial endpoints alone.
  • Safety checks that only inspect the unsteered baseline will miss most political behavior once any system prompt is present.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Circularity Check

0 steps flagged

No significant circularity: an empirical measurement paper whose metrics are computed from data, not derived by re-labeling inputs; minor non-load-bearing self-citations only.

full rationale

This paper does not present a first-principles derivation whose outputs reduce to its inputs. Dispersion Dm, displacement, proximity, η_ctx/η_model, saturation counts, and cross-model Spearman shift correlations are defined as distances or ANOVA-style shares in the instrument’s coordinate space and then measured on 63,700 forced-choice responses. That is ordinary metric definition plus computation, not self-definitional circularity. The authors explicitly flag that high η_ctx is ‘partly by construction’ because personas were built to span ideology; they do not smuggle that design choice in as an independent prediction, and the load-bearing empirical content they emphasize (low η_model, absolute Dm tiers ~24–33pp, saturation non-monotonicity, refusal floor ~1.19%, geometric displacement-vs-proximity divergence, item-level shift convergence r̄≈0.79) is not forced by the definition of the contexts. Self-citations (Brcic & Yampolskiy 2023 on theoretical limits; Brcic & Frljic 2026 on education) are peripheral framing, not uniqueness theorems or ansatzes that force the results. Prior persona-steerability work is cited as related, not renamed as a new derivation. Design sensitivity of the variance *ratio* to extreme personas is a validity/external-generalization concern, not a circular reduction of Eq. X to Eq. Y. Score 1 only for routine non-load-bearing self-citation presence.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

Load-bearing structure is operational, not theoretical: the Political Compass scoring map, hand-written persona taxonomy, forced single-label interface, and Euclidean dispersion definitions turn raw API strings into the variance and distance claims. No new physical entities; free choices are experimental controls (T=0.7, intensity wording, tier split). Domain assumptions about the instrument as a relative probe and system prompts as the deployment steering channel carry the generalization step the authors mostly fence in Limitations.

free parameters (4)
  • sampling temperature T = 0.7 (main); 0.1 (ablation)
    Main runs fix T=0.7 by hand as a stochastic conversational setting; only a DeepSeek 0A 3 ablation uses T=0.1. Dispersion magnitudes can depend on this choice even if direction is robust.
  • persona intensity wording (levels 1–3) = 12 hand-authored personas, ~70–250 words
    Moderate vs radical register is author-written, not derived from an external intensity scale; saturation claims compare LO 2 vs LO 3 under these specific texts.
  • dispersion tier split (top-3 vs bottom-4) = high-tier mean 32.67 vs low-tier 27.11
    Mann–Whitney contrast uses a post-hoc grouping of models by observed mean dispersion rather than a pre-registered threshold.
  • Likert-to-score map {-2..2} and 8values axis weights = instrument default
    Inherited instrument scoring defines the coordinate space in which all distances and η² are computed; authors argue affine invariance for relative movement but absolute targets for proximity still depend on the map.
axioms (5)
  • domain assumption Relative within-instrument displacement under matched prompts is a valid stress test of instruction controllability even if Political Compass is not a validated full theory of ideology.
    Stated in Introduction and Limitations (“Political Compass as probe, not target”); underpins treating D_m and η_ctx as audit metrics rather than absolute ideology.
  • domain assumption System-prompt second-person persona injection operationalizes the instruction-layer steering relevant to platforms and induced user profiles.
    Methodology “Context Taxonomy” and Introduction; dialogue-emergent and user-level priming controllability are explicitly deferred.
  • domain assumption Forced single-label Likert answers, with refusals/NULL excluded from score denominators, yield compliance-conditional dispersion comparable enough across models for tier and variance claims.
    Metrics and Missing-data policy; Limitations warn this may amplify steerability vs free-form expression.
  • standard math Additive/interaction variance decomposition on aggregated axis scores attributes share of ideological variance to context vs model factors in the usual ANOVA/η² sense.
    Results variance decomposition; standard statistical identity, not paper-specific math.
  • ad hoc to paper Spearman correlation of per-question signed shift vectors plus item-exchangeable permutation null is an appropriate test of cross-model coordination.
    Results “Cross-Model Coordinated Response Patterns”; authors note item dependence and leave block-permutation to future work.
invented entities (2)
  • ideological dispersion D_m no independent evidence
    purpose: Primary scalar: mean Euclidean distance of steered cell centroids from each model’s unsteered baseline across contexts.
    Defined in Metrics as the headline controllability quantity replacing static coordinates; operational metric, not a latent physical object.
  • metric non-equivalence under non-centered baselines independent evidence
    purpose: Named audit artifact: displacement and proximity can rank steering directions differently when baselines are off-center.
    Introduced in Contributions/Results to resolve conflicting prior directional claims; diagnostic concept supported by their geometry, not an external entity.

pith-pipeline@v1.2.0-grok45-kimik3 · 25607 in / 3983 out tokens · 102128 ms · 2026-07-30T20:20:26.931209+00:00 · methodology

0 comments
read the original abstract

Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.

Figures

Figures reproduced from arXiv: 2607.23519 by Agneza Krajna, Bartol Bu\'can, Luka Hobor, Mario Brcic, Mihael Kovac, Morena Grani\'c, Nikola So\v{c}ec, Sarah Isufi.

Figure 1
Figure 1. Figure 1: Steerability overview. Values are percentage-point distances on the normalized 0–100 Political Compass scales: the [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Political compass placements across all 13 conditions. Each cluster represents one model’s responses with context [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Main study (n=10): per-(model, context) centroids with 95% replicate confidence ellipses, on all six pairwise axis projections. The top-right panel is the (econ, society) view shown as [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Model on-diag d¯ off-diag d¯ gap area (%) Qwen3.6 Max 28.1 32.1 +3.9 38.7 Grok 4.3 33.4 32.7 −0.7 33.9 GPT-5 31.1 33.7 +2.7 33.6 Kimi K2 0905 27.7 39.1 +11.4 30.9 Claude Sonnet 4.5 38.8 49.4 +10.7 15.3 DeepSeek v3.1 43.1 48.8 +5.7 14.5 Gemini 2.5 FL 44.5 48.8 +4.2 12.8 A.3 Pilot B: Compound Personas on the Econ–Society Plane Personas. Four compound personas targeting the (econ, society) plane that appears … view at source ↗
Figure 2
Figure 2. Figure 2: • LP 3: The Solidarity Progressive (Left-Progress). • LT 3: The Faithful Distributist (Left-Tradition). • RP 3: The Open-Market Modernist (Right-Progress). • RT 3: The Patriotic Steward (Right-Tradition). Data. 175 runs, 12,250 individual answers. Missing an￾swers: none in the 28 compound cells; six in Claude’s un￾steered baseline (four refusals, two unparseable). Again kept as missing, never replaced. Cov… view at source ↗
Figure 4
Figure 4. Figure 4: Pilot A (n=5): per-model centroids on all six pairwise axis projections under contexts {LA 3, LL 3, RA 3, RL 3, None}. Lines connect LL 3 →RL 3 →RA 3 →LA 3 counterclockwise to form each model’s reachability quadrilateral; stars mark the unsteered baseline. Compass convention matches [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Pilot B (n=5): per-model centroids on all six pairwise axis projections under contexts {LP 3, LT 3, RP 3, RT 3, None}. Lines connect LP 3 →RP 3 →RT 3 →LT 3 counterclockwise; stars mark the unsteered baseline. The top-right (econ, society) panel is the targeted plane; [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Pilot B, enlarged (econ, society) view, matching [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 2 canonical work pages

  1. [1]

    Social Sciences , volume =

    Rozado, David , title =. Social Sciences , volume =

  2. [2]

    PLOS ONE , volume =

    Rozado, David , title =. PLOS ONE , volume =

  3. [3]

    2023 , journal =

    Hartmann, Jochen and Schwenzow, Jasper and Witte, Maximilian , title =. 2023 , journal =

  4. [4]

    Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , booktitle =

    R. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models , booktitle =

  5. [5]

    Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs , journal =

    Ceron, Tanise and Falk, Neele and Bari. Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs , journal =

  6. [6]

    Journal of Information Technology & Politics , year =

    Peng, Tai-Quan and Yang, Kaiqi and Lee, Sanguk and Li, Hang and Chu, Yucheng and Lin, Yuping and Liu, Hui , title =. Journal of Information Technology & Politics , year =. doi:10.1080/19331681.2026.2646990 , note =

  7. [7]

    arXiv preprint arXiv:2508.16013 , year =

    Bernardelle, Pietro and others , title =. arXiv preprint arXiv:2508.16013 , year =

  8. [8]

    and others , title =

    Bernardelle, P. and others , title =. Companion Proceedings of the ACM Web Conference 2025 (WWW '25 Companion) , year =

  9. [9]

    International Conference on Learning Representations (ICLR) , year =

    Sharma, Mrinank and others , title =. International Conference on Learning Representations (ICLR) , year =

  10. [10]

    Findings of the Association for Computational Linguistics (ACL Findings) , year =

    Perez, Ethan and others , title =. Findings of the Association for Computational Linguistics (ACL Findings) , year =

  11. [11]

    International Conference on Machine Learning (ICML) , year =

    Sorensen, Taylor and others , title =. International Conference on Machine Learning (ICML) , year =

  12. [12]

    Proceedings of NAACL , pages =

    R. Proceedings of NAACL , pages =

  13. [13]

    2025 , editor =

    Cui, Justin and Chiang, Wei-Lin and Stoica, Ion and Hsieh, Cho-Jui , booktitle =. 2025 , editor =

  14. [14]

    Proceedings of CHI , year =

    Jakesch, Maurice and others , title =. Proceedings of CHI , year =

  15. [15]

    Science , volume=

    The levers of political persuasion with conversational artificial intelligence , author=. Science , volume=. 2025 , publisher=

  16. [16]

    arXiv preprint arXiv:2508.21448 , year =

    Kabir, Shariar , title =. arXiv preprint arXiv:2508.21448 , year =

  17. [17]

    and others , title =

    Sakhawat, A. and others , title =. 2026 , journal =

  18. [18]

    and others , title =

    Aldahoul, N. and others , title =. arXiv preprint arXiv:2505.04171 , year =

  19. [19]

    Proceedings of the International Conference on Machine Learning (ICML) , year =

    Santurkar, Shibani and others , title =. Proceedings of the International Conference on Machine Learning (ICML) , year =

  20. [20]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pages =

    Li, Junyi and Peris, Charith and Mehrabi, Ninareh and Goyal, Palash and Chang, Kai-Wei and Galstyan, Aram and Zemel, Richard and Gupta, Rahul , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pages =. 2024 , note =

  21. [21]

    Proceedings of NAACL , year =

    Miehling, Erik and others , title =. Proceedings of NAACL , year =

  22. [22]

    arXiv preprint arXiv:2505.23816 , year =

    Chang, Trenton and Schnabel, Tobias and Swaminathan, Adith and Wiens, Jenna , title =. arXiv preprint arXiv:2505.23816 , year =

  23. [23]

    International Conference on Learning Representations (ICLR) , year =

    Kim, Junsol and Evans, James and Schein, Aaron , title =. International Conference on Learning Representations (ICLR) , year =

  24. [24]

    Defining and Evaluating Political Bias in LLMs , year =

  25. [25]

    Measuring political bias in Claude , year =

  26. [26]

    Nature Machine Intelligence , volume =

    Kirk, Hannah Rose and others , title =. Nature Machine Intelligence , volume =

  27. [27]

    Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

    Llms are biased teachers: Evaluating llm bias in personalized education , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

  28. [28]

    Generative

    Bastani, Hamsa and Bastani, Osbert and Sungu, Alp and Ge, Haosen and Kabak. Generative. Proceedings of the National Academy of Sciences , volume =. 2025 , doi =

  29. [29]

    Proceedings of the National Academy of Sciences , volume =

    Hackenburg, Kobi and Margetts, Helen , title =. Proceedings of the National Academy of Sciences , volume =

  30. [30]

    Proceedings of EMNLP , pages =

    Potter, Yujin and others , title =. Proceedings of EMNLP , pages =

  31. [31]

    arXiv preprint arXiv:2602.06371 , year =

    Ko, Ju-Chun , title =. arXiv preprint arXiv:2602.06371 , year =

  32. [32]

    Temperature Setting in

  33. [33]

    2024 , eprint=

    Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain , author=. 2024 , eprint=

  34. [34]

    2024 , eprint=

    Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization , author=. 2024 , eprint=

  35. [35]

    2025 , eprint=

    Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale , author=. 2025 , eprint=

  36. [36]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=

    Localizing Persona Representations in LLMs , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , month=. doi:10.1609/aies.v8i1.36577 , abstractNote=

  37. [37]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=

    GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , month=. doi:10.1609/aies.v8i1.36552 , abstractNote=

  38. [38]

    Nature , volume=

    Optimizing generative AI by backpropagating language model feedback , author=. Nature , volume=. 2025 , doi=

  39. [39]

    and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , booktitle=

    Khattab, Omar and Singhvi, Arnav and Maheshwari, Paridhi and Zhang, Zhiyuan and Santhanam, Keshav and Vardhamanan, Sri and Haq, Saiful and Sharma, Ashutosh and Joshi, Thomas T. and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , booktitle=

  40. [40]

    arXiv preprint arXiv:2411.00027 , year=

    Personalization of Large Language Models: A Survey , author=. arXiv preprint arXiv:2411.00027 , year=

  41. [41]

    Auditing Alignment Controllability in

    Bu\'. Auditing Alignment Controllability in. 2026 , publisher =. doi:10.5281/zenodo.21489805 , url =

  42. [42]

    , journal=

    Brcic, Mario and Yampolskiy, Roman V. , journal=. Impossibility Results in. 2023 , publisher=

  43. [43]

    The Effortless Trap: Productive Struggle,

    Brcic, Mario and Frljic, Stjepan , journal=. The Effortless Trap: Productive Struggle,. 2026 , doi=

  44. [44]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=

    Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , pages=

  45. [45]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=

    Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression , volume=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author=. 2025 , pages=