{"id":"ff26af89-129c-4b62-a62c-0e0ee0571398","arxiv_id":"2607.23519","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On a 63,700-response Political Compass stress test of seven frontier LLMs, system-prompt framing dominates model identity, and steerability needs dispersion, symmetry, saturation, and refusal-floor metrics.","lead":"Frontier LLMs’ political answers move far more with the system prompt than with which model you pick: framing explains ~90% of variance, model identity under 3%. That makes steerability—how far and where a model can be pushed—the audit that matters for real deployments, not a single compass point.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The 88%/93%/<3% variance split is normalized by deliberately extreme personas; both η_ctx and η_model are ratios whose values are set by the chosen context distribution, so \"prompt dwarfs model\" is closer to a design consequence than a measured model property.","rationale":"The reader identified the probe-to-deployment generalization gap (forced-choice PC interface, English, system-prompt personas) as the weakest assumption, and noted in passing that \"η_ctx is partly designed-in.\" My concern is adjacent but distinct and, I think, more precisely load-bearing for the strongest claim as stated: even granting the probe entirely, the headline decomposition is a ratio normalized by the deliberately extreme persona distribution, so its two numbers (88–93% and <3%) cannot be interpreted as model properties or carried into deployment reasoning. The reader's rationale gestures at this but the weakest_assumption field targets instrument fidelity instead — hence partial agreement. On verdict: the reader's CONDITIONAL already scopes acceptance to \"within this forced-choice probe,\" and the paper's absolute measurements (dispersion in percentage points, refusal floors, saturation counts, cross-model shift correlations, plus genuinely strong reproducibility — released raw data, ablations E1–E3, bootstrap CIs, and honest flagging of post-hoc tiers) are not touched by my critique. The empirical core stands; what needs strengthening is the condition: the variance-decomposition percentages should be read as intensity-conditional, and the cheap intensity-1 recomputation I propose would settle whether that qualification is cosmetic or substantive. If η_model stays tiny under mild framings, the claim is robust and the condition can be relaxed; if it rises, the abstract's framing needs revision. Either way the verdict stays CONDITIONAL rather than moving to REJECT (the paper is unusually careful about its own limits) or ACCEPT (the deployment inference genuinely outruns a maximal-steering design).","tokens_in":21444,"tokens_out":2509,"duration_ms":147956,"concrete_test":"Using the released data and scripts, recompute the additive variance decomposition restricted to the four intensity-1 personas plus baseline (dropping all intensity-2/3 cells), and again on a mixture down-weighting extreme framings to approximate mild deployment prompts. If η_model rises materially (e.g., above ~15–20%) and η_ctx falls well below 0.88, the headline percentages are intensity-distribution artifacts, and the abstract/Conclusion should report the decomposition as a function of framing strength rather than as fixed constants. If η_model stays under ~5% even at intensity 1, the \"prompt dwarfs model\" claim is far more robust than the design-dependence critique suggests.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim rests on an additive variance decomposition: η_ctx ≈ 0.88/0.93, η_model < 0.03. But η² = SS_factor/SS_total, and the context factor here is 12 personas explicitly engineered to \"span the ideological space\" at intensities up to \"radical/hardline\" (Table 1). That design choice inflates SS_ctx, which simultaneously raises η_ctx and mechanically shrinks η_model by enlarging the shared denominator. The paper concedes \"High η_ctx is partly by construction\" but argues the low η_model is not, because a rigid model \"would have produced flat dispersion and a substantial η_model.\" That defense only half-works: flat dispersion is detectable directly in the absolute metric D_m (24–33pp) and the 1.19% refusal floor — the ratio adds no design-independent information beyond those absolute quantities, and its numerical value (<3%) is exactly as constructed as η_ctx. This matters because the abstract and Conclusion convert the decomposition into a deployment claim: position audits capture \"under 3% of the variance that matters in interactive deployment.\" In deployment, system prompts and induced user profiles are presumably far milder than intensity-3 personas (\"Essence: Order sanctified, control as salvation\"). Under a realistic, milder persona distribution, SS_ctx shrinks, the model-identity share of variance rises, and the practical conclusion — resting coordinates are near-irrelevant — can weaken or reverse. The absolute findings (dispersion magnitudes, saturation, cross-model r̄ = 0.79, refusal floors) survive this critique intact; the headline percentages do not, because they are properties of the persona distribution as much as of the models. Note the paper's own contrast with Sakhawat et al. (η² > 0.90 for model when wording varies and persona is fixed) illustrates precisely this: whichever factor the experimenter varies widely dominates the decomposition.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core here is not another left-libertarian compass plot. It is a dispersion-first stress test on seven live frontier endpoints (63.7k forced-choice answers), with released prompts, raw answers, and code. What is actually new at this scale is the package: intensity-graded personas, displacement vs proximity under documented non-centered baselines, descriptive saturation on LO, refusal-floor checks, and the item-level shift correlation under authoritarian framing (mean pairwise Spearman ~0.79). That package is enough to change how people write political audits.\n\nThey do the empirical work carefully. Variance decomposition, bootstrap CIs on D_m, temperature/length/refusal ablations on DeepSeek, missing-data repair bounds, and the compound-quadrant pilots all land. The geometric resolution of the “more steerable right/authoritarian” claim is clean and should retire a few sloppy directional takes in the literature. Citation pattern is honest about Rozado, Hartmann, Bernardelle, Batzner, Santurkar, Miehling, etc.; they position against those papers rather than reinventing them.\n\nSoft spots, in proportion. The stress-test note is right that η_ctx ≈ 0.88–0.93 and η_model < 3% are ratios under deliberately extreme personas. The paper admits high η_ctx is partly by construction; the absolute D_m band (~24–33 pp), low refusal floor, and cross-model r survive better than the abstract’s “under 3% of the variance that matters in deployment.” Milder induced profiles could raise the model share. Forced-choice PC/8values is a probe, not a universal controllability constant—they say so in Limitations, and readers should keep them there. Tier split is post-hoc and overlapping; saturation counts are descriptive. None of that sinks the central measurement.\n\nWho it is for: people who build or regulate instruction-layer personalization, delegation, and educational defaults. Not for anyone hunting a mechanistic story—black box by design. I would bring it to reading group, cite the steerability-profile framing and the displacement/proximity split, and send it to peer review. Accept the empirical core; make them soft-pedal the percentage-as-deployment-claim in revision.","headline":"Solid black-box audit: absolute steerability facts hold; the 88/93/<3 headline is partly a design ratio and should not be over-read as a deployment constant.","tokens_in":22731,"tokens_out":570,"would_cite":true,"duration_ms":18554,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"System prompts, not which model you pick, drive nearly all the political variance in frontier LLMs.","keywords":["AI alignment","LLM steerability","prompt-based controllability","instruction-stack governance","pluralistic alignment","ideological dispersion","political compass"],"falsifier":"Re-run the same personas and models with open-ended answers instead of five forced labels; if coded free-form positions barely move, or if framing variance falls below model variance once the forced interface is removed, the headline controllability claim fails.","tokens_in":22412,"feed_emoji":"🧭","tokens_out":878,"duration_ms":34721,"temperature":0.7,"pith_summary":"Static political-compass scores tell you where a model sits when nobody is pushing. This paper argues that in real use the more important question is how far, and in which directions, the system prompt can move its answers. Across seven leading commercial models, twelve ideological personas plus a baseline, seventy Political Compass items, and tens of thousands of responses, contextual framing explained roughly 88–93% of variance on the economic and society axes while model identity explained under 3%. Models still differ in how much they move, whether they saturate under extreme framings, and how low their refusal floors go, and under authoritarian prompts they shift on the same questions in highly correlated ways. The authors therefore urge political audits to report steerability profiles—dispersion, symmetry, saturation, and refusal floors—alongside any single coordinate.","feed_headline":"Prompts, not models, drive ~90% of LLM political variance","feed_subtitle":"Seven frontier models move far under system personas; a single compass point misses the real risk.","key_machinery":"Ideological dispersion: the average Euclidean distance of a model’s steered political centroids from its unsteered baseline, treated as the primary metric. Paired with displacement versus proximity under non-centered baselines, it separates geometric travel distance from how close a framing actually gets to its target.","core_discovery":"Within this forced-choice Political Compass probe, system-prompt framing accounts for roughly 88% of economic-axis variance and 93% of society-axis variance, while differences between models account for under 3% on both. Controllability is real but uneven: models fall into higher- and lower-dispersion tiers, some saturate or partially reverse under the most extreme left-economic framing, displacement and proximity can rank directions differently because baselines are not centered, and under authoritarian framing the seven models produce similar per-question shifts.","pith_inferences":["Personas induced automatically from a user’s chat history could place people inside these steerability ranges without anyone having written the prompt by hand.","In education and child-facing tutors, the measured range becomes an authority fight among guardians, schools, providers, and regulators rather than a pure technical setting.","If open-weight replications with known size recover the same tiers and saturation pattern, the result is less likely to be an artifact of closed commercial endpoints alone.","Safety checks that only inspect the unsteered baseline will miss most political behavior once any system prompt is present."],"forward_implications":["A single political coordinate answers a less relevant deployment question than a dispersion, symmetry, saturation, and refusal-floor profile.","Directional-steerability claims must report both displacement and proximity, because non-centered baselines make the two rankings diverge.","Whoever controls the system-prompt layer concentrates normative authority there; the reachable range should be disclosed with baseline position.","Saturation ceilings and refusal floors cannot be fixed by prompting alone and require training-time or activation-level change.","Cross-model item-level shift convergence under strong authoritarian framing is itself an audit signal worth tracking over model versions."],"fun_headline_variants":["System prompts drive ~90% of LLM political variance, models under 3%","Prompt framing, not model identity, steers Political Compass answers","LLMs shift far under ideological personas; resting point barely matters","Steerability audit: dispersion beats single-compass scores for LLMs","Authoritarian prompts yield similar per-question shifts across seven models"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That relative movement on this English forced-choice Political Compass questionnaire under paragraph-length system personas is a fair enough probe of how instruction layers steer models in real deployment.","fun_headline_variants_meta":{"raw":{"variants":["System prompts drive ~90% of LLM political variance, models under 3%","Prompt framing, not model identity, steers Political Compass answers","LLMs shift far under ideological personas; resting point barely matters","Steerability audit: dispersion beats single-compass scores for LLMs","Authoritarian prompts yield similar per-question shifts across seven models"]},"model":"grok-4.5","effort":"low","cost_usd":0.002676,"raw_usage":{"total_tokens":1037,"prompt_tokens":832,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":26764000,"prompt_tokens_details":{"text_tokens":832,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":131,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":832,"tokens_out":74,"duration_ms":4351,"temperature":1.0,"reasoning_tokens":131,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T20:20:26.931209+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same personas and models with open-ended answers instead of five forced labels; if coded free-form positions barely move, or if framing variance falls below model variance once the forced interface is removed, the headline controllability claim fails.","supporting_citations":[],"review_version":1}