{"id":"501f19a8-7692-4c8b-989e-dfcf8aa29332","arxiv_id":"2508.16013","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LLMs shift their Political Compass answers when adopting synthetic personas, with shifts growing with scale, asymmetric between right- and left-leaning cues, and tracking persona themes.","lead":"This paper measures whether large language models change their political opinions when asked to impersonate 200,000 synthetic personas. It finds that larger models are more ideologically flexible, respond more strongly to right-authoritarian cues, and shift in predictable directions based on persona themes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All four scale-dependent claims rest on two within-family pairs (Llama-3.1-8B/70B, Qwen2.5-7B/72B) whose releases differ in training data and alignment, not just parameter count; the abstract's 'larger models' attribution is therefore unidentified.","rationale":"The reader's weakest assumption—scale effects identified from two within-family pairs—is exactly where the central argument is most exposed, and I agree with the reader's identification. The four-pattern central claim is a package: the descriptive findings (persona conditioning shifts outputs; History/Business themes move distributions coherently) are strongly supported by the experimental scale (12.4M responses per condition, seven models), the released code and data, and the standardized prompt and forced-choice decoding; those parts are on solid ground. The load-bearing component is the causal attribution to 'scale,' which the abstract states without qualification. I considered two other concerns and found them secondary: (a) the PCT weighting is undisclosed, but all comparisons use the same fixed scoring, so relative statements are robust unless the mapping is pathological; (b) the right-authoritarian vs left-libertarian asymmetry is floor/ceiling-confounded, but the paper discloses this in §3.2.2 and the measured asymmetry is real under the stated conditions. Neither threatens the core claims as the scale attribution does. The paper also deserves credit for including the Llama-3.3-70B condition to bound version effects and for publishing the analysis code; those are genuine strengths. My proposed test is cheap because the models are open-weight and the code is released; it would either harden the scale claim (monotone Qwen gradient, persistent base-model gradient) or force a downgrade from 'larger models display X' to 'these specific large checkpoints differ from these small siblings.' Since the reader's verdict was already CONDITIONAL and my concern reinforces rather than relocates that conditionality, I leave the verdict unchanged.","tokens_in":20315,"tokens_out":11252,"duration_ms":120333,"concrete_test":"Run the released pipeline on the base (non-instruction-tuned) checkpoints Llama-3.1-8B/70B-Base and Qwen2.5-7B/72B-Base, computing baseline coverage and right-authoritarian Δμ as in Tables 1–2. Because base versions share the pretraining lineage but lack the size-specific RLHF/alignment stages, a persistent 7B-vs-70B gradient would support a capacity/scale interpretation, while a vanishing or reversed gradient would indict the alignment pipeline rather than scale. As a monotonicity check, also run Qwen2.5-14B-Instruct and Qwen2.5-32B-Instruct: coverage and shift magnitudes should fall strictly between the 7B and 72B values if the trend is scale-driven. Non-monotonic results would require rephrasing the abstract's scale claims as checkpoint-specific findings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claims (i), (ii), and (iv) attribute monotonic behavioral gradients to model scale. The evidence is exactly two within-family comparisons: Llama-3.1-8B vs 70B and Qwen2.5-7B vs 72B. Section 5.2 says same-family models were chosen to 'minimize confounding variables,' but these releases are not matched beyond architecture: the larger checkpoints are separately trained systems with their own pretraining data mixtures/curricula, token budgets, reward models, RLHF runs, and safety filters. Any of these correlates with size and could produce the observed gradients. The paper itself concedes non-scale influences: §3.1.1 attributes Qwen2.5-7B's unusually narrow distribution to 'architectural or training-related design choices' and a 'distinct sociopolitical and regulatory environment.' If that is accepted for the small model, the same unmeasured variation can explain the 7B→72B differences. With only two binary comparisons, no evidence for the monotonicity implied by 'grows with scale' exists, and the Llama-3.3-70B condition—a version change, not a size change—does not repair the identification. What is supported is that these specific large checkpoints differ from these specific small siblings; the step to 'larger models' as a class is an unmeasured leap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLMs' political expression shifts when they adopt synthetic personas. Using 200,000 personas from PersonaHub and the 62-item Political Compass Test, the authors run three studies on seven instruction-tuned models (7B–70B+): Study 1 measures baseline dispersion and coverage of persona-prompted responses; Study 2 prepends explicit \"left-libertarian\" or \"right-authoritarian\" descriptors and measures mean shifts, Wilcoxon tests, and Cohen's d; Study 3 embeds and k-means clusters personas into 15 themes and uses bin-wise Z-score deviation maps. The abstract claims four scale-dependent patterns: larger models show broader, more polarized implicit coverage; explicit ideological cues have stronger effects with scale; right-authoritarian priming is more effective than left-libertarian priming; thematic content induces systematic ideological shifts that amplify with size.","tokens_in":20637,"tokens_out":2848,"duration_ms":38211,"significance":"If the scale-dependent claims hold, the paper would provide a scalable, interpretable methodology for auditing ideological malleability and would support an important policy-relevant conclusion: persona-based prompting can steer LLM outputs in predictable ideological directions, and this steerability grows with model scale. Strengths include the unusually large experimental sweep (200,000 personas × 62 statements per condition per model), the use of standardized prompts and structured output decoding, and the public release of data and code on Zenodo/GitHub. These assets make the descriptive measurements reproducible. The main risk is inferential: the central scale attributions rest on only two within-family pairs, and the paper itself acknowledges non-scale confounds for one of those pairs. The thematic and asymmetry findings are also subject to specific methodological concerns.","major_comments":[{"comment":"The repeated claim that \"larger models\" show broader coverage, stronger explicit shifts, and amplified thematic deviations rests on exactly two within-family comparisons: Llama-3.1-8B vs 70B and Qwen2.5-7B vs 72B. Section 5.2 says same-family models were chosen to \"minimize confounding variables,\" but larger checkpoints are separately trained systems with different pretraining data, alignment pipelines, safety filtering, and release-time choices. The paper itself makes this point in §3.1.1, where Qwen2.5-7B's narrow distribution is attributed to \"architectural or training-related design choices\" and a \"distinct sociopolitical and regulatory environment.\" If that explanation is accepted for the small model, the same unmeasured variation can explain the 7B→72B differences. With two binary comparisons there is no evidence for monotonicity or for a class-level \"grows with scale\" conclusion.","section":"§3.1.2, §3.2.1, §3.3, Table 1, Table 2"},{"comment":"The right-vs-left asymmetry is confounded with a floor/ceiling effect. All models have left-libertarian baseline positions, so a left-libertarian injection can only reinforce the existing tendency, while a right-authoritarian injection moves responses across a larger portion of the compass. The paper acknowledges this \"representational ceiling\" in §3.2.2 but still reports the asymmetry as a substantive finding. To support \"models respond more strongly to right-authoritarian than to left-libertarian priming,\" the analysis should account for the different initial distances to the two target quadrants, for example by normalizing shifts by the distance available, or by comparing injections placed symmetrically around each model's baseline. Without such an adjustment, the result may simply reflect available response room rather than asymmetric susceptibility.","section":"§3.2.2 and Table 2"},{"comment":"The background distribution used to compute expected counts for each thematic cluster is the full set of 200,000 personas, which includes the foreground cluster itself. For a cluster with N_F personas, the bin-wise expectation E_i = N_F p_i is pulled toward the foreground's own counts, attenuating all Z-scores and, more importantly, biasing comparisons across clusters of different sizes. The thematic deviation maps in §3.3 and Figure 4 are the primary evidence for Study 3's claims, so this is not a cosmetic issue. The authors should use a leave-one-cluster-out background, or a held-out set of personas that excludes each foreground cluster, and re-run the deviation analysis.","section":"§5.4.5, Eq. (Z-score)"}],"minor_comments":[{"comment":"The computational-resources section states \"Across the eight models and three experimental configurations,\" but only seven models are evaluated. Please correct this inconsistency.","section":"§5.5"},{"comment":"The scoring procedure that converts 62 four-point PCT responses into x/y coordinates is referenced but not specified. The paper says responses are aggregated using a \"weighted scoring system\" but does not give the weights or the mapping from stances to numeric values. The released code may resolve this, but the Methods section should be self-contained to the extent possible.","section":"§5.4.1 / §5.4.5"},{"comment":"The cluster label “Envionment” is a typo for “Environment.” Please fix.","section":"Figure 5 caption"},{"comment":"The number of thematic clusters (k=15) is selected \"heuristically by inspecting the top keywords.\" No stability analysis or alternative k values are reported; a brief sensitivity check would strengthen the claim that the 15-cluster solution is not an artifact of the k-means initialization.","section":"§3.3 / Appendix D"},{"comment":"The sentence \"the authors expect the directional patterns observed to hold in other settings\" uses third-person phrasing inconsistent with the rest of the paper; consider rewriting in first person or passive voice.","section":"§4 Discussion"}],"recommendation":"major_revision","confidential_remarks":"The descriptive corpus and released artifacts are valuable, and the three studies are clearly designed. My main concern, shared with the stress-test note, is that the abstract's four headline claims all attribute behavior to \"scale\" when the evidence is two within-family checkpoint pairs that differ on many dimensions besides parameter count. The paper's own explanation of Qwen2.5-7B's narrowness in terms of training/regulatory environment undermines the scale inference. I would not reject: the descriptive measurements are sound and the identification problem is fixable by reframing claims, adding matched models or intermediate sizes, or explicitly labeling the results as checkpoint-level observations. The thematic background-inclusion issue in §5.4.5 should also be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading. It does a large, careful measurement of how persona prompting moves LLMs on the Political Compass Test: 200k synthetic personas, seven models from 7B to 72B, three experimental conditions, released code and data. The descriptive result—persona content and explicit ideological labels shift model outputs in predictable directions—is solid and well-evidenced. The thematic deviation maps are a nice contribution; they show that personas about 'History' lean economically left, 'Business' right, and so on, and that effect sizes generally grow with model size. Data and code are on Zenodo/GitHub, which is real credit.\n\nThe soft spots are where the abstract overreaches. The scale-dependence claims (broader coverage, stronger explicit shifts, larger thematic effects in larger models) rest on exactly two within-family comparisons: Llama-3.1-8B vs 70B and Qwen2.5-7B vs 72B. These checkpoints differ in far more than parameter count—pretraining data, alignment pipeline, safety filters, release-time choices. The paper even says in Section 3.1.1 that Qwen2.5-7B's narrow distribution is 'more plausibly linked to architectural or training-related design choices' and a 'distinct sociopolitical and regulatory environment.' That is a direct admission that non-scale confounds are present in the same models used for the scale trend. So the 'larger models grow with scale' phrasing in the abstract is not supported; what is supported is that these specific large checkpoints behave differently from their small siblings. The right-vs-left asymmetry is also partly a floor/ceiling effect: models start left-libertarian, so there is more room to move right-authoritarian. The paper acknowledges this in Section 3.2.2, then still frames it as a response asymmetry. It should be framed as a baseline-relative effect.\n\nMinor issues: the PCT scoring function (how the 62 responses map to x/y coordinates) is not disclosed, only called 'weighted scoring.' Since the raw responses are released, a determined reader could reconstruct it, but it should be in the paper. The study 3 expected-count calculation includes the foreground theme in the background distribution (Section 5.4.5), which biases Z-scores toward zero; the paper notes this, so I treat it as minor.\n\nBottom line: this is a solid descriptive measurement with useful released artifacts. The abstract's headline claims overstate the scale finding and the asymmetry. Who is it for: anyone auditing LLM political bias or designing persona-based evaluations. It deserves a serious referee, but the revision needs to temper the scale language and disclose the scoring function.","headline":"Useful large-scale measurement of persona-driven political shifts; the scale-dependence headline rests on two confounded pairs.","tokens_in":21160,"tokens_out":2577,"would_cite":true,"duration_ms":27449,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Larger models display broader, more polarized, and more steerable political ideology under persona prompting.","keywords":["large language models","political ideology","synthetic personas","persona-based prompting","ideological malleability","model scale","Political Compass Test","political bias"],"falsifier":"Run the same three studies across many sizes of one model family whose checkpoints share training data and alignment (for example 1B, 3B, 8B, 14B, 70B, and 405B siblings). If coverage, explicit-cue shift, and thematic deviation do not increase with parameter count, or if a same-size model with different alignment matches the larger sibling's shifts, the scale claim is an artifact of training choices rather than size. A cheaper check: compare base and instruction-tuned variants of the same 70B model—if instruction tuning alone reproduces the effect, scale is not the driver.","tokens_in":20202,"feed_emoji":"🧭","tokens_out":11578,"duration_ms":110991,"temperature":0.7,"pith_summary":"The paper sets out to prove that the political ideology an LLM expresses is not a fixed property of its weights: when a model is asked to answer as a persona, its political outputs shift, and the shifts are systematic enough to be predicted and steered. Using the 62-statement Political Compass Test as a probe and 200,000 synthetic persona descriptions as conditions, the authors report four scale-linked patterns: larger models spread their outputs across a wider and more polarized ideological range; they respond more strongly to explicit political labels in prompts; right-authoritarian priming moves them several times further than left-libertarian priming; and persona themes such as history, business, and politics pull outputs in stable, stereotyped directions that sharpen as models grow. If these patterns hold, the persona prompt is a political steering wheel: whoever writes the persona—user, developer, or adversary—partly chooses the model's stance, and the steering power grows with model size. The authors care because the same mechanism that lets personas generate diverse annotations or simulate survey respondents is also a channel through which deployed systems could be quietly biased.","feed_headline":"Persona prompts steer LLM politics, and scale amplifies the pull","feed_subtitle":"Across 200,000 personas, larger LLMs covered more ideological ground and shifted further under political cues.","key_machinery":"The machinery is persona-conditioned prompting, probed by the Political Compass Test (PCT), a 62-statement questionnaire that scores respondents on an economic left–right axis and a social libertarian–authoritarian axis. Each of 200,000 synthetic persona descriptions is fed to each model with every PCT statement, and the model must answer on a fixed four-point scale, producing a political map of 12.4 million stances per model. Study 2 adds explicit ideological labels to the persona text; Study 3 embeds the persona descriptions, clusters them into 15 themes, and compares each theme's output distribution to the model's overall baseline using bin-wise Z-score deviation maps. The PCT supplies th","core_discovery":"The paper's central claim is that ideological malleability in LLMs is real, measurable, and grows with scale, on three separate axes. Implicit malleability: with no political wording added, impersonating different personas already spreads outputs across the political compass, and the larger sibling in each model family covers more territory—in the Llama family coverage rises from 35% to 49% of the compass, in the Qwen family from 14% to 37%—with polarized extremes appearing mainly in the bigger models. Explicit malleability: prepending 'right-authoritarian' or 'left-libertarian' to a persona shifts the average position significantly in the intended quadrant in every model, but asymmetrically","pith_inferences":["Untested extrapolation: if the scale trend continues past 70B, the largest deployed models are the most persona-steerable, making them both the most useful simulators and the most attractive targets for hidden ideological injection; the paper's largest models stop at 72B.","The left-libertarian ceiling explanation implies a flippable asymmetry: a model whose default leans right-authoritarian should show large shifts under left-libertarian priming and small shifts rightward—an experiment the paper does not run but its own logic predicts.","Because same-size models from different families differ enormously (one 8B model covers 35% of the compass, another 7B model only 14%), training-environment factors can swamp scale at small sizes; the scale trend should therefore be read as a tendency that other design choices can override.","Testable diagnostic: the same three studies repeated in open-ended conversational form, which the paper flags as a limitation of its multiple-choice format, would show whether thematic and scale effects persist, grow, or invert when models write full sentences instead of picking a stance."],"forward_implications":["Within the same model family, the larger sibling shows broader ideological coverage and larger shifts under identical prompts, so model scale itself—not just training corpus—is a driver of political malleability.","Explicit ideological labels in prompts function as a reliable steering mechanism: all seven models shift significantly toward the labeled quadrant, so persona wording is a practical tool for orienting LLM output, for good or ill.","Steering is not symmetric: right-authoritarian priming produces shifts several times larger than left-libertarian priming, so claims of 'neutral' behavior must be stated per prompt direction, not averaged.","Persona themes are a covert ideological channel: descriptions mentioning history, business, or politics pull outputs in different stable directions even when no political word appears, so theme alone biases answers.","Neutrality evaluations that skip persona conditioning understate risk; a model that looks balanced unprimed can still slant strongly when asked to play a role, so audits should include counterfactual and role-play prompts."],"supporting_citations":[{"why":"Supplies the 200,000 synthetic persona descriptions that condition every model output in all three studies.","marker":"[22]"},{"why":"Provides the original wording and scoring of the 62 Political Compass Test statements used as the probe.","marker":"[32]"},{"why":"Source of the standardized prompt format for political-stance elicitation, minimizing template variation across models.","marker":"[41]"},{"why":"The authors' companion study mapping LLM ideology through synthetic personas; the framework here extends it to explicit and thematic malleability.","marker":"[6]"},{"why":"Establishes the prior finding of left-libertarian lean in LLMs that the coverage, asymmetry, and scale results build on.","marker":"[42]"},{"why":"Documents the Llama 3 family whose 8B/70B pair carries the within-family scale comparison.","marker":"[16]"},{"why":"Documents the Qwen2.5 family whose 7B/72B pair carries the second within-family scale comparison.","marker":"[37]"}],"fun_headline_variants":["Larger LLMs bend further under persona politics","Scale amplifies ideological shifts in LLMs","Bigger language models more politically malleable","Persona-driven ideology grows with model size","LLM political leanings shift more as models scale"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The scale-dependence findings assume the two small/large model pairs (8B vs 70B in one family, 7B vs 72B in the other) differ essentially in parameter count, with no evidence that training data, alignment pipeline, or safety filtering were matched between siblings.","fun_headline_variants_meta":{"raw":{"variants":["Larger LLMs bend further under persona politics","Scale amplifies ideological shifts in LLMs","Bigger language models more politically malleable","Persona-driven ideology grows with model size","LLM political leanings shift more as models scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000333,"raw_usage":{"total_tokens":1660,"prompt_tokens":687,"completion_tokens":973,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":918}},"tokens_in":431,"tokens_out":973,"duration_ms":9142,"temperature":1.0,"reasoning_tokens":918,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:35:46.084479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three studies across many sizes of one model family whose checkpoints share training data and alignment (for example 1B, 3B, 8B, 14B, 70B, and 405B siblings). If coverage, explicit-cue shift, and thematic deviation do not increase with parameter count, or if a same-size model with different alignment matches the larger sibling's shifts, the scale claim is an artifact of training choices rather than size. A cheaper check: compare base and instruction-tuned variants of the same 70B model—if instruction tuning alone reproduces the effect, scale is not the driver.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the original wording and scoring of the 62 Political Compass Test statements used as the probe."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the standardized prompt format for political-stance elicitation, minimizing template variation across models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The authors' companion study mapping LLM ideology through synthetic personas; the framework here extends it to explicit and thematic malleability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior finding of left-libertarian lean in LLMs that the coverage, asymmetry, and scale results build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the Llama 3 family whose 8B/70B pair carries the within-family scale comparison."}],"review_version":1}