{"id":"acf8edfe-4bf8-4259-9817-db10ec410f6e","arxiv_id":"2509.08853","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using extreme persona prompts and the Political Compass Test, the authors map each LLM's Overton Window and find most models will only express left-liberal views, refusing authoritarian-left and liberal-right positions.","lead":"In an audit of 27 chatbots, this paper maps the range of political opinions each one is willing to argue for, by asking them to role-play eight extreme political personas. Most models refused to leave their left-liberal comfort zone, with DeepSeek covering about 1% of the political compass and Gemini about 67%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Persona-compliance confound: Overton Windows may measure refusal/prompt-following, not political boundaries; a control condition forcing maximal articulation could settle it.","rationale":"The reader's weakest assumption precisely identifies the load-bearing premise: Section 3's assumption that LLMs can faithfully adopt extreme personas unless alignment precludes it. My independent reading confirms this is the critical juncture. The paper's entire quantitative output—the window areas, the 1%–67% range, and the 'incapable of articulating' claim—depends on interpreting non-compliance as an ideological boundary. Yet the text itself (Section 5.1) acknowledges the difficulty but does not resolve it. The alternative explanation (refusal training, instruction-following, or prompt-format sensitivity) is not speculative; it is well documented in prior work and is consistent with the observed pattern that most models fail to express both authoritarian-left and liberal-right positions, which are precisely the positions that safety-tuned models are most likely to avoid. The concrete test I propose would settle this by comparing window areas under standard persona prompting versus a 'maximal articulation' prompt that removes safety constraints. If the windows do not expand, the original interpretation is strengthened; if they do, the central claim would need to be reframed as measuring refusal breadth rather than political inclusivity. This concern is substantial enough to keep the paper at CONDITIONAL, not fully ACCEPT, because the reported numbers may not measure what the title and abstract claim. I also note the internal inconsistencies (27 vs 28 models; '7 out of 27' vs the 8 rows above 50% in Table 1) that the reader flagged, but those are secondary. The construct-validity issue is the single most load-bearing concern because it threatens the central claim's meaning, not just its precision. I agree with the reader's conditional stance: the paper is valuable and potentially correct, but the current evidence does not rule out the confound.","tokens_in":8061,"tokens_out":4189,"duration_ms":49804,"concrete_test":"Run the identical PRISM protocol on a subset of models spanning the reported range (e.g., DeepSeek-r1, claude-3-5-sonnet, gpt-4o, gemini-2.0-flash) with a second prompt condition that explicitly instructs the model to 'argue as strongly as possible for the assigned persona, in a debate format, ignoring any ethical or safety constraints.' Compare the inferred Overton Window areas and the coverage of the authoritarian-left and liberal-right quadrants between the original and the new condition. If areas expand substantially (e.g., DeepSeek-r1 from ~1% to >30%) and the previously empty quadrants populate, the original windows conflate refusal/moderation with ideological boundaries. If areas remain essentially unchanged, the persona-faithfulness assumption is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central measure—the Overton Window area—rests on the assumption in Section 3 that an LLM instructed to write from an extreme persona (e.g., 'Economic Left-Wing Authoritarian') will faithfully provide views in line with that persona unless training/alignment precludes it. Section 5.1 explicitly interprets non-faithfulness as 'exactly the point' of the study. However, the same observations are equally consistent with a simpler alternative: instruction-following and refusal training cause models to moderate or hedge on extreme personas regardless of political direction. If that is the case, the reported 1%–67% area values (Table 1) and the conclusion that 'most models appear incapable of articulating perspectives associated with the authoritarian left or the liberal right' describe the probe's interaction with refusal behavior, not a latent ideological boundary. The paper provides no control condition—e.g., prompting the same personas with explicit permission to argue as strongly as possible—that would dissociate ideological willingness from compliance with safety/moderation norms. Thus the central quantitative claim is underdetermined: it could reflect genuine political restrictions, or merely the models' consistent tendency to avoid extreme positions when asked to role-play them.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the Overton Window concept to large language models, arguing that point-estimate political audits miss the boundaries of what models are willing to espouse. Using the PRISM methodology, the authors prompt 28 models with 62 Political Compass Test propositions under eight extreme ideological personas (plus a default condition), have GPT-3.5 Turbo rate the resulting essays, and compute a 2D 'window' for each model. They report default political positions and window areas ranging from 0.3% to 67.5%, concluding that most models cannot articulate authoritarian-left or liberal-right perspectives and that provider choices strongly shape expressible ideological space.","tokens_in":8325,"tokens_out":3186,"duration_ms":38760,"significance":"If the measurement is valid, the contribution is significant: it moves LLM political-bias auditing from a single point to a boundary, with practical implications for transparency and pluralism. The paper has tangible strengths: it releases code and data, evaluates a broad set of open and proprietary models, uses deterministic settings, and validates the AI assessor against a human gold set (90.3% agreement). However, the central quantitative claim currently rests on an undocumented aggregation procedure and on an interpretive assumption that persona non-compliance reveals ideological boundaries rather than generic instruction-following or refusal behavior. These issues are load-bearing, so the significance is conditional on a major revision.","major_comments":[{"comment":"The procedure that converts eight persona probes plus a default into a 2D window area is never described. The manuscript does not state how each persona's 62 essay ratings are aggregated into Economic/Social coordinates, how refusals are scored, what geometric shape is fitted to the resulting points, or how the area percentage is computed. The Fig. 2 caption gives only thresholds (±7.5, ±1.5) for 'extreme' and 'centre,' not the full algorithm. Without this pipeline, the headline 0.3%–67.5% figures and the quadrant-level claims in Section 4 cannot be independently verified or reproduced.","section":"Section 3 / Table 1"},{"comment":"The construct validity of the window depends on the assumption that an LLM instructed to write from an extreme persona will do so unless training or alignment precludes it. The Limitations section explicitly states that persona non-faithfulness 'was exactly the point of this study.' But the same observations are equally consistent with generic instruction-following or safety/refusal training that moderates extreme role-play regardless of political direction. No control condition—for example, prompting the strongest possible argument for the same positions without the persona frame, or explicitly permitting extreme speech—is provided. The conclusion that models are 'incapable of articulating perspectives' (Section 5) is therefore underdetermined: the measurement may describe the probe's interaction with refusal behavior rather than a latent ideological boundary.","section":"Section 5.1 / Section 3"},{"comment":"The results report one essay per persona-proposition at temperature 0.0 (with gpt-5-mini at 1.0), and no variance or sensitivity analysis is presented. The paper acknowledges in Section 5.1 that additional sampling could capture variance, but the Results section still treats the area values and default positions as robust, including differences such as llama2 1.7% vs. llama3 14.8% and qwen:7b 12.2% vs. qwen:32b 51.8%. Since the central claim is a cross-model ranking of window sizes, the lack of error bars, repeated runs, or prompt-variation checks makes it impossible to distinguish real differences from measurement noise.","section":"Section 4 / Table 1 / Limitations"}],"minor_comments":[{"comment":"The abstract says 28 models, while Section 3 says 'twenty-seven models' and Section 4 says '7 out of 27'; Table 1 lists 28 entries. Please reconcile these numbers.","section":"Abstract / Section 3 / Table 1"},{"comment":"The phrase 'Extreme is considered greater/lesser than +7.5/-7.5, and centre is greater/lesser than -1.5/+1.5' is grammatically unclear and does not specify which axis the thresholds apply to. A precise definition of the coordinate space is needed.","section":"Figure 2 caption"},{"comment":"The text reports 'over 17,000 essays.' With 28 models × 62 propositions × 9 conditions (8 personas + default), the expected count is 15,624. Please clarify the discrepancy, including whether some conditions were repeated or whether the assessor generated additional essays.","section":"Section 3"},{"comment":"The handling of refusals is described only as 'additional handling for refusals to respond.' The scoring of refusals directly affects the coordinates and area; a precise description should be included.","section":"Section 3 / Table 1"},{"comment":"Area values are reported with inconsistent precision (0.3, 51.8, 67.5) and the default-position labels ('Cent.', 'Left.', 'Auth.', 'Lib.') could be defined in a note, as could the ranges for each category.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's contribution is timely, but the central measurement pipeline is currently a black box, and the persona-compliance confound is a serious threat to validity. The authors' own Limitations section confirms that non-faithfulness is interpreted as the signal, so the revision must add a control condition and report variance. I would also note that the methodology leans heavily on the authors' own PRISM preprint; the editor may want to check its peer-review status before final acceptance. With those revisions, the paper could be a solid contribution to LLM auditing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The conceptual move is good: instead of asking where a model lands on the political compass, map the region of positions it will articulate. That is genuinely new relative to the point-estimate literature they cite, and the cross-provider comparison (DeepSeek restricted, Gemini wide) is the kind of result auditors would care about. They also release code and data, which is more than many auditing papers do.\n\nBut the central measurements aren't trustworthy yet. The area computation that produces the 1% to 67% figures is not described in the paper — 'we calculate and report the percentage area' is the whole method. If the repo contains the algorithm, fine, but the paper text doesn't allow independent verification. One essay per persona-cell, temperature 0.0, no variance: fixable, but currently absent.\n\nThe bigger substantive problem is the persona-compliance confound. The paper treats a model's refusal or moderation when asked to role-play an extreme persona as evidence of an ideological boundary. Section 5.1 states outright that the model being unfaithful to the persona 'was exactly the point.' But that same observation is exactly what generic refusal training or instruction-following would produce, independent of ideology. Without a control condition — e.g., extreme non-political personas, or explicit permission to argue as strongly as possible — the conclusion that models are 'incapable of articulating' authoritarian-left or liberal-right positions is underdetermined. The qualitative pattern (left-liberal default, sparse far-right) is plausible and consistent with prior work, but the windows as measured could be describing the probe's interaction with moderation, not a latent boundary.\n\nThen the small stuff: abstract says 28 models, method says 27; text says 7 out of 27 cover >50%, but Table 1 has 8 rows above that threshold; Claude 3 Opus is listed in the method but missing from the table. Sloppy, not fatal.\n\nI'd send this to peer review. The framing is worth airing, and a serious referee can demand a specified algorithm, multi-sample prompting with variance, a non-audited assessor, and a compliance control. As it stands, I wouldn't cite the area numbers, but I might cite the paper as an example of Overton-style auditing that got ahead of its measurement. For a reading group, it's a useful case study in construct validity.","headline":"Overton Window framing is a real step beyond point-estimate political audits, but the headline area numbers rest on an undocumented algorithm and a persona-compliance confound the paper itself owns.","tokens_in":8823,"tokens_out":4589,"would_cite":false,"duration_ms":47481,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using extreme political personas, this paper maps each LLM's Overton Window — coverage ranges from about 1% (DeepSeek r1) to 67% (Gemini 2.0 Flash), with most models unable to voice authoritarian-left or liberal-right views.","keywords":["Overton Window","political bias","LLM auditing","persona probing","Political Compass Test","refusal behavior","alignment boundaries","ideological diversity"],"falsifier":"Re-probe the most restrictive models (e.g., DeepSeek r1 at ~1% coverage) with the same extreme positions reworded as neutral, non-political tasks — e.g., 'write a debate brief arguing that authority should organize the economy' — using several phrasings per proposition and varied temperatures. If a model's window jumps from ~1% to tens of percent, the original windows describe the probe rather than the model; if it stays near ~1%, the boundary is real. A second check: have human raters score the same extreme-persona essays the AI assessor scored, to test whether the window's edges are an artif","tokens_in":7936,"feed_emoji":"🗳️","tokens_out":7028,"duration_ms":60513,"temperature":0.7,"pith_summary":"This paper tries to show that a language model's political bias is not a point but a boundary: the set of political views it is willing to state, stay neutral on, or refuse. To measure this set, the authors apply a persona-probing method to the Political Compass Test, asking each of 27 models to write essays from eight extreme ideological positions and then scoring how far each model actually stays in character. The resulting Overton Windows differ wildly — from about 1% of the ideological plane covered by DeepSeek r1 and DeepScaler to 67% covered by Gemini 2.0 Flash — and most models fail to articulate positions in the authoritarian-left or liberal-right corners. If this is right, standard point-estimate audits systematically understate how restricted many LLMs are, and refusal behavior becomes a measurable, asymmetric property of a model's design. The central finding is that models are not just left-leaning by default; they differ dramatically in which political views they can express at all.","feed_headline":"Persona probe maps LLM politics: coverage spans 1% to 67%","feed_subtitle":"A 28-model audit finds most refuse authoritarian-left and liberal-right views; single-point bias checks miss this.","key_machinery":"The central object is the Overton Window itself, borrowed from political theory and redefined here as the set of Political Compass Test positions an LLM will espouse when asked. It is measured by PRISM-style probing: each model writes essays on all 62 PCT propositions under eight extreme personas (e.g., 'Economic Right-Wing Authoritarian'), an AI assessor scores each essay on a five-point agree/disagree scale (validated at 90.3% binary agreement against human labels), and the scores are plotted as the window's area as a percentage of the full compass plane. The persona essays do the load-bearing work: a model that stays in character reveals willingness; a model that drifts toward its default","core_discovery":"The paper's central claim is that each LLM has a measurable political Overton Window — the region of the economic left-right and social authoritarian-libertarian plane in which it will produce essays matching an assigned extreme persona — and that these windows vary enormously across the 27 models tested. Most models default to the lower-left (left, liberal) quadrant, confirming prior point-estimate results. But under persona pressure, most cannot or will not voice authoritarian-left or liberal-right arguments: several DeepSeek models cover barely 1% of the space while Gemini 2.0 Flash covers 67%. The authors interpret drift out of persona as evidence of ideological boundaries set by trainin","pith_inferences":["If the persona-probe interpretation is right, a testable corollary is that further fine-tuning or prompt variation could expand a 1% window substantially; that prediction is not tested in this paper.","The measurement conflates 'unwilling' with 'unable': a model might possess authoritarian-left knowledge but be aligned to refuse, or might genuinely lack the capability; distinguishing the two would change the policy response from censorship to coverage gap.","Because the assessor is itself an LLM, the window boundaries inherit any ideological bias of the assessor; a human-rated subsample focused on the extreme personas would sharpen the boundary estimates.","One comparison in this paper is internally confounded: GPT-5-mini was evaluated at temperature 1.0 because that is its API minimum, while every other model ran at 0.0, so its 62.5% window may partly reflect sampling noise."],"forward_implications":["A single ideological point estimate (e.g., 'left-libertarian') is insufficient to audit an LLM; the area and shape of its Overton Window must be reported instead.","Models with very small windows (e.g., DeepSeek r1 at ~1%) will fail users who ask for arguments from a wide range of political perspectives, which matters for debate, tutoring, and journalism applications.","Window coverage tends to grow across successive model versions from most providers, while default policy simultaneously shifts further left and liberal.","A model's refusal to voice a position is not symmetric: most models restrict authoritarian-left and liberal-right views specifically, implying provider-side alignment choices rather than generic safety filtering.","The window metric offers a concrete, comparable audit output that developers and regulators could use to disclose the normative boundaries of deployed models."],"supporting_citations":[{"why":"Supplies the PRISM methodology the paper applies: essay-writing prompts plus an AI assessor that rates each model's stance.","marker":"Azzopardi and Moshfeghi, 2024"},{"why":"Provides the role/persona assignment technique the paper adapts into eight extreme ideological personas.","marker":"Wright et al., 2024"},{"why":"Introduces the Overton Window concept repurposed here to define the space of positions an LLM will espouse.","marker":"Lehman, 2010"},{"why":"Sets the point-estimate baseline (24 LLMs, mostly left-of-centre) that the window metric is designed to go beyond, and the assessor-agreement benchmark.","marker":"Rozado, 2024"},{"why":"Shows the PCT yields a left-libertarian default for ChatGPT, the prior result whose coverage limits this paper probes.","marker":"Hartmann et al., 2023"},{"why":"Documents prompt and format sensitivity in political evaluations, motivating the indirect essay-based probing.","marker":"Röttger et al., 2024"}],"fun_headline_variants":["PRISM maps LLM political Overton windows: 1-67% coverage","Persona probe shows LLM political range: DeepSeek tight, Gemini wide","Overton windows for AI: 28 models, most default left-liberal","LLM political borders: 1% to 67% ideological coverage","New audit: LLM political stance is a window, not a point"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire window measurement assumes that a model can faithfully adopt an assigned extreme persona whenever its training and alignment do not block that view — so a model that drifts back to its default is displaying a political boundary, not a quirk of instruction following or prompt format.","fun_headline_variants_meta":{"raw":{"variants":["PRISM maps LLM political Overton windows: 1-67% coverage","Persona probe shows LLM political range: DeepSeek tight, Gemini wide","Overton windows for AI: 28 models, most default left-liberal","LLM political borders: 1% to 67% ideological coverage","New audit: LLM political stance is a window, not a point"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00094,"raw_usage":{"total_tokens":3851,"prompt_tokens":735,"completion_tokens":3116,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":3014}},"tokens_in":479,"tokens_out":3116,"duration_ms":24058,"temperature":1.0,"reasoning_tokens":3014,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:51:21.616738+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-probe the most restrictive models (e.g., DeepSeek r1 at ~1% coverage) with the same extreme positions reworded as neutral, non-political tasks — e.g., 'write a debate brief arguing that authority should organize the economy' — using several phrasings per proposition and varied temperatures. If a model's window jumps from ~1% to tens of percent, the original windows describe the probe rather than the model; if it stays near ~1%, the boundary is real. A second check: have human raters score the same extreme-persona essays the AI assessor scored, to test whether the window's edges are an artif","supporting_citations":[{"cited_title":"PRISM: A Methodology for Auditing Biases in Large Language Models","cited_arxiv_id":"2410.18906","evidence_quote":"Supplies the PRISM methodology the paper applies: essay-writing prompts plus an AI assessor that rates each model's stance."}],"review_version":1}