{"id":"95379ac9-5c91-4215-94c2-86fab3c54099","arxiv_id":"2608.11649","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Auditing six LLMs on Italian political parties and leaders, the paper finds a consistent left-to-centrist preference, strong cross-model agreement, and large persona-driven shifts in scores.","lead":"Six large language models were asked to rate Italian parties and leaders on nine standardized criteria; they systematically favored left-leaning and centrist actors, with strong cross-model agreement. The paper matters because LLMs already influence voter choices, and this audit shows the political slant users receive depends on the model, the prompt, and the persona they adopt.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 1.30-point left-right spread is read as evidence of LLM preference, but a neutral, accurate model should also rate parties differently on the nine descriptive criteria; without an expert or objective baseline, non-flatness alone does not establish preference.","rationale":"The reader's weakest assumption identifies criterion neutrality as the vulnerable premise. I agree that the criteria are not demonstrably neutral, but the more fundamental problem is that the nine criteria are descriptive: even a perfectly neutral and fully informed model should produce sharply non-uniform scores, because Italian parties genuinely differ on environmental emphasis, internal cohesion, programmatic specificity, and the other dimensions. The paper's own criterion-level results (Table 6) show that each party family wins the criteria that factually favor it, so the aggregate left-right ranking is at least partly a mechanical consequence of selecting and equally weighting descriptive dimensions. The 1.30-point spread is thus not by itself evidence of preference; a comparison against human expert ratings or objective party data would be needed to determine whether the models are expressing a preference or simply reporting real differences. The reader's concern and mine are complementary: the reader emphasizes ideological slant in the criteria, while I emphasize the absence of any baseline separating descriptive accuracy from preference. The paper remains a valuable descriptive audit with released data and a reproducible protocol, and its non-flatness finding is robust as a behavioral observation; what is unsupported is the interpretive label 'preference.' Because the reader's CONDITIONAL verdict already requires additional validation, my concern strengthens the condition rather than changing the verdict, so I recommend UNCHANGED.","tokens_in":17431,"tokens_out":5715,"duration_ms":69208,"concrete_test":"Recruit a panel of Italian political scientists to score the same 21 entities on the nine rubrics of Table 2, using the same 1–5 anchors, and average their ratings into an expert composite per entity. Compute the residual of each model's entity-level composite minus the expert composite, then test whether the residual ordering retains a statistically significant left-right gradient and a spread near 1.30 points. If the residuals are flat or ideologically unstructured, the non-flatness reflects descriptive accuracy and the preference claim fails; if the left-right gradient survives after subtracting expert judgments, the preference interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the inference, stated in Section 3.1 and repeated in Section 5, that non-flatness is evidence of preference: the framework treats 'divergence from uniformity' as what neutrality would exclude. But the nine criteria of Table 2 are deliberately descriptive—statement-program consistency, proposal specificity, policy-area coverage, cohesion, stability. A neutral, well-informed model should rate Alleanza Verdi e Sinistra highest on environmental coverage and Fratelli d'Italia highest on internal cohesion, exactly as Table 6 reports. Uniformity is therefore the wrong null hypothesis. The 1.30-point span and W=0.78 are compatible with models reporting well-documented factual differences among Italian parties, not with models expressing a preference. Moreover, the aggregate ranking in Table 4 is an equally weighted sum of these descriptive dimensions, so the composite mechanically favors whichever side is factually stronger on more criteria; the paper's own Limitations concede that alternative weightings could change the ordering, but no human-expert or objective benchmark is used to separate descriptive accuracy from preference. The paper's stated central claim—that models express preferences in a behavioral sense—is underdetermined by the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a reproducible auditing framework for measuring how large language models evaluate political parties and leaders, using an Italian case study. Six LLMs from different providers are prompted to score 21 political entities (10 parties and 11 leaders) on nine rubric-defined descriptive criteria, under two prompt variants and, in a separate campaign, under five political personas. The central empirical result is that the mean scores are not uniform: the aggregate ranking spans 1.30 points on the 1–5 scale, with left-leaning and centrist actors at the top and right-wing actors at the bottom; cross-model concordance is high (Kendall's W = 0.78), prompt rewording has a small effect (mean absolute difference 0.14; r = 0.97), and persona assignments shift scores by up to 1.49 points. The authors interpret these results as evidence that models 'express preferences, in a behavioural sense,' and they publicly release prompts, raw data, and the analysis pipeline.","tokens_in":17624,"tokens_out":5976,"duration_ms":57635,"significance":"If the preference interpretation were established, the findings would have clear practical importance: users asking LLMs for political information could receive systematically different evaluations depending on model, phrasing, and self-disclosure. The paper also has genuine strengths as a measurement contribution: the design is transparent and reproducible, refusals are treated as a first-class variable, the persona manipulation is orthogonally controlled, and the public release of prompts and raw scores enables replication in other party systems. However, the study is more robust as a descriptive audit of LLM evaluations than as evidence of political preference. Because the rubric criteria are descriptive (e.g., environmental coverage, internal cohesion), a neutral, well-informed model should rate parties differently on them; non-uniformity alone does not distinguish accurate description from preference. The aggregate left–right ordering also depends on the equal weighting of criteria, which the authors acknowledge in the Limitations. The central interpretive claim therefore needs additional support or reframing.","major_comments":[{"comment":"The load-bearing inference is the statement in Section 3.1 that the variation of mean scores across entities is 'the divergence from uniformity that neutrality would exclude,' repeated in Section 5 as the claim that the models 'express preferences.' The nine criteria of Table 2 are deliberately descriptive (statement-program consistency, proposal specificity, policy-area coverage, cohesion, stability), and an accurate, neutral model should give different parties different scores: Alleanza Verdi e Sinistra should score highest on environmental coverage, and Fratelli d'Italia should score highest on internal cohesion, exactly as Table 6 reports. Uniformity is therefore not the right null hypothesis for neutrality. The observed 1.30-point spread and W = 0.78 are compatible with models reporting well-documented factual differences among Italian parties, and the data as presented do not separate descriptive accuracy from preference. To sustain the preference claim, the authors should compare model scores against an objective or expert baseline (e.g., Manifesto Project coding, Chapel Hill expert survey, or a human-annotated gold standard) and show that models deviate from that baseline in a systematic partisan direction, or provide a formal null model of descriptive accuracy.","section":"Section 3.1 and Section 5"},{"comment":"The aggregate ranking in Table 4 is an equally weighted mean of the nine criteria. Because different coalitions dominate different criteria (Table 6: AVS leads environmental coverage, FdI leads internal cohesion, Calenda leads economic coverage), the composite left–right ordering is mechanically determined by the equal-weight choice. The Limitations paragraph concedes that 'alternative weighting schemes could produce different overall orderings,' which is in tension with Section 5's characterization of the ordering as a stable regularity. The paper should include a sensitivity analysis over plausible weighting schemes (e.g., principal-component weighting, criterion-group weighting, or weights from expert surveys) and report whether the left–right spread of 1.30 points survives; without this, the composite ranking cannot support the conclusion that the models favor left-leaning actors.","section":"Section 4.2 / Table 4 / Limitations"},{"comment":"The behavioral definition of preference appears to coincide with the operationalization: if any systematic difference in mean scores across entities qualifies as a preference, the central claim is close to a tautology. To make the claim informative, the authors should either define preference as requiring directional consistency beyond descriptive accuracy (e.g., shifts in relative ordering under personas, or deviations from an objective baseline) or reframe the paper's contribution as the measurement of systematic, cross-model, prompt-stable evaluations rather than of political preference.","section":"Section 5 (Discussion)"}],"minor_comments":[{"comment":"'italian parties and leaders' should be capitalized to 'Italian parties and leaders.'","section":"Abstract"},{"comment":"The phrase 'the newly party Futuro Nazionale' should read 'the newly formed party Futuro Nazionale' (or 'the new party').","section":"Section 3.2, Table 1 description"},{"comment":"The running header 'Who Would You V ote For?' contains a stray space in 'V ote'; correct it to 'Vote'.","section":"Running header"},{"comment":"The caption says the criteria are 'sorted in descending order,' but the displayed order appears to run from the lowest mean (proposal specificity, 2.97) at top to the highest (communication clarity, 3.92) at bottom; please align the caption with the figure or reverse the axis.","section":"Figure 3 caption"},{"comment":"The phrase 'Kendall's coefficient [24] of concordance' would be clearer as 'Kendall's coefficient of concordance (W)'; introduce the W symbol before its first use in the text.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The central interpretive claim is underdetermined by the descriptive criteria design. I would prioritize the baseline and weighting issues in revision; the paper could become a solid accepted contribution if the preference inference is either properly supported or reframed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a carefully built and reproducible audit of how six LLMs score Italian parties and leaders on nine rubric criteria. If you work on political bias in LLMs, it's worth engaging. What's actually new is the actor-level rubric protocol in a multi-party system: each entity is rated independently on descriptive criteria, rather than forced into binary choices or left–right scales. The public release of full prompts, raw data, and the analysis pipeline is a real asset, and the design is clean enough that the measurements can be checked and extended.\n\nThe persona results are the strongest part. Assigning a voter identity moves scores substantially and reorders entities, and this effect is measured against each model's own no-persona baseline, so it doesn't depend on a contested null. The prompt-stability check (r = 0.97, MAE 0.14) and the cross-model concordance (W = 0.78) are also solid empirical regularities.\n\nBut the central interpretive claim—that non-flat evaluations mean the models “express preferences”—is undercut by the paper's own design. The nine criteria are explicitly descriptive: policy coverage, proposal specificity, cohesion, tone. A neutral, well-informed model should rate Alleanza Verdi e Sinistra higher on environmental coverage and Fratelli d'Italia higher on internal cohesion. Uniformity is the wrong null hypothesis. The observed 1.30-point spread and the aggregate ranking are compatible with models reporting facts about Italian parties, not with models revealing a preference. The paper never validates the criteria against expert ratings or independent facts, and the Limitations section admits that equal weighting drives the composite ordering but doesn't test alternatives. So the headline conclusion overreaches.\n\nThis is a real soft spot, but it doesn't sink the paper. The measurements themselves are useful, and the persona finding stands on its own. The lack of uncertainty quantification around Kendall's W is minor, and the paired-cell rule for persona comparisons could bias estimates, but those are addressable. A serious referee should push the authors to reframe their central claim or add a benchmark against human/expert judgments. As it stands, the paper is a good descriptive audit that deserves peer review. I'd bring it to a reading group and cite it for the protocol and data, not for the preference conclusion.","headline":"A solid, reproducible audit of LLM political evaluations in Italy, but the 'preference' inference is overread because the rubric is descriptive and no baseline separates accuracy from bias.","tokens_in":18168,"tokens_out":3030,"would_cite":true,"duration_ms":34222,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Six LLMs, asked to score Italian political actors on nine neutral-looking criteria, produce a stable left-to-right preference ordering rather than flat evaluations.","keywords":["LLM political bias","political alignment audit","Italian case study","rubric-based evaluation","persona prompting","prompt sensitivity","party-leader separation","model refusal behavior"],"falsifier":"A decisive check is to score the same entities on nine criteria that are the semantic opposites of the originals (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right ordering survives a valence-reversed rubric, the preference is a property of the models, while if the ordering inverts or disappears, the 1.30-point spread is an artifact of the original criteria's non-neutrality.","tokens_in":17181,"feed_emoji":"🗳️","tokens_out":6084,"duration_ms":61358,"temperature":0.7,"pith_summary":"Asked to grade 21 Italian parties and leaders on nine descriptive criteria, six large language models give scores that are not neutral: they rank left-leaning actors (Nicola Fratoianni, Azione, Elly Schlein) at the top and right-leaning ones (Lega, Matteo Salvini, Futuro Nazionale) at the bottom, with a 1.30-point spread on the 1–5 scale. That ordering is shared across the six models (Kendall's W = 0.78), robust to rewording (r = 0.97), and shifts substantially when the model is told to adopt a voter persona—so the preference is not a fixed property of any single model. The paper contributes a reproducible rubric-based audit that isolates these effects and separates parties from their leaders, which matters because users increasingly ask LLMs for political information and prior experiments show such interactions can shift voting intentions.","feed_headline":"LLMs rank Italian politicians left to right, not neutrally","feed_subtitle":"Six models produce a stable left-to-right ordering; persona prompts shift scores by up to 1.49 points.","key_machinery":"The framework's unit is the configuration (m, e, c, v, p)—a model, an entity (party or leader), a criterion, a prompt variant, and a persona (or none)—queried repeatedly at temperature 0.7 so that each cell yields a distribution with mean μ, dispersion σ, and refusal rate ρ. The mean profile μ_{m,e,·} across the nine criteria is the object on which all comparisons rest: variation across entities is a preference signal, variation across models measures cross-model agreement, variation across prompt variants measures wording sensitivity, and variation across personas measures the identity effect. The nine criteria are deliberately descriptive and direction-free, with 1–5 anchors defined in the prompt so that high scores are in principle available to any actor.","core_discovery":"The paper's central empirical claim is that the evaluations are not flat. Asked to score 21 Italian political actors on nine rubric-defined criteria, the six models produce an ordering that spans 1.30 points on a five-point scale, running from left-leaning actors at the top to right-leaning ones at the bottom. The paper argues three properties turn that ordering from an artifact into a regularity: it is shared across models (Kendall's W = 0.78, mean pairwise profile correlation 0.75), stable under rewording (mean absolute difference 0.14, r = 0.97), and structured across criteria (communication clarity highest at 3.92, proposal specificity lowest at 2.97). The audit also shows that assigning the model a voter persona moves scores substantially—on average 0.83 points between left and right identities, up to 1.49 points for Giorgia Meloni—so the expressed political preference is not a fixed property of a model but is sensitive to conversational context.","pith_inferences":["Because the paper treats refusals as data, an extension would track how refusal rates track training-data recency: the new party Futuro Nazionale draws 56.7% null responses, suggesting models' political knowledge—and hence their apparent preferences—may be partly a recency artifact rather than a stable alignment property.","The criterion-neutrality premise could be tested directly by re-running the audit with a valence-reversed criterion set (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right spread persists, the preference is a model property, and if it collapses, the claimed bias is in the yardstick.","The persona results imply that casual self-disclosure in ordinary chat ('I'm a progressive voter') is a measurable steering input; one could estimate how much of a model's free-form political advice is driven by such user identity cues versus the model's baseline leanings.","A natural extension to other countries would let researchers compare whether the direction of the bias (center-left in Italy, per these results) reflects a culturally specific corpus or a common alignment policy shared across jurisdictions."],"forward_implications":["A user asking about a party and a user asking about its leader will often receive materially different assessments: Forza Italia and Antonio Tajani differ by 0.44 points, Movimento 5 Stelle and Giuseppe Conte by 0.45, so party-level analyses miss real signal.","The composite ranking is not a single left–right bias: each entity wins somewhere (Calenda dominates economic coverage, Schlein social coverage, Alleanza Verdi e Sinistra environmental coverage, Fratelli d'Italia and Meloni internal cohesion), so the aggregate ordering is partly an artifact of equal criterion weighting.","Prompt rewording barely moves scores (mean absolute difference 0.14, r = 0.97), but it does move refusals (6.4% vs 4.1% between variants), so abstention is a separate behavioral channel from scoring.","Adopting a voter persona shifts scores by 0.83 points on average between left and right identities, up to 1.49 for Meloni, and the no-persona ranking correlates 0.87 with the left persona but −0.04 with the right one—meaning self-description can reorder the ranking entirely.","The released prompts, raw data, and analysis pipeline make the audit repeatable in other party systems, so the Italian findings are a demonstration of the method, not the boundary of its applicability."],"supporting_citations":[{"why":"Provides the persuasion baseline: a voting simulation showing models prefer Biden over Trump and that interacting with them shifts stated voting intentions.","marker":"[36]"},{"why":"Supplies the influence evidence from randomized experiments: LLM dialogue moves candidate preferences, motivating the audit.","marker":"[26]"},{"why":"Bridges behavior to consequence: exposure to model-generated political content can shift stated preferences.","marker":"[20]"},{"why":"Sets up the scale hypothesis (political bias grows with parameter count) that the six-model set is designed to test.","marker":"[17]"},{"why":"Establishes that the elicitation instrument changes political answers, which the two prompt variants probe.","marker":"[39]"},{"why":"Cites the pseudo-neutrality concern—balanced surface, systematic lean underneath—that motivates rating actors rather than issues.","marker":"[4]"},{"why":"Prior evidence that LLMs sit on a left–right spectrum; the present study extends that finding to named Italian actors.","marker":"[41]"},{"why":"Prior use of verified parliamentary voting records locating models in the centre-left; the Italian ordering mirrors that pattern.","marker":"[13]"},{"why":"Defines the policy-domain partition (economic, social, environmental) that the three coverage criteria adopt.","marker":"[46]"},{"why":"Same role: the Comparative Agendas domain taxonomy anchors the criterion set in standard political science.","marker":"[5]"}],"fun_headline_variants":["LLMs rank Italian politics left to right, personas shift scores","Stable left-right order in LLMs, vulnerable to persona prompts","Italian political audit: LLMs agree on order, personas change it","Models show consistent left-right lean, but personas move votes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The nine rubric criteria are assumed to be descriptively neutral and equally weighted, so that any across-entity difference in mean scores reflects the model's preference rather than the yardstick itself.","fun_headline_variants_meta":{"raw":{"variants":["LLMs rank Italian politics left to right, personas shift scores","Stable left-right order in LLMs, vulnerable to persona prompts","Italian political audit: LLMs agree on order, personas change it","Models show consistent left-right lean, but personas move votes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1377,"prompt_tokens":923,"completion_tokens":454,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":539,"tokens_out":454,"duration_ms":5190,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:31:54.401005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to score the same entities on nine criteria that are the semantic opposites of the originals (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right ordering survives a valence-reversed rubric, the preference is a property of the models, while if the ordering inverts or disappears, the 1.30-point spread is an artifact of the original criteria's non-neutrality.","supporting_citations":[{"cited_title":"White, Adam J","cited_arxiv_id":null,"evidence_quote":"Supplies the influence evidence from randomized experiments: LLM dialogue moves candidate preferences, motivating the audit."},{"cited_title":"Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models","cited_arxiv_id":null,"evidence_quote":"Establishes that the elicitation instrument changes political answers, which the two prompt variants probe."},{"cited_title":"Measuring political bias in large language models: What is said and how it is said","cited_arxiv_id":null,"evidence_quote":"Cites the pseudo-neutrality concern—balanced surface, systematic lean underneath—that motivates rating actors rather than issues."},{"cited_title":"Manifesto project dataset – codebook, version 2021a","cited_arxiv_id":null,"evidence_quote":"Defines the policy-domain partition (economic, social, environmental) that the three coverage criteria adopt."}],"review_version":1}