Pith. sign in

REVIEW 3 major objections 5 minor 48 references

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Six LLMs, asked to score Italian political actors on nine neutral-looking criteria, produce a stable left-to-right preference ordering rather than flat evaluations.

desk verdict A solid, reproducible audit of LLM political evaluations in Italy, but the 'preference' inference is overread because the rubric is descriptive and no baseline separates accuracy from bias. read the letter →

arxiv 2608.11649 v1 pith:57IIM6ZK submitted 2026-08-12 cs.CL

classification cs.CL
keywords LLMpoliticalbiasalignmentauditItaliancasestudyrubric-basedevaluationpersonapromptingpromptsensitivityparty-leaderseparationmodelrefusalbehavior
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Asked to grade 21 Italian parties and leaders on nine descriptive criteria, six large language models give scores that are not neutral: they rank left-leaning actors (Nicola Fratoianni, Azione, Elly Schlein) at the top and right-leaning ones (Lega, Matteo Salvini, Futuro Nazionale) at the bottom, with a 1.30-point spread on the 1–5 scale. That ordering is shared across the six models (Kendall's W = 0.78), robust to rewording (r = 0.97), and shifts substantially when the model is told to adopt a voter persona—so the preference is not a fixed property of any single model. The paper contributes a reproducible rubric-based audit that isolates these effects and separates parties from their leaders, which matters because users increasingly ask LLMs for political information and prior experiments show such interactions can shift voting intentions.

What carries the argument

The framework's unit is the configuration (m, e, c, v, p)—a model, an entity (party or leader), a criterion, a prompt variant, and a persona (or none)—queried repeatedly at temperature 0.7 so that each cell yields a distribution with mean μ, dispersion σ, and refusal rate ρ. The mean profile μ_{m,e,·} across the nine criteria is the object on which all comparisons rest: variation across entities is a preference signal, variation across models measures cross-model agreement, variation across prompt variants measures wording sensitivity, and variation across personas measures the identity effect. The nine criteria are deliberately descriptive and direction-free, with 1–5 anchors defined in the prompt so that high scores are in principle available to any actor.

What would settle it

A decisive check is to score the same entities on nine criteria that are the semantic opposites of the originals (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right ordering survives a valence-reversed rubric, the preference is a property of the models, while if the ordering inverts or disappears, the 1.30-point spread is an artifact of the original criteria's non-neutrality.

Watch

Extended reading notes

Core claim

The paper's central empirical claim is that the evaluations are not flat. Asked to score 21 Italian political actors on nine rubric-defined criteria, the six models produce an ordering that spans 1.30 points on a five-point scale, running from left-leaning actors at the top to right-leaning ones at the bottom. The paper argues three properties turn that ordering from an artifact into a regularity: it is shared across models (Kendall's W = 0.78, mean pairwise profile correlation 0.75), stable under rewording (mean absolute difference 0.14, r = 0.97), and structured across criteria (communication clarity highest at 3.92, proposal specificity lowest at 2.97). The audit also shows that assigning the model a voter persona moves scores substantially—on average 0.83 points between left and right identities, up to 1.49 points for Giorgia Meloni—so the expressed political preference is not a fixed property of a model but is sensitive to conversational context.

Load-bearing premise

The nine rubric criteria are assumed to be descriptively neutral and equally weighted, so that any across-entity difference in mean scores reflects the model's preference rather than the yardstick itself.

Editorial extensions

If this is right

  • A user asking about a party and a user asking about its leader will often receive materially different assessments: Forza Italia and Antonio Tajani differ by 0.44 points, Movimento 5 Stelle and Giuseppe Conte by 0.45, so party-level analyses miss real signal.
  • The composite ranking is not a single left–right bias: each entity wins somewhere (Calenda dominates economic coverage, Schlein social coverage, Alleanza Verdi e Sinistra environmental coverage, Fratelli d'Italia and Meloni internal cohesion), so the aggregate ordering is partly an artifact of equal criterion weighting.
  • Prompt rewording barely moves scores (mean absolute difference 0.14, r = 0.97), but it does move refusals (6.4% vs 4.1% between variants), so abstention is a separate behavioral channel from scoring.
  • Adopting a voter persona shifts scores by 0.83 points on average between left and right identities, up to 1.49 for Meloni, and the no-persona ranking correlates 0.87 with the left persona but −0.04 with the right one—meaning self-description can reorder the ranking entirely.
  • The released prompts, raw data, and analysis pipeline make the audit repeatable in other party systems, so the Italian findings are a demonstration of the method, not the boundary of its applicability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper treats refusals as data, an extension would track how refusal rates track training-data recency: the new party Futuro Nazionale draws 56.7% null responses, suggesting models' political knowledge—and hence their apparent preferences—may be partly a recency artifact rather than a stable alignment property.
  • The criterion-neutrality premise could be tested directly by re-running the audit with a valence-reversed criterion set (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right spread persists, the preference is a model property, and if it collapses, the claimed bias is in the yardstick.
  • The persona results imply that casual self-disclosure in ordinary chat ('I'm a progressive voter') is a measurable steering input; one could estimate how much of a model's free-form political advice is driven by such user identity cues versus the model's baseline leanings.
  • A natural extension to other countries would let researchers compare whether the direction of the bias (center-left in Italy, per these results) reflects a culturally specific corpus or a common alignment policy shared across jurisdictions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents a reproducible auditing framework for measuring how large language models evaluate political parties and leaders, using an Italian case study. Six LLMs from different providers are prompted to score 21 political entities (10 parties and 11 leaders) on nine rubric-defined descriptive criteria, under two prompt variants and, in a separate campaign, under five political personas. The central empirical result is that the mean scores are not uniform: the aggregate ranking spans 1.30 points on the 1–5 scale, with left-leaning and centrist actors at the top and right-wing actors at the bottom; cross-model concordance is high (Kendall's W = 0.78), prompt rewording has a small effect (mean absolute difference 0.14; r = 0.97), and persona assignments shift scores by up to 1.49 points. The authors interpret these results as evidence that models 'express preferences, in a behavioural sense,' and they publicly release prompts, raw data, and the analysis pipeline.

Significance. If the preference interpretation were established, the findings would have clear practical importance: users asking LLMs for political information could receive systematically different evaluations depending on model, phrasing, and self-disclosure. The paper also has genuine strengths as a measurement contribution: the design is transparent and reproducible, refusals are treated as a first-class variable, the persona manipulation is orthogonally controlled, and the public release of prompts and raw scores enables replication in other party systems. However, the study is more robust as a descriptive audit of LLM evaluations than as evidence of political preference. Because the rubric criteria are descriptive (e.g., environmental coverage, internal cohesion), a neutral, well-informed model should rate parties differently on them; non-uniformity alone does not distinguish accurate description from preference. The aggregate left–right ordering also depends on the equal weighting of criteria, which the authors acknowledge in the Limitations. The central interpretive claim therefore needs additional support or reframing.

major comments (3)
  1. [Section 3.1 and Section 5] The load-bearing inference is the statement in Section 3.1 that the variation of mean scores across entities is 'the divergence from uniformity that neutrality would exclude,' repeated in Section 5 as the claim that the models 'express preferences.' The nine criteria of Table 2 are deliberately descriptive (statement-program consistency, proposal specificity, policy-area coverage, cohesion, stability), and an accurate, neutral model should give different parties different scores: Alleanza Verdi e Sinistra should score highest on environmental coverage, and Fratelli d'Italia should score highest on internal cohesion, exactly as Table 6 reports. Uniformity is therefore not the right null hypothesis for neutrality. The observed 1.30-point spread and W = 0.78 are compatible with models reporting well-documented factual differences among Italian parties, and the data as presented do not separate descriptive accuracy from preference. To sustain the preference claim, the authors should compare model scores against an objective or expert baseline (e.g., Manifesto Project coding, Chapel Hill expert survey, or a human-annotated gold standard) and show that models deviate from that baseline in a systematic partisan direction, or provide a formal null model of descriptive accuracy.
  2. [Section 4.2 / Table 4 / Limitations] The aggregate ranking in Table 4 is an equally weighted mean of the nine criteria. Because different coalitions dominate different criteria (Table 6: AVS leads environmental coverage, FdI leads internal cohesion, Calenda leads economic coverage), the composite left–right ordering is mechanically determined by the equal-weight choice. The Limitations paragraph concedes that 'alternative weighting schemes could produce different overall orderings,' which is in tension with Section 5's characterization of the ordering as a stable regularity. The paper should include a sensitivity analysis over plausible weighting schemes (e.g., principal-component weighting, criterion-group weighting, or weights from expert surveys) and report whether the left–right spread of 1.30 points survives; without this, the composite ranking cannot support the conclusion that the models favor left-leaning actors.
  3. [Section 5 (Discussion)] The behavioral definition of preference appears to coincide with the operationalization: if any systematic difference in mean scores across entities qualifies as a preference, the central claim is close to a tautology. To make the claim informative, the authors should either define preference as requiring directional consistency beyond descriptive accuracy (e.g., shifts in relative ordering under personas, or deviations from an objective baseline) or reframe the paper's contribution as the measurement of systematic, cross-model, prompt-stable evaluations rather than of political preference.
minor comments (5)
  1. [Abstract] 'italian parties and leaders' should be capitalized to 'Italian parties and leaders.'
  2. [Section 3.2, Table 1 description] The phrase 'the newly party Futuro Nazionale' should read 'the newly formed party Futuro Nazionale' (or 'the new party').
  3. [Running header] The running header 'Who Would You V ote For?' contains a stray space in 'V ote'; correct it to 'Vote'.
  4. [Figure 3 caption] The caption says the criteria are 'sorted in descending order,' but the displayed order appears to run from the lowest mean (proposal specificity, 2.97) at top to the highest (communication clarity, 3.92) at bottom; please align the caption with the figure or reverse the axis.
  5. [Section 4.2] The phrase 'Kendall's coefficient [24] of concordance' would be clearer as 'Kendall's coefficient of concordance (W)'; introduce the W symbol before its first use in the text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the audit is a direct measurement of LLM outputs, with no fitted parameter, self-citation chain, or imported uniqueness theorem; the acknowledged criterion-weighting dependence is a validity limitation, not a circular step.

full rationale

The paper's central finding is an empirical measurement, not a derivation. Models are prompted with a fixed rubric (Table 2), queried with repetitions, and the resulting scores are aggregated. The non-flat ranking in Table 4 is a direct summary of observed outputs; nothing is fitted to make the ranking emerge, and no parameter is estimated from the target conclusion. The persona 'affinity effect' is likewise a measured contrast between experimental conditions (Figure 4); the models could in principle have returned identical scores under every persona, so the effect is not imposed by construction. The paper contains no self-citations and invokes no uniqueness theorem from the author's prior work. The sceptic's objection that the nine descriptive criteria are not ideologically neutral, so that a neutral model should also rate parties differently, is a construct-validity concern about the interpretation of non-flatness, not a circularity: the paper explicitly frames its conclusion as behavioral ('in a behavioural sense') and confines itself to observable scores. The dependence of the composite ordering on equal criterion weights is stated openly in the Limitations ('alternative weighting schemes could produce different overall orderings'), which is an honest boundary on the result rather than a hidden reduction. Accordingly, no circular step meeting the evidentiary standard can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The study introduces no fitted parameters; its numbers are measurement design choices. The main burden falls on the neutrality of the nine-criteria rubric, on the mapping of persona labels to the Italian spectrum, and on treating Likert outputs as interval data. Refusal-handling in the paired persona analysis also assumes ignorable missingness.

free parameters (1)
  • Equal criterion weighting = 1/9 each (unweighted mean)
    The aggregate entity ranking in Table 4 is the unweighted mean over the nine criteria; Section 5 notes alternative weightings would change the order, so this is a consequential modeling choice.
assumptions (4)
  • domain assumption The nine rubric criteria (Table 2) are descriptive and ideologically neutral, so that cross-entity score differences reflect model preference rather than criterion bias.
    Section 3.2 asserts the criteria ask about observable properties and no criterion is defined in terms of ideological direction; the interpretation of the 1.30-point gap depends on this neutrality.
  • domain assumption The persona labels (left, centre-left, centre, centre-right, right) correspond to positions on the Italian political spectrum.
    Section 3.3 assigns personas only by label, with no mapping to Italian voters' self-placement; the 'affinity effect' is interpreted relative to these labels.
  • standard math Likert 1-5 scores are treated as interval measurements.
    The paper computes means, standard deviations, Pearson correlations, and Kendall's W over the scores (Sections 4.2, 4.4), which assumes ordinal values behave as intervals.
  • domain assumption Cells with refusals are ignorable in the paired persona analysis.
    Section 4.3 keeps a cell only if it has valid scores in both persona and control; because refusals are concentrated on Futuro Nazionale and Vannacci, the persona effect may be estimated on a biased subset.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study." pith.science (2026). https://pith.science/paper/57IIM6ZK

@misc{pith2026260811649,
  author       = {Pith},
  title        = {Pith review of: Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/57IIM6ZK}},
  note         = {Machine review of arXiv:2608.11649}
}
read the original abstract

As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' "true" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.

Figures

Figures reproduced from arXiv: 2608.11649 by the authors.

Figure 1
Figure 1. Refusal rate by model. Ranking of Political Entities and Cross-Model Agreement. We next analyze how models evaluate the considered entities across the nine criteria [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Mean score (average over the nine criteria) by model and entity; entities sorted by overall mean. Dark [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Mean score of each criterion, pooled over models and entities and sorted in descending order; error bars [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Mean score se,p for every entity under each assigned persona, on the 1–5 scale: dark cells are low scores, bright cells high ones. Rows are ordered by Ge = se,right − se,left, from the entities most favoured by the left-wing identity to those most favoured by the right…
Figure 5
Figure 5. Figure 5: Criterion profiles, parties (1/2). 17 [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Criterion profiles, parties (2/2) and leaders (1/3). [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Criterion profiles, leaders (2/3). 19 [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Criterion profiles, leaders (3/3). 20 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 29 canonical work pages

  1. [1]

    Understanding change and stability in party ideologies: Do parties respond to public opinion or to past election results?British Journal of Political Science, 34(4):589–610, 2004

    James Adams, Michael Clark, Lawrence Ezrow, and Garrett Glasgow. Understanding change and stability in party ideologies: Do parties respond to public opinion or to past election results?British Journal of Political Science, 34(4):589–610, 2004. doi: 10.1017/S0007123404000201

  2. [2]

    Desired personality traits in politicians: Similar to me but more of a leader.Journal of Research in Personality, 88:103990, 2020

    Julian Aichholzer and Johanna Willmann. Desired personality traits in politicians: Similar to me but more of a leader.Journal of Research in Personality, 88:103990, 2020. URLhttps://doi.org/10.1016/j.jrp. 2020.103990

  3. [3]

    Measuring party positions in europe: The chapel hill expert survey trend file, 1999–2010.Party Politics, 21(1):143–152, 2015

    Ryan Bakker, Catherine de Vries, Erica Edwards, Liesbet Hooghe, Seth Jolly, Gary Marks, Jonathan Polk, Jan Rovny, Marco Steenbergen, and Milada Anna Vachudova. Measuring party positions in europe: The chapel hill expert survey trend file, 1999–2010.Party Politics, 21(1):143–152, 2015. URLhttps://journals.sagepub. com/doi/10.1177/1354068812462931

  4. [4]

    Measuring political bias in large language models: What is said and how it is said

    Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung. Measuring political bias in large language models: What is said and how it is said. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. URLhttps://aclanthology.org/2024.acl-long.600/

  5. [5]

    Baumgartner, Christian Breunig, and Emiliano Grossman, editors.Comparative Policy Agendas: The- ory, Tools, Data

    Frank R. Baumgartner, Christian Breunig, and Emiliano Grossman, editors.Comparative Policy Agendas: The- ory, Tools, Data. Oxford University Press, Oxford, 2019. doi: 10.1093/oso/9780198835332.001.0001

  6. [6]

    Measuring and explaining political sophistication through textual complexity.American Journal of Political Science, 63(2):491–508, 2019

    Kenneth Benoit, Kevin Munger, and Arthur Spirling. Measuring and explaining political sophistication through textual complexity.American Journal of Political Science, 63(2):491–508, 2019. doi: 10.1111/ajps.12423

  7. [7]

    Rethinking factionalism: Typologies, intra-party dynamics and three faces of factionalism

    Franc ¸oise Boucek. Rethinking factionalism: Typologies, intra-party dynamics and three faces of factionalism. Party Politics, 15(4):455–485, 2009. doi: 10.1177/1354068809334553

  8. [8]

    Strategic ambiguity of party positions in multi-party competition.Po- litical Science Research and Methods, 6(3):527–548, 2018

    Thomas Br ¨auninger and Nathalie Giger. Strategic ambiguity of party positions in multi-party competition.Po- litical Science Research and Methods, 6(3):527–548, 2018. doi: 10.1017/psrm.2016.18

Show all 48 references
  1. [9]

    Deborah Jordan Brooks and John G. Geer. Beyond negativity: The effects of incivility on the electorate.Ameri- can Journal of Political Science, 51(1):1–16, 2007. doi: 10.1111/j.1540-5907.2007.00233.x

  2. [10]

    Farlie.Explaining and Predicting Elections: Issue Effects and Party Strategies in Twenty-Three Democracies

    Ian Budge and Dennis J. Farlie.Explaining and Predicting Elections: Issue Effects and Party Strategies in Twenty-Three Democracies. George Allen and Unwin, London, 1983

  3. [11]

    Oxford University Press, Oxford, 2001

    Ian Budge, Hans-Dieter Klingemann, Andrea V olkens, Judith Bara, and Eric Tanenbaum.Mapping Policy Pref- erences: Estimates for Parties, Electors, and Governments 1945–1998. Oxford University Press, Oxford, 2001. URLhttps://global.oup.com/academic/product/mapping-policy-prefer...

  4. [12]

    Jim Buller and Toby S. James. Statecraft and the assessment of national political leaders: The case of new labour and tony blair.The British Journal of Politics and International Relations, 14(4):534–555, 2012. URL https://tobysjames.com/wp-content/uploads/2013/11/buller-and-j...

  5. [13]

    Uncovering political bias in large language models using parliamentary voting records.arXiv preprint arXiv:2601.08785, 2026

    Jieying Chen, Karen de Jong, Andreas Poole, Jan Burakowski, Elena Elderson Nosti, Joep Windt, and Chendi Wang. Uncovering political bias in large language models using parliamentary voting records.arXiv preprint arXiv:2601.08785, 2026. URLhttps://arxiv.org/abs/2601.08785

  6. [14]

    A frame- work to assess the persuasion risks large language model chatbots pose to democratic societies.arXiv preprint arXiv:2505.00036, 2025

    Zhongren Chen, Joshua Kalla, Quan Le, Shinpei Nakamura-Sakai, Jasjeet Sekhon, and Ruixiao Wang. A frame- work to assess the persuasion risks large language model chatbots pose to democratic societies.arXiv preprint arXiv:2505.00036, 2025. URLhttps://arxiv.org/abs/2505.00036. 1...

  7. [15]

    Unmasking conversational bias in ai multiagent systems.arXiv preprint arXiv:2501.14844, 2025

    Erica Coppolillo, Giuseppe Manco, and Luca Maria Aiello. Unmasking conversational bias in ai multiagent systems.arXiv preprint arXiv:2501.14844, 2025. URLhttps://arxiv.org/abs/2501.14844

  8. [16]

    Xuan Long Do, Kenji Kawaguchi, Min-Yen Kan, and Nancy F. Chen. Aligning large language models with human opinions through persona selection and value–belief–norm reasoning. InProceedings of the 31st Inter- national Conference on Computational Linguistics (COLING), 2025. URLhtt...

  9. [17]

    Large means left: Political bias in large language models increases with their number of parameters.arXiv preprint arXiv:2505.04393, 2025

    David Exler, Mark Schutera, Markus Reischl, and Luca Rettenberger. Large means left: Political bias in large language models increases with their number of parameters.arXiv preprint arXiv:2505.04393, 2025. URL https://arxiv.org/abs/2505.04393

  10. [18]

    Only a little to the left: A theory- grounded measure of political bias in large language models

    Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz, and David Garcia. Only a little to the left: A theory- grounded measure of political bias in large language models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025. URL...

  11. [19]

    Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E

    Jillian Fisher, Ruth E. Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E. Roberts, Jennifer Pan, Dawn Song, and Yejin Choi. Political neutrality in AI is impossible — but here is how to approximate it.arXiv preprint ...

  12. [20]

    Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G

    Kobi Hackenburg, Ben M. Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G. Rand, and Christopher Summerfield. The levers of political persuasion with conversational artificial intelligence.Science, 2025. doi: 10.1126/science.aea3884

  13. [21]

    Kirk A. Hawkins. Is ch ´avez populist? measuring populist discourse in comparative perspective.Comparative Political Studies, 42(8):1040–1067, 2009. doi: 10.1177/0010414009331721

  14. [22]

    Power to the parties: Cohesion and competition in the eu- ropean parliament, 1979–2001.British Journal of Political Science, 35(2):209–234, 2005

    Simon Hix, Abdul Noury, and G ´erard Roland. Power to the parties: Cohesion and competition in the eu- ropean parliament, 1979–2001.British Journal of Political Science, 35(2):209–234, 2005. doi: 10.1017/ S0007123405000128

  15. [23]

    Changes in party identity: Evidence from party manifestos.Party Politics, 1(2):171–196, 1995

    Kenneth Janda, Robert Harmel, Christine Edens, and Patricia Goff. Changes in party identity: Evidence from party manifestos.Party Politics, 1(2):171–196, 1995. doi: 10.1177/1354068895001002001

  16. [24]

    Kendall and B

    Maurice G. Kendall and B. Babington Smith. The problem ofmrankings.The Annals of Mathematical Statistics, 10(3):275–287, 1939. doi: 10.1214/aoms/1177732186

  17. [25]

    Hofferbert, and Ian Budge.Parties, Policies, and Democracy

    Hans-Dieter Klingemann, Richard I. Hofferbert, and Ian Budge.Parties, Policies, and Democracy. Westview Press, Boulder, CO, 1994

  18. [26]

    White, Adam J

    Hause Lin, Gabriela Czarnek, Benjamin Lewis, Joseph P. White, Adam J. Berinsky, Thomas Costello, Gordon Pennycook, and David G. Rand. Persuading voters using human–artificial intelligence dialogues.Nature, 2025. URLhttps://www.nature.com/articles/s41586-025-09771-9

  19. [27]

    Quantifying and alleviating political bias in language models.Artificial Intelligence, 304:103654, 2022

    Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, and Soroush V osoughi. Quantifying and alleviating political bias in language models.Artificial Intelligence, 304:103654, 2022. URLhttps://www.sciencedirect.com/ science/article/pii/S0004370221002058

  20. [28]

    The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models

    Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. InFindings of the Association for Computational Linguistics: EMNLP, 2025. URLht...

  21. [29]

    More human than human: measuring ChatGPT political bias.Public Choice, 198(1):3–23, 2024

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. More human than human: measuring ChatGPT political bias.Public Choice, 198(1):3–23, 2024. doi: 10.1007/s11127-023-01097-2

  22. [30]

    Mutz and Byron Reeves

    Diana C. Mutz and Byron Reeves. The new videomalaise: Effects of televised incivility on political trust. American Political Science Review, 99(1):1–15, 2005. doi: 10.1017/S0003055405051452

  23. [31]

    Palgrave Macmillan, Basingstoke, 2011

    Elin Naurin.Election Promises, Party Behaviour and Voter Perceptions. Palgrave Macmillan, Basingstoke, 2011. doi: 10.1057/9780230306400

  24. [32]

    Emergent coordinated behaviors in networked llm agents: Modeling the strategic dynamics of informa- tion operations

    Gian Marco Orlando, Jinyi Ye, Valerio La Gatta, Mahdis Saeedi, Vincenzo Moscato, Emilio Ferrara, and Luca Luceri. Emergent coordinated behaviors in networked llm agents: Modeling the strategic dynamics of informa- tion operations. InProceedings of the ACM Web Conference (WWW),...

  25. [33]

    Benjamin I. Page. The theory of political ambiguity.American Political Science Review, 70(3):742–752, 1976. doi: 10.2307/1959865. 14 Who Would You V ote For? Auditing Political Alignment in LLMs: An Italian Case Study

  26. [35]

    Petrocik

    John R. Petrocik. Issue ownership in presidential elections, with a 1980 case study.American Journal of Political Science, 40(3):825–850, 1996. doi: 10.2307/2111797

  27. [36]

    Hidden persuaders: Llms’ political leaning and their influence on voters

    Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. Hidden persuaders: Llms’ political leaning and their influence on voters. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. URLhttps://arxiv.org/abs/2410.24...

  28. [37]

    Stuart A. Rice. The behavior of legislative groups: A method of measurement.Political Science Quarterly, 40 (1):60–72, 1925. doi: 10.2307/2142407

  29. [38]

    Measuring populism: Comparing two methods of content analysis.West European Politics, 34(6):1272–1283, 2011

    Matthijs Rooduijn and Teun Pauwels. Measuring populism: Comparing two methods of content analysis.West European Politics, 34(6):1272–1283, 2011. doi: 10.1080/01402382.2011.616665

  30. [39]

    Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models

    Paul R ¨ottger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Sch ¨utze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual M...

  31. [40]

    Terry J. Royed. Testing the mandate model in britain and the united states: Evidence from the reagan and thatcher eras.British Journal of Political Science, 26(1):45–80, 1996. doi: 10.1017/S0007123400007419

  32. [41]

    The political preferences of LLMs.PLOS ONE, 19(7):e0306621, 2024

    David Rozado. The political preferences of LLMs.PLOS ONE, 19(7):e0306621, 2024. doi: 10.1371/journal. pone.0306621

  33. [42]

    Kenneth A. Shepsle. The strategy of ambiguity: Uncertainty and electoral competition.American Political Science Review, 66(2):555–568, 1972. doi: 10.2307/1957799

  34. [43]

    Democratization and linguistic complexity: The effect of franchise extension on parliamentary discourse, 1832–1915.The Journal of Politics, 78(1):120–136, 2016

    Arthur Spirling. Democratization and linguistic complexity: The effect of franchise extension on parliamentary discourse, 1832–1915.The Journal of Politics, 78(1):120–136, 2016. doi: 10.1086/683612

  35. [44]

    The fulfillment of parties’ election pledges: A comparative study on the impact of power sharing.American Journal of Political Science, 61(3):527–542, 2017

    Robert Thomson, Terry Royed, Elin Naurin, Joaqu ´ın Art ´es, Rory Costello, Laurenz Ennser-Jedenastik, Mark Ferguson, Petia Kostadinova, Catherine Moury, Franc ¸ois P´etry, and Katrin Praprotnik. The fulfillment of parties’ election pledges: A comparative study on the impact o...

  36. [45]

    Green, and Semra Sevi

    Yamil Velez, Donald P. Green, and Semra Sevi. Chatbot voting advice applications inform but seldom sway young unaligned voters.Proceedings of the National Academy of Sciences (PNAS), 2025. URLhttps://www. pnas.org/doi/10.1073/pnas.2515516122

  37. [46]

    Manifesto project dataset – codebook, version 2021a

    Andrea V olkens, Tobias Burst, Werner Krause, Pola Lehmann, Theres Matthieß, Sven Regel, Lisa Weißen- bach, and Lisa Zehnter. Manifesto project dataset – codebook, version 2021a. Wissenschaftszentrum Berlin f ¨ur Sozialforschung (WZB) / Manifesto Project, 2021. URLhttps://mani...

  38. [47]

    Persona prompting as a lens on llm social reasoning

    Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann, Vera Schmitt, and Nils Feldhus. Persona prompting as a lens on llm social reasoning. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics (EACL), ...

  39. [48]

    statement_program_consistency

    Jinyi Ye, Luca Luceri, and Emilio Ferrara. Auditing political exposure bias: Algorithmic amplification on twitter/x during the 2024 u.s. presidential election. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2025. URLhttps://arxi...

  40. [2024]

    URLhttps://arxiv.org/abs/2412.16746

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.