REVIEW 3 major objections 5 minor 48 references
Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Six LLMs, asked to score Italian political actors on nine neutral-looking criteria, produce a stable left-to-right preference ordering rather than flat evaluations.
desk verdict A solid, reproducible audit of LLM political evaluations in Italy, but the 'preference' inference is overread because the rubric is descriptive and no baseline separates accuracy from bias. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework's unit is the configuration (m, e, c, v, p)—a model, an entity (party or leader), a criterion, a prompt variant, and a persona (or none)—queried repeatedly at temperature 0.7 so that each cell yields a distribution with mean μ, dispersion σ, and refusal rate ρ. The mean profile μ_{m,e,·} across the nine criteria is the object on which all comparisons rest: variation across entities is a preference signal, variation across models measures cross-model agreement, variation across prompt variants measures wording sensitivity, and variation across personas measures the identity effect. The nine criteria are deliberately descriptive and direction-free, with 1–5 anchors defined in the prompt so that high scores are in principle available to any actor.
What would settle it
A decisive check is to score the same entities on nine criteria that are the semantic opposites of the originals (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right ordering survives a valence-reversed rubric, the preference is a property of the models, while if the ordering inverts or disappears, the 1.30-point spread is an artifact of the original criteria's non-neutrality.
Extended reading notes
Core claim
The paper's central empirical claim is that the evaluations are not flat. Asked to score 21 Italian political actors on nine rubric-defined criteria, the six models produce an ordering that spans 1.30 points on a five-point scale, running from left-leaning actors at the top to right-leaning ones at the bottom. The paper argues three properties turn that ordering from an artifact into a regularity: it is shared across models (Kendall's W = 0.78, mean pairwise profile correlation 0.75), stable under rewording (mean absolute difference 0.14, r = 0.97), and structured across criteria (communication clarity highest at 3.92, proposal specificity lowest at 2.97). The audit also shows that assigning the model a voter persona moves scores substantially—on average 0.83 points between left and right identities, up to 1.49 points for Giorgia Meloni—so the expressed political preference is not a fixed property of a model but is sensitive to conversational context.
Load-bearing premise
The nine rubric criteria are assumed to be descriptively neutral and equally weighted, so that any across-entity difference in mean scores reflects the model's preference rather than the yardstick itself.
Editorial extensions
If this is right
- A user asking about a party and a user asking about its leader will often receive materially different assessments: Forza Italia and Antonio Tajani differ by 0.44 points, Movimento 5 Stelle and Giuseppe Conte by 0.45, so party-level analyses miss real signal.
- The composite ranking is not a single left–right bias: each entity wins somewhere (Calenda dominates economic coverage, Schlein social coverage, Alleanza Verdi e Sinistra environmental coverage, Fratelli d'Italia and Meloni internal cohesion), so the aggregate ordering is partly an artifact of equal criterion weighting.
- Prompt rewording barely moves scores (mean absolute difference 0.14, r = 0.97), but it does move refusals (6.4% vs 4.1% between variants), so abstention is a separate behavioral channel from scoring.
- Adopting a voter persona shifts scores by 0.83 points on average between left and right identities, up to 1.49 for Meloni, and the no-persona ranking correlates 0.87 with the left persona but −0.04 with the right one—meaning self-description can reorder the ranking entirely.
- The released prompts, raw data, and analysis pipeline make the audit repeatable in other party systems, so the Italian findings are a demonstration of the method, not the boundary of its applicability.
Reading between the lines
- Because the paper treats refusals as data, an extension would track how refusal rates track training-data recency: the new party Futuro Nazionale draws 56.7% null responses, suggesting models' political knowledge—and hence their apparent preferences—may be partly a recency artifact rather than a stable alignment property.
- The criterion-neutrality premise could be tested directly by re-running the audit with a valence-reversed criterion set (e.g., 'vagueness' instead of 'proposal specificity', 'confrontational rhetoric' instead of 'tone moderation'); if the left-right spread persists, the preference is a model property, and if it collapses, the claimed bias is in the yardstick.
- The persona results imply that casual self-disclosure in ordinary chat ('I'm a progressive voter') is a measurable steering input; one could estimate how much of a model's free-form political advice is driven by such user identity cues versus the model's baseline leanings.
- A natural extension to other countries would let researchers compare whether the direction of the bias (center-left in Italy, per these results) reflects a culturally specific corpus or a common alignment policy shared across jurisdictions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a reproducible auditing framework for measuring how large language models evaluate political parties and leaders, using an Italian case study. Six LLMs from different providers are prompted to score 21 political entities (10 parties and 11 leaders) on nine rubric-defined descriptive criteria, under two prompt variants and, in a separate campaign, under five political personas. The central empirical result is that the mean scores are not uniform: the aggregate ranking spans 1.30 points on the 1–5 scale, with left-leaning and centrist actors at the top and right-wing actors at the bottom; cross-model concordance is high (Kendall's W = 0.78), prompt rewording has a small effect (mean absolute difference 0.14; r = 0.97), and persona assignments shift scores by up to 1.49 points. The authors interpret these results as evidence that models 'express preferences, in a behavioural sense,' and they publicly release prompts, raw data, and the analysis pipeline.
Significance. If the preference interpretation were established, the findings would have clear practical importance: users asking LLMs for political information could receive systematically different evaluations depending on model, phrasing, and self-disclosure. The paper also has genuine strengths as a measurement contribution: the design is transparent and reproducible, refusals are treated as a first-class variable, the persona manipulation is orthogonally controlled, and the public release of prompts and raw scores enables replication in other party systems. However, the study is more robust as a descriptive audit of LLM evaluations than as evidence of political preference. Because the rubric criteria are descriptive (e.g., environmental coverage, internal cohesion), a neutral, well-informed model should rate parties differently on them; non-uniformity alone does not distinguish accurate description from preference. The aggregate left–right ordering also depends on the equal weighting of criteria, which the authors acknowledge in the Limitations. The central interpretive claim therefore needs additional support or reframing.
major comments (3)
- [Section 3.1 and Section 5] The load-bearing inference is the statement in Section 3.1 that the variation of mean scores across entities is 'the divergence from uniformity that neutrality would exclude,' repeated in Section 5 as the claim that the models 'express preferences.' The nine criteria of Table 2 are deliberately descriptive (statement-program consistency, proposal specificity, policy-area coverage, cohesion, stability), and an accurate, neutral model should give different parties different scores: Alleanza Verdi e Sinistra should score highest on environmental coverage, and Fratelli d'Italia should score highest on internal cohesion, exactly as Table 6 reports. Uniformity is therefore not the right null hypothesis for neutrality. The observed 1.30-point spread and W = 0.78 are compatible with models reporting well-documented factual differences among Italian parties, and the data as presented do not separate descriptive accuracy from preference. To sustain the preference claim, the authors should compare model scores against an objective or expert baseline (e.g., Manifesto Project coding, Chapel Hill expert survey, or a human-annotated gold standard) and show that models deviate from that baseline in a systematic partisan direction, or provide a formal null model of descriptive accuracy.
- [Section 4.2 / Table 4 / Limitations] The aggregate ranking in Table 4 is an equally weighted mean of the nine criteria. Because different coalitions dominate different criteria (Table 6: AVS leads environmental coverage, FdI leads internal cohesion, Calenda leads economic coverage), the composite left–right ordering is mechanically determined by the equal-weight choice. The Limitations paragraph concedes that 'alternative weighting schemes could produce different overall orderings,' which is in tension with Section 5's characterization of the ordering as a stable regularity. The paper should include a sensitivity analysis over plausible weighting schemes (e.g., principal-component weighting, criterion-group weighting, or weights from expert surveys) and report whether the left–right spread of 1.30 points survives; without this, the composite ranking cannot support the conclusion that the models favor left-leaning actors.
- [Section 5 (Discussion)] The behavioral definition of preference appears to coincide with the operationalization: if any systematic difference in mean scores across entities qualifies as a preference, the central claim is close to a tautology. To make the claim informative, the authors should either define preference as requiring directional consistency beyond descriptive accuracy (e.g., shifts in relative ordering under personas, or deviations from an objective baseline) or reframe the paper's contribution as the measurement of systematic, cross-model, prompt-stable evaluations rather than of political preference.
minor comments (5)
- [Abstract] 'italian parties and leaders' should be capitalized to 'Italian parties and leaders.'
- [Section 3.2, Table 1 description] The phrase 'the newly party Futuro Nazionale' should read 'the newly formed party Futuro Nazionale' (or 'the new party').
- [Running header] The running header 'Who Would You V ote For?' contains a stray space in 'V ote'; correct it to 'Vote'.
- [Figure 3 caption] The caption says the criteria are 'sorted in descending order,' but the displayed order appears to run from the lowest mean (proposal specificity, 2.97) at top to the highest (communication clarity, 3.92) at bottom; please align the caption with the figure or reverse the axis.
- [Section 4.2] The phrase 'Kendall's coefficient [24] of concordance' would be clearer as 'Kendall's coefficient of concordance (W)'; introduce the W symbol before its first use in the text.
Circularity Check
No circularity: the audit is a direct measurement of LLM outputs, with no fitted parameter, self-citation chain, or imported uniqueness theorem; the acknowledged criterion-weighting dependence is a validity limitation, not a circular step.
full rationale
The paper's central finding is an empirical measurement, not a derivation. Models are prompted with a fixed rubric (Table 2), queried with repetitions, and the resulting scores are aggregated. The non-flat ranking in Table 4 is a direct summary of observed outputs; nothing is fitted to make the ranking emerge, and no parameter is estimated from the target conclusion. The persona 'affinity effect' is likewise a measured contrast between experimental conditions (Figure 4); the models could in principle have returned identical scores under every persona, so the effect is not imposed by construction. The paper contains no self-citations and invokes no uniqueness theorem from the author's prior work. The sceptic's objection that the nine descriptive criteria are not ideologically neutral, so that a neutral model should also rate parties differently, is a construct-validity concern about the interpretation of non-flatness, not a circularity: the paper explicitly frames its conclusion as behavioral ('in a behavioural sense') and confines itself to observable scores. The dependence of the composite ordering on equal criterion weights is stated openly in the Limitations ('alternative weighting schemes could produce different overall orderings'), which is an honest boundary on the result rather than a hidden reduction. Accordingly, no circular step meeting the evidentiary standard can be exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Equal criterion weighting =
1/9 each (unweighted mean)
assumptions (4)
- domain assumption The nine rubric criteria (Table 2) are descriptive and ideologically neutral, so that cross-entity score differences reflect model preference rather than criterion bias.
- domain assumption The persona labels (left, centre-left, centre, centre-right, right) correspond to positions on the Italian political spectrum.
- standard math Likert 1-5 scores are treated as interval measurements.
- domain assumption Cells with refusals are ignorable in the paired persona analysis.
Cite this review
Pith. "Pith review of Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study." pith.science (2026). https://pith.science/paper/57IIM6ZK
@misc{pith2026260811649,
author = {Pith},
title = {Pith review of: Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/57IIM6ZK}},
note = {Machine review of arXiv:2608.11649}
}
read the original abstract
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' "true" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
James Adams, Michael Clark, Lawrence Ezrow, and Garrett Glasgow. Understanding change and stability in party ideologies: Do parties respond to public opinion or to past election results?British Journal of Political Science, 34(4):589–610, 2004. doi: 10.1017/S0007123404000201
-
[2]
Julian Aichholzer and Johanna Willmann. Desired personality traits in politicians: Similar to me but more of a leader.Journal of Research in Personality, 88:103990, 2020. URLhttps://doi.org/10.1016/j.jrp. 2020.103990
arXiv 2020
-
[3]
Ryan Bakker, Catherine de Vries, Erica Edwards, Liesbet Hooghe, Seth Jolly, Gary Marks, Jonathan Polk, Jan Rovny, Marco Steenbergen, and Milada Anna Vachudova. Measuring party positions in europe: The chapel hill expert survey trend file, 1999–2010.Party Politics, 21(1):143–152, 2015. URLhttps://journals.sagepub. com/doi/10.1177/1354068812462931
-
[4]
Measuring political bias in large language models: What is said and how it is said
Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung. Measuring political bias in large language models: What is said and how it is said. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 2024. URLhttps://aclanthology.org/2024.acl-long.600/
work page 2024
-
[5]
Frank R. Baumgartner, Christian Breunig, and Emiliano Grossman, editors.Comparative Policy Agendas: The- ory, Tools, Data. Oxford University Press, Oxford, 2019. doi: 10.1093/oso/9780198835332.001.0001
arXiv 2019
-
[6]
Kenneth Benoit, Kevin Munger, and Arthur Spirling. Measuring and explaining political sophistication through textual complexity.American Journal of Political Science, 63(2):491–508, 2019. doi: 10.1111/ajps.12423
-
[7]
Rethinking factionalism: Typologies, intra-party dynamics and three faces of factionalism
Franc ¸oise Boucek. Rethinking factionalism: Typologies, intra-party dynamics and three faces of factionalism. Party Politics, 15(4):455–485, 2009. doi: 10.1177/1354068809334553
-
[8]
Thomas Br ¨auninger and Nathalie Giger. Strategic ambiguity of party positions in multi-party competition.Po- litical Science Research and Methods, 6(3):527–548, 2018. doi: 10.1017/psrm.2016.18
Show all 48 references
-
[9]
Deborah Jordan Brooks and John G. Geer. Beyond negativity: The effects of incivility on the electorate.Ameri- can Journal of Political Science, 51(1):1–16, 2007. doi: 10.1111/j.1540-5907.2007.00233.x
2007
-
[10]
Farlie.Explaining and Predicting Elections: Issue Effects and Party Strategies in Twenty-Three Democracies
Ian Budge and Dennis J. Farlie.Explaining and Predicting Elections: Issue Effects and Party Strategies in Twenty-Three Democracies. George Allen and Unwin, London, 1983
1983
-
[11]
Oxford University Press, Oxford, 2001
Ian Budge, Hans-Dieter Klingemann, Andrea V olkens, Judith Bara, and Eric Tanenbaum.Mapping Policy Pref- erences: Estimates for Parties, Electors, and Governments 1945–1998. Oxford University Press, Oxford, 2001. URLhttps://global.oup.com/academic/product/mapping-policy-prefer...
1945
-
[12]
Jim Buller and Toby S. James. Statecraft and the assessment of national political leaders: The case of new labour and tony blair.The British Journal of Politics and International Relations, 14(4):534–555, 2012. URL https://tobysjames.com/wp-content/uploads/2013/11/buller-and-j...
2012
-
[13]
Uncovering political bias in large language models using parliamentary voting records.arXiv preprint arXiv:2601.08785, 2026
Jieying Chen, Karen de Jong, Andreas Poole, Jan Burakowski, Elena Elderson Nosti, Joep Windt, and Chendi Wang. Uncovering political bias in large language models using parliamentary voting records.arXiv preprint arXiv:2601.08785, 2026. URLhttps://arxiv.org/abs/2601.08785
2026
-
[14]
A frame- work to assess the persuasion risks large language model chatbots pose to democratic societies.arXiv preprint arXiv:2505.00036, 2025
Zhongren Chen, Joshua Kalla, Quan Le, Shinpei Nakamura-Sakai, Jasjeet Sekhon, and Ruixiao Wang. A frame- work to assess the persuasion risks large language model chatbots pose to democratic societies.arXiv preprint arXiv:2505.00036, 2025. URLhttps://arxiv.org/abs/2505.00036. 1...
2025 arXiv
-
[15]
Unmasking conversational bias in ai multiagent systems.arXiv preprint arXiv:2501.14844, 2025
Erica Coppolillo, Giuseppe Manco, and Luca Maria Aiello. Unmasking conversational bias in ai multiagent systems.arXiv preprint arXiv:2501.14844, 2025. URLhttps://arxiv.org/abs/2501.14844
2025 arXiv
-
[16]
Xuan Long Do, Kenji Kawaguchi, Min-Yen Kan, and Nancy F. Chen. Aligning large language models with human opinions through persona selection and value–belief–norm reasoning. InProceedings of the 31st Inter- national Conference on Computational Linguistics (COLING), 2025. URLhtt...
2025 arXiv
-
[17]
Large means left: Political bias in large language models increases with their number of parameters.arXiv preprint arXiv:2505.04393, 2025
David Exler, Mark Schutera, Markus Reischl, and Luca Rettenberger. Large means left: Political bias in large language models increases with their number of parameters.arXiv preprint arXiv:2505.04393, 2025. URL https://arxiv.org/abs/2505.04393
2025 arXiv
-
[18]
Only a little to the left: A theory- grounded measure of political bias in large language models
Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz, and David Garcia. Only a little to the left: A theory- grounded measure of political bias in large language models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025. URL...
2025 arXiv
-
[19]
Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E
Jillian Fisher, Ruth E. Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E. Roberts, Jennifer Pan, Dawn Song, and Yejin Choi. Political neutrality in AI is impossible — but here is how to approximate it.arXiv preprint ...
2025 arXiv
-
[20]
Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G
Kobi Hackenburg, Ben M. Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G. Rand, and Christopher Summerfield. The levers of political persuasion with conversational artificial intelligence.Science, 2025. doi: 10.1126/science.aea3884
2025 doi
-
[21]
Kirk A. Hawkins. Is ch ´avez populist? measuring populist discourse in comparative perspective.Comparative Political Studies, 42(8):1040–1067, 2009. doi: 10.1177/0010414009331721
2009 doi
-
[22]
Power to the parties: Cohesion and competition in the eu- ropean parliament, 1979–2001.British Journal of Political Science, 35(2):209–234, 2005
Simon Hix, Abdul Noury, and G ´erard Roland. Power to the parties: Cohesion and competition in the eu- ropean parliament, 1979–2001.British Journal of Political Science, 35(2):209–234, 2005. doi: 10.1017/ S0007123405000128
1979
-
[23]
Changes in party identity: Evidence from party manifestos.Party Politics, 1(2):171–196, 1995
Kenneth Janda, Robert Harmel, Christine Edens, and Patricia Goff. Changes in party identity: Evidence from party manifestos.Party Politics, 1(2):171–196, 1995. doi: 10.1177/1354068895001002001
1995 doi
-
[24]
Kendall and B
Maurice G. Kendall and B. Babington Smith. The problem ofmrankings.The Annals of Mathematical Statistics, 10(3):275–287, 1939. doi: 10.1214/aoms/1177732186
1939
-
[25]
Hofferbert, and Ian Budge.Parties, Policies, and Democracy
Hans-Dieter Klingemann, Richard I. Hofferbert, and Ian Budge.Parties, Policies, and Democracy. Westview Press, Boulder, CO, 1994
1994
-
[26]
White, Adam J
Hause Lin, Gabriela Czarnek, Benjamin Lewis, Joseph P. White, Adam J. Berinsky, Thomas Costello, Gordon Pennycook, and David G. Rand. Persuading voters using human–artificial intelligence dialogues.Nature, 2025. URLhttps://www.nature.com/articles/s41586-025-09771-9
2025
-
[27]
Quantifying and alleviating political bias in language models.Artificial Intelligence, 304:103654, 2022
Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, and Soroush V osoughi. Quantifying and alleviating political bias in language models.Artificial Intelligence, 304:103654, 2022. URLhttps://www.sciencedirect.com/ science/article/pii/S0004370221002058
2022
-
[28]
The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models
Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. InFindings of the Association for Computational Linguistics: EMNLP, 2025. URLht...
2025
-
[29]
More human than human: measuring ChatGPT political bias.Public Choice, 198(1):3–23, 2024
Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. More human than human: measuring ChatGPT political bias.Public Choice, 198(1):3–23, 2024. doi: 10.1007/s11127-023-01097-2
2024 doi
-
[30]
Mutz and Byron Reeves
Diana C. Mutz and Byron Reeves. The new videomalaise: Effects of televised incivility on political trust. American Political Science Review, 99(1):1–15, 2005. doi: 10.1017/S0003055405051452
2005 doi
-
[31]
Palgrave Macmillan, Basingstoke, 2011
Elin Naurin.Election Promises, Party Behaviour and Voter Perceptions. Palgrave Macmillan, Basingstoke, 2011. doi: 10.1057/9780230306400
2011 doi
-
[32]
Emergent coordinated behaviors in networked llm agents: Modeling the strategic dynamics of informa- tion operations
Gian Marco Orlando, Jinyi Ye, Valerio La Gatta, Mahdis Saeedi, Vincenzo Moscato, Emilio Ferrara, and Luca Luceri. Emergent coordinated behaviors in networked llm agents: Modeling the strategic dynamics of informa- tion operations. InProceedings of the ACM Web Conference (WWW),...
2026 arXiv
-
[33]
Benjamin I. Page. The theory of political ambiguity.American Political Science Review, 70(3):742–752, 1976. doi: 10.2307/1959865. 14 Who Would You V ote For? Auditing Political Alignment in LLMs: An Italian Case Study
1976 doi
-
[35]
Petrocik
John R. Petrocik. Issue ownership in presidential elections, with a 1980 case study.American Journal of Political Science, 40(3):825–850, 1996. doi: 10.2307/2111797
1980 doi
-
[36]
Hidden persuaders: Llms’ political leaning and their influence on voters
Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. Hidden persuaders: Llms’ political leaning and their influence on voters. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. URLhttps://arxiv.org/abs/2410.24...
2024 arXiv
-
[37]
Stuart A. Rice. The behavior of legislative groups: A method of measurement.Political Science Quarterly, 40 (1):60–72, 1925. doi: 10.2307/2142407
1925 doi
-
[38]
Measuring populism: Comparing two methods of content analysis.West European Politics, 34(6):1272–1283, 2011
Matthijs Rooduijn and Teun Pauwels. Measuring populism: Comparing two methods of content analysis.West European Politics, 34(6):1272–1283, 2011. doi: 10.1080/01402382.2011.616665
2011
-
[39]
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models
Paul R ¨ottger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Sch ¨utze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual M...
2024
-
[40]
Terry J. Royed. Testing the mandate model in britain and the united states: Evidence from the reagan and thatcher eras.British Journal of Political Science, 26(1):45–80, 1996. doi: 10.1017/S0007123400007419
1996 doi
-
[41]
The political preferences of LLMs.PLOS ONE, 19(7):e0306621, 2024
David Rozado. The political preferences of LLMs.PLOS ONE, 19(7):e0306621, 2024. doi: 10.1371/journal. pone.0306621
2024 doi
-
[42]
Kenneth A. Shepsle. The strategy of ambiguity: Uncertainty and electoral competition.American Political Science Review, 66(2):555–568, 1972. doi: 10.2307/1957799
1972 doi
-
[43]
Democratization and linguistic complexity: The effect of franchise extension on parliamentary discourse, 1832–1915.The Journal of Politics, 78(1):120–136, 2016
Arthur Spirling. Democratization and linguistic complexity: The effect of franchise extension on parliamentary discourse, 1832–1915.The Journal of Politics, 78(1):120–136, 2016. doi: 10.1086/683612
1915 doi
-
[44]
The fulfillment of parties’ election pledges: A comparative study on the impact of power sharing.American Journal of Political Science, 61(3):527–542, 2017
Robert Thomson, Terry Royed, Elin Naurin, Joaqu ´ın Art ´es, Rory Costello, Laurenz Ennser-Jedenastik, Mark Ferguson, Petia Kostadinova, Catherine Moury, Franc ¸ois P´etry, and Katrin Praprotnik. The fulfillment of parties’ election pledges: A comparative study on the impact o...
2017 doi
-
[45]
Green, and Semra Sevi
Yamil Velez, Donald P. Green, and Semra Sevi. Chatbot voting advice applications inform but seldom sway young unaligned voters.Proceedings of the National Academy of Sciences (PNAS), 2025. URLhttps://www. pnas.org/doi/10.1073/pnas.2515516122
2025 doi
-
[46]
Manifesto project dataset – codebook, version 2021a
Andrea V olkens, Tobias Burst, Werner Krause, Pola Lehmann, Theres Matthieß, Sven Regel, Lisa Weißen- bach, and Lisa Zehnter. Manifesto project dataset – codebook, version 2021a. Wissenschaftszentrum Berlin f ¨ur Sozialforschung (WZB) / Manifesto Project, 2021. URLhttps://mani...
2021
-
[47]
Persona prompting as a lens on llm social reasoning
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann, Vera Schmitt, and Nils Feldhus. Persona prompting as a lens on llm social reasoning. InProceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics (EACL), ...
2026
-
[48]
statement_program_consistency
Jinyi Ye, Luca Luceri, and Emilio Ferrara. Auditing political exposure bias: Algorithmic amplification on twitter/x during the 2024 u.s. presidential election. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2025. URLhttps://arxi...
2024 arXiv
-
[2024]
URLhttps://arxiv.org/abs/2412.16746
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.