REVIEW 4 major objections 4 minor 22 references
Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper claims that, in 2001–2025, three of four large language models voting on 5,555 adopted UN General Assembly resolutions sit closest to Russia among the UN Security Council's permanent five, and all four sit farthest from the Unite
desk verdict Solid, honest measurement paper whose headline Russia-closest ranking is specification-dependent; the US-farthest result is robust. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a dynamic ordinal ideal-point model of UN voting. For each actor-session, a latent utility Z = beta_v * theta_it + epsilon maps onto support, abstain, or opposition through two item-specific cutpoints; beta_v measures how strongly a resolution discriminates between positions, and theta_it is the actor's location on a single latent dimension interpreted as proximity to the US-led liberal international order. The model uses recurring resolutions to keep the scale comparable across sessions and a dynamic prior to smooth each actor's trajectory over time. The machinery matters because it separates an actor's tendency to support many resolutions (which raw agreement score
What would settle it
Take the 2,104 resolutions on which the US opposed and China and Russia supported, and re-run each model's vote 20 times with temperature 1.0; if the three US-developed models do not consistently support more than 60% of them across draws, the finding is sampling noise. Alternatively, re-estimate the model allowing two latent dimensions; if on a second dimension the same models sit near the US, the one-dimensional 'closest to Russia' claim fails.
Extended reading notes
Core claim
The central discovery is a comparison of revealed geopolitical positions. Treating each model as a respondent to 5,555 divisive, recorded, adopted UN General Assembly resolutions, the paper estimates a latent position for each model on the same one-dimensional scale used for states. In the 2001–2025 period, GPT-5, Claude Sonnet, and Gemini have mean ideal points that place them nearest to Russia among the permanent five, while DeepSeek sits nearest to France; every model is farthest from the United States. The paper also shows the result is not a pure artifact of the estimator: on 2,104 resolutions where the US voted no while China and Russia voted yes, the three models supported the resolut
Load-bearing premise
The load-bearing premise is that a single fixed latent dimension—estimated from ordinal voting patterns with item-specific thresholds—is the correct definition of geopolitical proximity, and that one temperature-1.0 response per resolution is a stable revealed preference rather than a random draw.
Editorial extensions
If this is right
- If the paper is right, the geopolitical orientation of an LLM cannot be inferred from its developer's home country: American systems (GPT-5, Claude Sonnet, Gemini) sit farthest from the US on the recovered dimension, and the Chinese-developed DeepSeek is closest to France, not China.
- Geopolitical 'alignment' is measurement-dependent: support rates, ternary similarity, and latent ideal points can rank the same actors differently (e.g., the three assent-heavy models are nearest to China by raw S-score but nearest to Russia by ideal point), so any single audit statistic can mislead.
- LLMs used as political coders or simulated respondents carry model-specific attitudes that can change research conclusions; replacing one model with another moved support by almost 60 percentage points, so political-science workflows should record exact prompts and identifiers and use multiple models.
- Governments integrating models into policy or diplomatic work cannot infer alignment from hosting or ownership; sovereign AI does not guarantee a model reproduces a national position, and task-specific evaluation against real decision corpora is needed.
- The observed US gap is tied to a recurring mechanism: models endorse broadly framed humanitarian, developmental, self-determination, and arms-control provisions that the US often opposes as part of larger diplomatic packages; this is visible in the resolution texts themselves, not only in estimated latent positions.
Reading between the lines
- A testable extension would be to vary the prompt's national-role instruction: if models were told to act as a US delegate, the modern Russia-closest ranking would probably collapse, which would show how much of the finding is an artifact of the role-neutral 'substantive content' instruction rather than a stable model disposition.
- The one-dimensional scale may hide issue-specific coalitions: a model can be near Russia on the aggregate dimension while diverging sharply on Ukraine, human-rights, or information-security votes; splitting resolutions into issue domains and re-estimating would reveal whether the Russia-proximity is uniform or concentrated.
- Because the agenda contains only adopted resolutions, assent-heavy models may be rewarding institutional norm language rather than indicating geopolitical preference; a parallel audit using failed drafts or proposed-but-rejected resolutions could show whether the US distance shrinks when the texts are not already endorsed by an Assembly majority.
- The S-score/ideal-point divergence (China vs Russia) implies that public 'closest to X' rankings of models are conditional on a measurement theory; reporters and auditors should treat any single-number alignment score as a model-dependent estimate, not a fact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper adapts the Bailey, Strezhnev, and Voeten (BSV) dynamic ordinal ideal-point model to measure the geopolitical preferences expressed by four LLMs (GPT-5, Claude Sonnet, Gemini, and DeepSeek) treated as respondents to 5,555 divisive, recorded, adopted UN General Assembly resolutions from regular sessions 1–80. Each model was prompted with the full resolution text under a single no-role instruction and asked to return Support, Abstain, or Oppose. The paper reports that in the twenty-first century GPT-5, Claude Sonnet, and Gemini are, in the ideal-point space, closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. It also documents large differences in raw response rates and analyzes 2,104 resolutions where the United States opposed and China and Russia/USSR supported. The paper explicitly contrasts the ideal-point ranking with a raw ternary S-score, which ranks China first for the three assent-heavy models in the same period, and provides a replication of the BSV country estimates.
Significance. If the central finding is robust, the paper makes a valuable methodological and substantive contribution: it imports an established international-relations measurement framework into LLM auditing, and it provides concrete evidence that a model's expressed geopolitical position can differ markedly from its developer's home country. The paper has notable strengths: it uses an established estimator rather than an ad hoc survey; it reports a transparent agreement measure alongside the latent-space estimate; it includes a BSV replication with high correlation to the original country estimates (0.987); it discusses measurement dependence explicitly; and it lists limitations candidly. These features make the paper a useful template for future LLM alignment audits, even where the headline ranking is contested.
major comments (4)
- [§3.2, Figures 3 and 4] The central claim in the abstract and Section 3.2 that GPT-5, Claude Sonnet, and Gemini are 'closest to Russia' among the P5 in 2001–2025 is not robust across the paper's own measures. The session-level S-score in Figure 4 ranks China first for all three models in the same period, and the paper's text acknowledges this divergence. The ideal-point estimator is one defensible definition of proximity, but the abstract states the Russia-closest result without the 'ideal-point' qualifier. Because raw agreement and latent proximity are both reported and can rank actors differently, the headline should either be qualified or accompanied by a stronger argument that the latent dimension is the authoritative measure for the question asked.
- [§3.1–3.2, Figure 3, Table 4] Figure 3 reports mean absolute ideal-point distances as point estimates with no uncertainty quantification for the distances themselves. The paper shows 90% posterior intervals for model ideal points in Figure 2 and convergence diagnostics in Table 4, but these are not propagated into the distances or into the ranking of nearest P5 member. Without credible intervals on the distances or the posterior probability that Russia is the nearest P5 member, the ordering could be within estimation noise. The differences between Russia and the next nearest P5 member are not reported, so the reader cannot assess whether the ranking is decisive. The authors should report posterior intervals for distances and, ideally, the posterior probability of each P5 member being nearest.
- [§2.2, §3.1] The analysis treats exactly one temperature-1.0 response per model per resolution as a stable revealed preference. This is a load-bearing assumption: at temperature 1.0, LLM outputs are stochastic, and a single draw per item introduces sampling variance that is never quantified. If repeated draws change support on even a few hundred discriminating items, the nearest-P5 ordering could flip, especially given the divergence between ideal-point and raw agreement rankings. The paper should either include repeated sampling and report the distribution of rankings, or explicitly acknowledge that the result applies to one fixed draw and justify why that is sufficient for the substantive conclusion.
- [Abstract, §3.2] The abstract statement 'all four are farthest from the United States' is an overstatement. For DeepSeek, the modern S-score is lowest with China, not the United States, as the text in §3.2 itself notes ('DeepSeek is the exception'). The ideal-point estimate finds DeepSeek's largest latent distance is from the United States, but the unqualified wording in the abstract obscures the measurement dependence that the paper elsewhere emphasizes. The claim should be qualified to the latent ideal-point comparison, or the S-score exception should be reflected in the summary.
minor comments (4)
- [Table 4] The column layout is confusing: 'Median' and 'ˆR≤1.05' appear to share a column heading. Clarify whether the reported median is the median ideal point or the median R-hat, and label the diagnostic column explicitly.
- [§3.4] The BSV replication uses 49 model sessions and 42 sessions with all four models and all five P5 members, but the text does not explain why DeepSeek has fewer sessions meeting the R-hat threshold. A brief note would help readers interpret the replication's coverage.
- [§4] The limitations paragraph is thorough, but the 'sixth' limitation—that the experiment recovers expressed choices, not internal beliefs—could be moved earlier, since it qualifies how the phrase 'preference' should be read throughout the paper.
- [Appendix A] The prompt includes a line 'Respond with exactly one word' and the model is expected to output one of three tokens. Reporting the validation failure rate (if any) would strengthen confidence that the responses are not contaminated by off-protocol text.
Circularity Check
No circularity: ideal points are estimated from observed model votes, not fitted to the Russia-closest conclusion.
full rationale
The paper's derivation chain is self-contained: model responses to 5,555 UNGA resolutions are observed under a fixed prompt (§2.2), and the BSV-style dynamic ordinal ideal-point model (§3.1) is estimated from those responses together with country votes. The Russia-closest ranking is a posterior summary of estimated distances, not a quantity used to define or fit the model. The paper explicitly reports that the raw ternary S-score ranks China first for the same three models (§3.2, Figure 4), so it does not present the ideal-point result as the only possible measure or hide the divergence. The core method is imported from Bailey et al. (2017), an external source, and the robustness check in §3.4 re-estimates using the original BSV country data and sampler; no load-bearing claim rests on a self-citation. Potential concerns such as single temperature-1.0 responses per resolution, absence of posterior intervals on the P5 distance rankings, and the one-dimensional latent specification are robustness or identification issues, not circular reasoning. There is no equation in which an input is defined in terms of an output, no fitted parameter renamed as a prediction, and no self-citation used to forbid alternative interpretations. The central empirical findings therefore stand independently of the paper's own assumptions in the sense required for a circularity finding.
Assumptions & free parameters
free parameters (6)
- Item discrimination parameters β_v =
estimated jointly
- Item-specific cutpoints (two per item) =
estimated
- Actor-session ideal points θ_it =
estimated posteriors
- Smoothing parameter =
0.5
- Latent space dimension =
1
- Analysis universe =
5,555 adopted, divisive, regular-session resolutions
assumptions (6)
- domain assumption Ordinal item-response assumptions: normal errors, conditional independence, monotone cutpoints.
- domain assumption A unidimensional latent scale is adequate and comparable across states and LLMs.
- ad hoc to paper A single English no-role prompt with one temperature-1.0 response measures a stable model revealed preference.
- domain assumption Automated text matching reliably identifies recurring resolutions and bridges the intertemporal scale.
- domain assumption The UN Digital Library bulk voting data version 5 is accurate and complete.
- domain assumption The sign orientation (US positive, Russia negative) and the interpretation as a US-led liberal-order dimension.
Cite this review
Pith. "Pith review of Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data." pith.science (2026). https://pith.science/paper/OZUMQMQK
@misc{pith2026260725526,
author = {Pith},
title = {Pith review of: Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/OZUMQMQK}},
note = {Machine review of arXiv:2607.25526}
}
read the original abstract
How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings show that a model's expressed geopolitical position can differ markedly from that of its developer's home country, especially in international politics, where state actions can diverge from the stated principles prevalent in the texts on which models are trained.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
P., Busby, E
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis , 31(3):337--351
2023
-
[2]
A., Strezhnev, A., and Voeten, E
Bailey, M. A., Strezhnev, A., and Voeten, E. (2017). Estimating dynamic state preferences from united nations voting data. Journal of Conflict Resolution , 61(2):430--456
2017
-
[3]
Ball, M. M. (1951). Bloc voting in the general assembly. International Organization , 5(1):3--31
1951
-
[4]
Bladon, S. and Bent, B. (2026). It's the humans, not the data: Geopolitical bias in llms originates in post-training, amplified by the language of the prompt. arXiv preprint arXiv:2605.23825
arXiv 2026
-
[5]
and Innerarity, D
Bosoer, L. and Innerarity, D. (2025). Unpacking ai sovereignty. STG Policy Papers 2025/18, European University Institute, Florence School of Transnational Governance
2025
-
[6]
Buyl, M., Rogiers, A., Noels, S., Bied, G., Dominguez-Catena, I., Heiter, E., Johary, I., Mara, A.-C., Romero, R., Lijffijt, J., and De Bie, T. (2026). Large language models reflect the ideology of their creators. npj Artificial Intelligence , 2:7
2026
-
[7]
Faisal, F. and Anastasopoulos, A. (2023). Geographic and geopolitical biases of language models. arXiv preprint arXiv:2212.10408
arXiv 2023
-
[8]
Gartzke, E. (1998). Kant we all just get along? opportunity, willingness, and the origins of the democratic peace. American Journal of Political Science , 42(1):1--27
1998
Show all 22 references
-
[9]
Lijphart, A. (1963). The analysis of bloc voting in the general assembly: A critique and a proposal. American Political Science Review , 57(4):902--917
1963
-
[10]
Moon, B. E. (1985). Consensus or compliance? foreign-policy change and external dependence. International Organization , 39(2):297--329
1985
-
[11]
Russett, B. M. (1966). Discovering voting groups in the united nations. American Political Science Review , 60:327--339
1966
-
[12]
Y., Loukachevitch, N., Panchenko, A., and Tutubalina, E
Salnikov, M., Korzh, D., Lazichny, I., Karimov, E., Iudin, A., Oseledets, I., Rogov, O. Y., Loukachevitch, N., Panchenko, A., and Tutubalina, E. (2025). Geopolitical biases in llms: What are the ``good'' and the ``bad'' countries according to contemporary language models. arXi...
2025 arXiv
-
[13]
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023). Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning , volume 202, pages 29971--30004
2023
-
[14]
Signorino, C. S. and Ritter, J. M. (1999). Tau-b or not tau-b: Measuring the similarity of foreign policy positions. International Studies Quarterly , 43(1):115--144
1999
-
[15]
and Voeten, E
Strezhnev, A. and Voeten, E. (2013). United nations general assembly voting data. Version 7.1
2013
-
[16]
S., and Kizilcec, R
Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus , 3(9):pgae346
2024
-
[17]
Ai opportunities action plan: Government response
UK Department for Science, Innovation and Technology (2025). Ai opportunities action plan: Government response. Technical report, Government of the United Kingdom
2025
-
[18]
United nations digital library bulk voting data
United Nations (2026). United nations digital library bulk voting data. Version 5, dated 6 February 2026
2026
-
[19]
Vengroff, R. (1976). Instability and foreign policy behavior: Black africa in the u.n. American Journal of Political Science , 20(3):425--438
1976
-
[20]
Voeten, E. (2000). Clashes in the assembly. International Organization , 54(2):185--215
2000
-
[21]
Walker, C. P. and Timoneda, J. C. (2025). Is chatgpt conservative or liberal? a novel approach to assess ideological stances and biases in generative llms. Political Science Research and Methods , pages 1--15
2025
-
[22]
B., Faulborn, M., and Garc \'i a, D
Weidmann, N. B., Faulborn, M., and Garc \'i a, D. (2026). Large language models are democracy coders with attitudes. PS: Political Science & Politics , 59(1):17--23
2026
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.