Pith. sign in

REVIEW 4 major objections 4 minor 22 references

Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data

T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper claims that, in 2001–2025, three of four large language models voting on 5,555 adopted UN General Assembly resolutions sit closest to Russia among the UN Security Council's permanent five, and all four sit farthest from the Unite

desk verdict Solid, honest measurement paper whose headline Russia-closest ranking is specification-dependent; the US-farthest result is robust. read the letter →

arxiv 2607.25526 v1 pith:OZUMQMQK submitted 2026-07-28 cs.CY

classification cs.CY
keywords largelanguagemodelsgeopoliticalalignmentUnitedNationsvotingidealpointsforeign-policypreferencesAIgovernancemeasurementstrategyUNGeneralAssembly
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper takes a method developed in international relations for recovering hidden preferences from UN voting—a dynamic ideal-point model that maps support, abstention, and opposition onto a single latent scale—and applies it to four large language models (GPT-5, Claude Sonnet, Gemini, and DeepSeek), which each voted on the full texts of 5,555 divisive, adopted UN General Assembly resolutions from 1946 to 2025. The paper's central finding is that, in the modern period (2001–2025), GPT-5, Claude Sonnet, and Gemini are closest among the five permanent Security Council members to Russia, DeepSeek is closest to France, and all four are farthest from the United States. On the 2,104 resolutions the US opposed while China and Russia supported, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The conclusion is that a model's expressed geopolitical position can differ sharply from its developer's home country, and that the measured alignment depends on which statistical definition of proximity is used: raw agreement, a ternary similarity score, and a latent ideal point can rank the same actors differently. The author argues this cautions against treating any single alignment score—or any model provider's national origin—as a reliable guide to a model's geopolitical orientation.

What carries the argument

The central mechanism is a dynamic ordinal ideal-point model of UN voting. For each actor-session, a latent utility Z = beta_v * theta_it + epsilon maps onto support, abstain, or opposition through two item-specific cutpoints; beta_v measures how strongly a resolution discriminates between positions, and theta_it is the actor's location on a single latent dimension interpreted as proximity to the US-led liberal international order. The model uses recurring resolutions to keep the scale comparable across sessions and a dynamic prior to smooth each actor's trajectory over time. The machinery matters because it separates an actor's tendency to support many resolutions (which raw agreement score

What would settle it

Take the 2,104 resolutions on which the US opposed and China and Russia supported, and re-run each model's vote 20 times with temperature 1.0; if the three US-developed models do not consistently support more than 60% of them across draws, the finding is sampling noise. Alternatively, re-estimate the model allowing two latent dimensions; if on a second dimension the same models sit near the US, the one-dimensional 'closest to Russia' claim fails.

Watch

Extended reading notes

Core claim

The central discovery is a comparison of revealed geopolitical positions. Treating each model as a respondent to 5,555 divisive, recorded, adopted UN General Assembly resolutions, the paper estimates a latent position for each model on the same one-dimensional scale used for states. In the 2001–2025 period, GPT-5, Claude Sonnet, and Gemini have mean ideal points that place them nearest to Russia among the permanent five, while DeepSeek sits nearest to France; every model is farthest from the United States. The paper also shows the result is not a pure artifact of the estimator: on 2,104 resolutions where the US voted no while China and Russia voted yes, the three models supported the resolut

Load-bearing premise

The load-bearing premise is that a single fixed latent dimension—estimated from ordinal voting patterns with item-specific thresholds—is the correct definition of geopolitical proximity, and that one temperature-1.0 response per resolution is a stable revealed preference rather than a random draw.

Editorial extensions

If this is right

  • If the paper is right, the geopolitical orientation of an LLM cannot be inferred from its developer's home country: American systems (GPT-5, Claude Sonnet, Gemini) sit farthest from the US on the recovered dimension, and the Chinese-developed DeepSeek is closest to France, not China.
  • Geopolitical 'alignment' is measurement-dependent: support rates, ternary similarity, and latent ideal points can rank the same actors differently (e.g., the three assent-heavy models are nearest to China by raw S-score but nearest to Russia by ideal point), so any single audit statistic can mislead.
  • LLMs used as political coders or simulated respondents carry model-specific attitudes that can change research conclusions; replacing one model with another moved support by almost 60 percentage points, so political-science workflows should record exact prompts and identifiers and use multiple models.
  • Governments integrating models into policy or diplomatic work cannot infer alignment from hosting or ownership; sovereign AI does not guarantee a model reproduces a national position, and task-specific evaluation against real decision corpora is needed.
  • The observed US gap is tied to a recurring mechanism: models endorse broadly framed humanitarian, developmental, self-determination, and arms-control provisions that the US often opposes as part of larger diplomatic packages; this is visible in the resolution texts themselves, not only in estimated latent positions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be to vary the prompt's national-role instruction: if models were told to act as a US delegate, the modern Russia-closest ranking would probably collapse, which would show how much of the finding is an artifact of the role-neutral 'substantive content' instruction rather than a stable model disposition.
  • The one-dimensional scale may hide issue-specific coalitions: a model can be near Russia on the aggregate dimension while diverging sharply on Ukraine, human-rights, or information-security votes; splitting resolutions into issue domains and re-estimating would reveal whether the Russia-proximity is uniform or concentrated.
  • Because the agenda contains only adopted resolutions, assent-heavy models may be rewarding institutional norm language rather than indicating geopolitical preference; a parallel audit using failed drafts or proposed-but-rejected resolutions could show whether the US distance shrinks when the texts are not already endorsed by an Assembly majority.
  • The S-score/ideal-point divergence (China vs Russia) implies that public 'closest to X' rankings of models are conditional on a measurement theory; reporters and auditors should treat any single-number alignment score as a model-dependent estimate, not a fact.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper adapts the Bailey, Strezhnev, and Voeten (BSV) dynamic ordinal ideal-point model to measure the geopolitical preferences expressed by four LLMs (GPT-5, Claude Sonnet, Gemini, and DeepSeek) treated as respondents to 5,555 divisive, recorded, adopted UN General Assembly resolutions from regular sessions 1–80. Each model was prompted with the full resolution text under a single no-role instruction and asked to return Support, Abstain, or Oppose. The paper reports that in the twenty-first century GPT-5, Claude Sonnet, and Gemini are, in the ideal-point space, closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. It also documents large differences in raw response rates and analyzes 2,104 resolutions where the United States opposed and China and Russia/USSR supported. The paper explicitly contrasts the ideal-point ranking with a raw ternary S-score, which ranks China first for the three assent-heavy models in the same period, and provides a replication of the BSV country estimates.

Significance. If the central finding is robust, the paper makes a valuable methodological and substantive contribution: it imports an established international-relations measurement framework into LLM auditing, and it provides concrete evidence that a model's expressed geopolitical position can differ markedly from its developer's home country. The paper has notable strengths: it uses an established estimator rather than an ad hoc survey; it reports a transparent agreement measure alongside the latent-space estimate; it includes a BSV replication with high correlation to the original country estimates (0.987); it discusses measurement dependence explicitly; and it lists limitations candidly. These features make the paper a useful template for future LLM alignment audits, even where the headline ranking is contested.

major comments (4)
  1. [§3.2, Figures 3 and 4] The central claim in the abstract and Section 3.2 that GPT-5, Claude Sonnet, and Gemini are 'closest to Russia' among the P5 in 2001–2025 is not robust across the paper's own measures. The session-level S-score in Figure 4 ranks China first for all three models in the same period, and the paper's text acknowledges this divergence. The ideal-point estimator is one defensible definition of proximity, but the abstract states the Russia-closest result without the 'ideal-point' qualifier. Because raw agreement and latent proximity are both reported and can rank actors differently, the headline should either be qualified or accompanied by a stronger argument that the latent dimension is the authoritative measure for the question asked.
  2. [§3.1–3.2, Figure 3, Table 4] Figure 3 reports mean absolute ideal-point distances as point estimates with no uncertainty quantification for the distances themselves. The paper shows 90% posterior intervals for model ideal points in Figure 2 and convergence diagnostics in Table 4, but these are not propagated into the distances or into the ranking of nearest P5 member. Without credible intervals on the distances or the posterior probability that Russia is the nearest P5 member, the ordering could be within estimation noise. The differences between Russia and the next nearest P5 member are not reported, so the reader cannot assess whether the ranking is decisive. The authors should report posterior intervals for distances and, ideally, the posterior probability of each P5 member being nearest.
  3. [§2.2, §3.1] The analysis treats exactly one temperature-1.0 response per model per resolution as a stable revealed preference. This is a load-bearing assumption: at temperature 1.0, LLM outputs are stochastic, and a single draw per item introduces sampling variance that is never quantified. If repeated draws change support on even a few hundred discriminating items, the nearest-P5 ordering could flip, especially given the divergence between ideal-point and raw agreement rankings. The paper should either include repeated sampling and report the distribution of rankings, or explicitly acknowledge that the result applies to one fixed draw and justify why that is sufficient for the substantive conclusion.
  4. [Abstract, §3.2] The abstract statement 'all four are farthest from the United States' is an overstatement. For DeepSeek, the modern S-score is lowest with China, not the United States, as the text in §3.2 itself notes ('DeepSeek is the exception'). The ideal-point estimate finds DeepSeek's largest latent distance is from the United States, but the unqualified wording in the abstract obscures the measurement dependence that the paper elsewhere emphasizes. The claim should be qualified to the latent ideal-point comparison, or the S-score exception should be reflected in the summary.
minor comments (4)
  1. [Table 4] The column layout is confusing: 'Median' and 'ˆR≤1.05' appear to share a column heading. Clarify whether the reported median is the median ideal point or the median R-hat, and label the diagnostic column explicitly.
  2. [§3.4] The BSV replication uses 49 model sessions and 42 sessions with all four models and all five P5 members, but the text does not explain why DeepSeek has fewer sessions meeting the R-hat threshold. A brief note would help readers interpret the replication's coverage.
  3. [§4] The limitations paragraph is thorough, but the 'sixth' limitation—that the experiment recovers expressed choices, not internal beliefs—could be moved earlier, since it qualifies how the phrase 'preference' should be read throughout the paper.
  4. [Appendix A] The prompt includes a line 'Respond with exactly one word' and the model is expected to output one of three tokens. Reporting the validation failure rate (if any) would strengthen confidence that the responses are not contaminated by off-protocol text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: ideal points are estimated from observed model votes, not fitted to the Russia-closest conclusion.

full rationale

The paper's derivation chain is self-contained: model responses to 5,555 UNGA resolutions are observed under a fixed prompt (§2.2), and the BSV-style dynamic ordinal ideal-point model (§3.1) is estimated from those responses together with country votes. The Russia-closest ranking is a posterior summary of estimated distances, not a quantity used to define or fit the model. The paper explicitly reports that the raw ternary S-score ranks China first for the same three models (§3.2, Figure 4), so it does not present the ideal-point result as the only possible measure or hide the divergence. The core method is imported from Bailey et al. (2017), an external source, and the robustness check in §3.4 re-estimates using the original BSV country data and sampler; no load-bearing claim rests on a self-citation. Potential concerns such as single temperature-1.0 responses per resolution, absence of posterior intervals on the P5 distance rankings, and the one-dimensional latent specification are robustness or identification issues, not circular reasoning. There is no equation in which an input is defined in terms of an output, no fitted parameter renamed as a prediction, and no self-citation used to forbid alternative interpretations. The central empirical findings therefore stand independently of the paper's own assumptions in the sense required for a circularity finding.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The paper imports an established Bayesian ordinal IRT model and a public UN voting archive. It introduces no new physical or conceptual entity; its main unverified imports are the model's statistical assumptions and the treatment of single LLM responses as stable choices.

free parameters (6)
  • Item discrimination parameters β_v = estimated jointly
    Each resolution's discrimination is estimated from the country and model vote matrix; the LLM placements depend on these values.
  • Item-specific cutpoints (two per item) = estimated
    Map latent utility into Support/Abstain/Oppose; estimated with the model rather than fixed independently.
  • Actor-session ideal points θ_it = estimated posteriors
    These include the LLM ideal points that are the paper's main output; they are parameters of the fitted model.
  • Smoothing parameter = 0.5
    Tuning constant in the dynamic prior over actor-session ideal points; no sensitivity analysis is reported.
  • Latent space dimension = 1
    The model assumes one geopolitical dimension; adding a second dimension could change the P5 rankings.
  • Analysis universe = 5,555 adopted, divisive, regular-session resolutions
    The criterion for divisive (at least two observed vote categories) and the exclusion of non-adopted and unanimous items shape the support rates and the estimated positions.
assumptions (6)
  • domain assumption Ordinal item-response assumptions: normal errors, conditional independence, monotone cutpoints.
    Stated in §3.1, Eq. 2, inherited from Bailey et al. (2017). The paper does not independently validate these assumptions for LLM respondents.
  • domain assumption A unidimensional latent scale is adequate and comparable across states and LLMs.
    Invoked throughout §3.2; the ranking is defined by distance on this single dimension.
  • ad hoc to paper A single English no-role prompt with one temperature-1.0 response measures a stable model revealed preference.
    The design in §2.2 uses one response per item with no repeated sampling; the paper itself notes this limitation.
  • domain assumption Automated text matching reliably identifies recurring resolutions and bridges the intertemporal scale.
    The dynamic model relies on 1,360 recurring-resolution matches (§3.1); the paper notes these are not manually validated to the standard of the original BSV bridges.
  • domain assumption The UN Digital Library bulk voting data version 5 is accurate and complete.
    All country votes are taken from this single public archive (§2.1).
  • domain assumption The sign orientation (US positive, Russia negative) and the interpretation as a US-led liberal-order dimension.
    The paper follows BSV's convention and interprets the resulting dimension conventionally (§3.1), but this interpretation is not directly tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data." pith.science (2026). https://pith.science/paper/OZUMQMQK

@misc{pith2026260725526,
  author       = {Pith},
  title        = {Pith review of: Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OZUMQMQK}},
  note         = {Machine review of arXiv:2607.25526}
}
read the original abstract

How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings show that a model's expressed geopolitical position can differ markedly from that of its developer's home country, especially in international politics, where state actions can diverge from the stated principles prevalent in the texts on which models are trained.

Figures

Figures reproduced from arXiv: 2607.25526 by the authors.

Figure 1
Figure 1. Model vote distributions Note: Shares pool the 5,555 divisive, recorded, adopted resolutions in regular UNGA sessions 1–80. Each model casts one valid vote per item. Session 80 is provisional. These totals already complicate any single ranking of “alignment.” In pooled ternary agreement, GPT-5’s closest country with at least 500 common votes is Timor-Leste (S = 0.982), followed by Seychelles and Cabo Verde. Claude S… view at source ↗
Figure 2
Figure 2. Model and permanent-five ideal points over time Note: Posterior means from the same BSV-style ordinal model. Gray bands are model 90% posterior intervals. Open circles mark model-session estimates with R >b 1.05. The dashed vertical line separates 2000 from 2001. Session 19 (opening year 1964) is absent; session 80 (2025) is provisional. The Soviet series represents the Russian P5 seat through session 46. 8 [PITH_F… view at source ↗
Figure 3
Figure 3. Model proximity to permanent-five members by century Note: Bars are mean session-level absolute ideal-point distances for 2001–2025; diamonds are means for 1946–2000. Lower values mean greater latent-space proximity. Periods contain 25 and 54 observed sessions, respectively. Identification diagnostics from [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Model agreement with permanent-five members by century Note: Bars are mean session-level ternary S-scores for 2001–2025; diamonds are means for 1946–2000. Higher values mean greater observed categorical agreement. Periods contain 25 and 54 regular sessions, respectivel…
Figure 5
Figure 5. Figure 5: Model proximity to permanent-five members in the BSV replication Note: Posterior means come from the original BSV paper-era country universe and sampler, augmented by four partially observed model respondents. Bars are mean session-level absolute distances for 2001–201…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

22 extracted references · 3 linked inside Pith

  1. [1]

    P., Busby, E

    Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis , 31(3):337--351

  2. [2]

    A., Strezhnev, A., and Voeten, E

    Bailey, M. A., Strezhnev, A., and Voeten, E. (2017). Estimating dynamic state preferences from united nations voting data. Journal of Conflict Resolution , 61(2):430--456

  3. [3]

    Ball, M. M. (1951). Bloc voting in the general assembly. International Organization , 5(1):3--31

  4. [4]

    and Bent, B

    Bladon, S. and Bent, B. (2026). It's the humans, not the data: Geopolitical bias in llms originates in post-training, amplified by the language of the prompt. arXiv preprint arXiv:2605.23825

  5. [5]

    and Innerarity, D

    Bosoer, L. and Innerarity, D. (2025). Unpacking ai sovereignty. STG Policy Papers 2025/18, European University Institute, Florence School of Transnational Governance

  6. [6]

    Buyl, M., Rogiers, A., Noels, S., Bied, G., Dominguez-Catena, I., Heiter, E., Johary, I., Mara, A.-C., Romero, R., Lijffijt, J., and De Bie, T. (2026). Large language models reflect the ideology of their creators. npj Artificial Intelligence , 2:7

  7. [7]

    and Anastasopoulos, A

    Faisal, F. and Anastasopoulos, A. (2023). Geographic and geopolitical biases of language models. arXiv preprint arXiv:2212.10408

  8. [8]

    Gartzke, E. (1998). Kant we all just get along? opportunity, willingness, and the origins of the democratic peace. American Journal of Political Science , 42(1):1--27

Show all 22 references
  1. [9]

    Lijphart, A. (1963). The analysis of bloc voting in the general assembly: A critique and a proposal. American Political Science Review , 57(4):902--917

  2. [10]

    Moon, B. E. (1985). Consensus or compliance? foreign-policy change and external dependence. International Organization , 39(2):297--329

  3. [11]

    Russett, B. M. (1966). Discovering voting groups in the united nations. American Political Science Review , 60:327--339

  4. [12]

    Y., Loukachevitch, N., Panchenko, A., and Tutubalina, E

    Salnikov, M., Korzh, D., Lazichny, I., Karimov, E., Iudin, A., Oseledets, I., Rogov, O. Y., Loukachevitch, N., Panchenko, A., and Tutubalina, E. (2025). Geopolitical biases in llms: What are the ``good'' and the ``bad'' countries according to contemporary language models. arXi...

  5. [13]

    Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023). Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning , volume 202, pages 29971--30004

  6. [14]

    Signorino, C. S. and Ritter, J. M. (1999). Tau-b or not tau-b: Measuring the similarity of foreign policy positions. International Studies Quarterly , 43(1):115--144

  7. [15]

    and Voeten, E

    Strezhnev, A. and Voeten, E. (2013). United nations general assembly voting data. Version 7.1

  8. [16]

    S., and Kizilcec, R

    Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus , 3(9):pgae346

  9. [17]

    Ai opportunities action plan: Government response

    UK Department for Science, Innovation and Technology (2025). Ai opportunities action plan: Government response. Technical report, Government of the United Kingdom

  10. [18]

    United nations digital library bulk voting data

    United Nations (2026). United nations digital library bulk voting data. Version 5, dated 6 February 2026

  11. [19]

    Vengroff, R. (1976). Instability and foreign policy behavior: Black africa in the u.n. American Journal of Political Science , 20(3):425--438

  12. [20]

    Voeten, E. (2000). Clashes in the assembly. International Organization , 54(2):185--215

  13. [21]

    Walker, C. P. and Timoneda, J. C. (2025). Is chatgpt conservative or liberal? a novel approach to assess ideological stances and biases in generative llms. Political Science Research and Methods , pages 1--15

  14. [22]

    B., Faulborn, M., and Garc \'i a, D

    Weidmann, N. B., Faulborn, M., and Garc \'i a, D. (2026). Large language models are democracy coders with attitudes. PS: Political Science & Politics , 59(1):17--23

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.