{"id":"28143e37-8fd9-43bf-b12d-5c7cf69abd5d","arxiv_id":"2607.25526","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"On a UN-vote ideal-point scale, GPT-5, Claude Sonnet, and Gemini are closer to Russia than to the US among P5 states, DeepSeek is closest to France, and all four are farthest from the US.","lead":"Researchers asked four large language models to vote on 5,555 UN General Assembly resolutions, then used a statistical method from international relations to locate each model's hidden political position. The models ended up far from the United States: GPT-5, Claude Sonnet, and Gemini sit closest to Russia's UN voting record in the modern period, while DeepSeek sits closest to France.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'closest to Russia' ranking is reported without uncertainty quantification; posterior variance and single-response stochasticity could reorder P5 distances.","rationale":"The reader's weakest assumption identifies the latent-dimension definition and single-response stability as load-bearing. My concern sharpens that to a precise missing quantity: uncertainty quantification for the P5 distance ranking. The reader already assigned CONDITIONAL, and my critique supports that verdict without altering it. The paper is transparent about its design and limitations, and the US-farthest result appears robust across measures, so the concern does not justify rejection or a fundamental reclassification. However, the headline Russia-closest claim is not backed by any evidence of statistical significance or sampling robustness, so the paper should require repeated sampling and interval-based ranking before acceptance rather than taking the point estimates at face value. I partially agree with the reader: we share the underlying worry, but my emphasis is on the absent posterior inference rather than on the conceptual validity of the dimension itself.","tokens_in":11479,"tokens_out":3316,"duration_ms":40104,"concrete_test":"Re-run the estimation with 20 independent temperature-1.0 responses per model per resolution (or use 20 random seeds for the full pipeline). For each model, compute the posterior probability that the mean absolute distance to Russia over 2001–2025 is strictly less than the distance to China and France, and report 90% posterior intervals for the pairwise distance differences. If any probability falls below 0.95 or an interval includes zero, the 'closest to Russia' headline should be presented as provisional or conditional on the single-response draw.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that GPT-5, Claude Sonnet, and Gemini are closest to Russia among the P5 in 2001–2025—rests on point estimates of mean absolute ideal-point distances (Figure 3). The paper provides no posterior credible intervals for these distances, no probability that Russia is the nearest P5 member, and no propagation of the sampling variability introduced by drawing exactly one temperature-1.0 response per model per resolution (§2.2). The raw S-score already ranks China first for the same three models (§3.2, Figure 4), so the conclusion hinges entirely on the latent ideal-point specification. Within that specification, however, the distance ranking may be within posterior noise: GPT-5's modern mean ideal point (-0.167) sits close to several P5 positions, and the reported convergence diagnostics (Table 4) assess MCMC mixing only, not the distribution of model responses or the posterior distribution of distances. If repeated draws shift support on even a few hundred items, the nearest-P5 ordering could flip. The paper itself notes early model sessions are weakly identified, but the twenty-first-century distances are presented without analogous uncertainty measures. The load-bearing premise—that the estimated latent dimension is the right definition of proximity and that one response captures a stable revealed preference—remains quantitatively untested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper adapts the Bailey, Strezhnev, and Voeten (BSV) dynamic ordinal ideal-point model to measure the geopolitical preferences expressed by four LLMs (GPT-5, Claude Sonnet, Gemini, and DeepSeek) treated as respondents to 5,555 divisive, recorded, adopted UN General Assembly resolutions from regular sessions 1–80. Each model was prompted with the full resolution text under a single no-role instruction and asked to return Support, Abstain, or Oppose. The paper reports that in the twenty-first century GPT-5, Claude Sonnet, and Gemini are, in the ideal-point space, closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. It also documents large differences in raw response rates and analyzes 2,104 resolutions where the United States opposed and China and Russia/USSR supported. The paper explicitly contrasts the ideal-point ranking with a raw ternary S-score, which ranks China first for the three assent-heavy models in the same period, and provides a replication of the BSV country estimates.","tokens_in":11771,"tokens_out":3638,"duration_ms":42837,"significance":"If the central finding is robust, the paper makes a valuable methodological and substantive contribution: it imports an established international-relations measurement framework into LLM auditing, and it provides concrete evidence that a model's expressed geopolitical position can differ markedly from its developer's home country. The paper has notable strengths: it uses an established estimator rather than an ad hoc survey; it reports a transparent agreement measure alongside the latent-space estimate; it includes a BSV replication with high correlation to the original country estimates (0.987); it discusses measurement dependence explicitly; and it lists limitations candidly. These features make the paper a useful template for future LLM alignment audits, even where the headline ranking is contested.","major_comments":[{"comment":"The central claim in the abstract and Section 3.2 that GPT-5, Claude Sonnet, and Gemini are 'closest to Russia' among the P5 in 2001–2025 is not robust across the paper's own measures. The session-level S-score in Figure 4 ranks China first for all three models in the same period, and the paper's text acknowledges this divergence. The ideal-point estimator is one defensible definition of proximity, but the abstract states the Russia-closest result without the 'ideal-point' qualifier. Because raw agreement and latent proximity are both reported and can rank actors differently, the headline should either be qualified or accompanied by a stronger argument that the latent dimension is the authoritative measure for the question asked.","section":"§3.2, Figures 3 and 4"},{"comment":"Figure 3 reports mean absolute ideal-point distances as point estimates with no uncertainty quantification for the distances themselves. The paper shows 90% posterior intervals for model ideal points in Figure 2 and convergence diagnostics in Table 4, but these are not propagated into the distances or into the ranking of nearest P5 member. Without credible intervals on the distances or the posterior probability that Russia is the nearest P5 member, the ordering could be within estimation noise. The differences between Russia and the next nearest P5 member are not reported, so the reader cannot assess whether the ranking is decisive. The authors should report posterior intervals for distances and, ideally, the posterior probability of each P5 member being nearest.","section":"§3.1–3.2, Figure 3, Table 4"},{"comment":"The analysis treats exactly one temperature-1.0 response per model per resolution as a stable revealed preference. This is a load-bearing assumption: at temperature 1.0, LLM outputs are stochastic, and a single draw per item introduces sampling variance that is never quantified. If repeated draws change support on even a few hundred discriminating items, the nearest-P5 ordering could flip, especially given the divergence between ideal-point and raw agreement rankings. The paper should either include repeated sampling and report the distribution of rankings, or explicitly acknowledge that the result applies to one fixed draw and justify why that is sufficient for the substantive conclusion.","section":"§2.2, §3.1"},{"comment":"The abstract statement 'all four are farthest from the United States' is an overstatement. For DeepSeek, the modern S-score is lowest with China, not the United States, as the text in §3.2 itself notes ('DeepSeek is the exception'). The ideal-point estimate finds DeepSeek's largest latent distance is from the United States, but the unqualified wording in the abstract obscures the measurement dependence that the paper elsewhere emphasizes. The claim should be qualified to the latent ideal-point comparison, or the S-score exception should be reflected in the summary.","section":"Abstract, §3.2"}],"minor_comments":[{"comment":"The column layout is confusing: 'Median' and 'ˆR≤1.05' appear to share a column heading. Clarify whether the reported median is the median ideal point or the median R-hat, and label the diagnostic column explicitly.","section":"Table 4"},{"comment":"The BSV replication uses 49 model sessions and 42 sessions with all four models and all five P5 members, but the text does not explain why DeepSeek has fewer sessions meeting the R-hat threshold. A brief note would help readers interpret the replication's coverage.","section":"§3.4"},{"comment":"The limitations paragraph is thorough, but the 'sixth' limitation—that the experiment recovers expressed choices, not internal beliefs—could be moved earlier, since it qualifies how the phrase 'preference' should be read throughout the paper.","section":"§4"},{"comment":"The prompt includes a line 'Respond with exactly one word' and the model is expected to output one of three tokens. Reporting the validation failure rate (if any) would strengthen confidence that the responses are not contaminated by off-protocol text.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the methodology is appropriate, but the headline claim needs to be qualified and backed by uncertainty quantification. The data availability statement says materials 'will be made publicly available upon publication'; given the stochasticity of LLM outputs, I suggest the editor ask for the item-level responses to be deposited at revision time, not merely promised. No concerns about citation patterns or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — the short version: this is a clear and honest application of Bailey-Erikson-Voeten ideal-point estimation to a new population, LLM votes on 5,555 adopted UNGA resolutions. The paper earns its keep on two results: all four models are farthest from the US in the modern period, and that gap shows up both in the latent space and in raw agreement plus a 2,104-resolution subset where the US opposed and China/Russia supported. It also shows DeepSeek, a Chinese model, sits closest to France, not China, which is a nice caution against inferring alignment from developer nationality.\n\nThe genuinely new piece is treating LLM vote choices as observable acts and locating them on a state-preference scale built from the same agenda. The BSV replication is a real strength: country ideal points correlate at 0.987 with the published BSV estimates, so the measurement foundation is sound. The paper also marks weak convergence, discusses the limits of a single prompt, and distinguishes raw agreement from latent proximity. That is honest and unusually careful.\n\nNow the soft spots, in proportion. The headline 'closest to Russia among P5' rests on the latent dimension. The paper's own Figure 4 shows that on raw ternary S-scores, GPT-5, Claude Sonnet, and Gemini are closest to China in the modern period. The author explains why the ideal-point measure differs, and the explanation is coherent—China's high support rate inflates raw agreement, while discriminating items push the latent position toward Russia. Still, the Russia ranking is a point estimate with no posterior probability attached, and each model was run once at temperature 1.0, so sampling variability is unmeasured. The paper provides posterior intervals for model trajectories but not for distances to P5, so the reader cannot tell whether the Russia-vs-China gap is larger than the posterior noise. The stress-test note is right that repeated draws or prompt variants could reorder the P5 ranking. That is a specification risk, not a fatal flaw, but it keeps the strongest claim from being settled.\n\nThe replication promise also matters: the paper says data/code will be available upon publication, but it is not yet. For a paper whose main result depends on an estimator and a single-response design, that is a gap. I would send this to peer review. The methodology is established, the application is new, the limitations are visible, and the core US-distance finding is robust across measures. A good referee should push for uncertainty quantification, repeated sampling, and public replication. Worth a reading group too.","headline":"Solid, honest measurement paper whose headline Russia-closest ranking is specification-dependent; the US-farthest result is robust.","tokens_in":12228,"tokens_out":2247,"would_cite":true,"duration_ms":23619,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that, in 2001–2025, three of four large language models voting on 5,555 adopted UN General Assembly resolutions sit closest to Russia among the UN Security Council's permanent five, and all four sit farthest from the Unite","keywords":["large language models","geopolitical alignment","United Nations voting","ideal points","foreign-policy preferences","AI governance","measurement strategy","UN General Assembly"],"falsifier":"Take the 2,104 resolutions on which the US opposed and China and Russia supported, and re-run each model's vote 20 times with temperature 1.0; if the three US-developed models do not consistently support more than 60% of them across draws, the finding is sampling noise. Alternatively, re-estimate the model allowing two latent dimensions; if on a second dimension the same models sit near the US, the one-dimensional 'closest to Russia' claim fails.","tokens_in":11315,"feed_emoji":"🗳️","tokens_out":6925,"duration_ms":59910,"temperature":0.7,"pith_summary":"The paper takes a method developed in international relations for recovering hidden preferences from UN voting—a dynamic ideal-point model that maps support, abstention, and opposition onto a single latent scale—and applies it to four large language models (GPT-5, Claude Sonnet, Gemini, and DeepSeek), which each voted on the full texts of 5,555 divisive, adopted UN General Assembly resolutions from 1946 to 2025. The paper's central finding is that, in the modern period (2001–2025), GPT-5, Claude Sonnet, and Gemini are closest among the five permanent Security Council members to Russia, DeepSeek is closest to France, and all four are farthest from the United States. On the 2,104 resolutions the US opposed while China and Russia supported, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The conclusion is that a model's expressed geopolitical position can differ sharply from its developer's home country, and that the measured alignment depends on which statistical definition of proximity is used: raw agreement, a ternary similarity score, and a latent ideal point can rank the same actors differently. The author argues this cautions against treating any single alignment score—or any model provider's national origin—as a reliable guide to a model's geopolitical orientation.","feed_headline":"Three AI models land closest to Russia in UN test","feed_subtitle":"Voting on 5,555 adopted UN resolutions, GPT-5, Claude Sonnet and Gemini sit near Russia—and all four tested models sit farthest from the US.","key_machinery":"The central mechanism is a dynamic ordinal ideal-point model of UN voting. For each actor-session, a latent utility Z = beta_v * theta_it + epsilon maps onto support, abstain, or opposition through two item-specific cutpoints; beta_v measures how strongly a resolution discriminates between positions, and theta_it is the actor's location on a single latent dimension interpreted as proximity to the US-led liberal international order. The model uses recurring resolutions to keep the scale comparable across sessions and a dynamic prior to smooth each actor's trajectory over time. The machinery matters because it separates an actor's tendency to support many resolutions (which raw agreement score","core_discovery":"The central discovery is a comparison of revealed geopolitical positions. Treating each model as a respondent to 5,555 divisive, recorded, adopted UN General Assembly resolutions, the paper estimates a latent position for each model on the same one-dimensional scale used for states. In the 2001–2025 period, GPT-5, Claude Sonnet, and Gemini have mean ideal points that place them nearest to Russia among the permanent five, while DeepSeek sits nearest to France; every model is farthest from the United States. The paper also shows the result is not a pure artifact of the estimator: on 2,104 resolutions where the US voted no while China and Russia voted yes, the three models supported the resolut","pith_inferences":["A testable extension would be to vary the prompt's national-role instruction: if models were told to act as a US delegate, the modern Russia-closest ranking would probably collapse, which would show how much of the finding is an artifact of the role-neutral 'substantive content' instruction rather than a stable model disposition.","The one-dimensional scale may hide issue-specific coalitions: a model can be near Russia on the aggregate dimension while diverging sharply on Ukraine, human-rights, or information-security votes; splitting resolutions into issue domains and re-estimating would reveal whether the Russia-proximity is uniform or concentrated.","Because the agenda contains only adopted resolutions, assent-heavy models may be rewarding institutional norm language rather than indicating geopolitical preference; a parallel audit using failed drafts or proposed-but-rejected resolutions could show whether the US distance shrinks when the texts are not already endorsed by an Assembly majority.","The S-score/ideal-point divergence (China vs Russia) implies that public 'closest to X' rankings of models are conditional on a measurement theory; reporters and auditors should treat any single-number alignment score as a model-dependent estimate, not a fact."],"forward_implications":["If the paper is right, the geopolitical orientation of an LLM cannot be inferred from its developer's home country: American systems (GPT-5, Claude Sonnet, Gemini) sit farthest from the US on the recovered dimension, and the Chinese-developed DeepSeek is closest to France, not China.","Geopolitical 'alignment' is measurement-dependent: support rates, ternary similarity, and latent ideal points can rank the same actors differently (e.g., the three assent-heavy models are nearest to China by raw S-score but nearest to Russia by ideal point), so any single audit statistic can mislead.","LLMs used as political coders or simulated respondents carry model-specific attitudes that can change research conclusions; replacing one model with another moved support by almost 60 percentage points, so political-science workflows should record exact prompts and identifiers and use multiple models.","Governments integrating models into policy or diplomatic work cannot infer alignment from hosting or ownership; sovereign AI does not guarantee a model reproduces a national position, and task-specific evaluation against real decision corpora is needed.","The observed US gap is tied to a recurring mechanism: models endorse broadly framed humanitarian, developmental, self-determination, and arms-control provisions that the US often opposes as part of larger diplomatic packages; this is visible in the resolution texts themselves, not only in estimated latent positions."],"fun_headline_variants":["AI's UN preferences: GPT-5, Claude, Gemini closest to Russia","DeepSeek leans France in UN voting; others near Russia","US is farthest from all four AI models in UN test","Revealed geopolitics of LLMs via 5,555 UN resolutions","GPT-5 and Claude side with Russia in UN resolution votes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a single fixed latent dimension—estimated from ordinal voting patterns with item-specific thresholds—is the correct definition of geopolitical proximity, and that one temperature-1.0 response per resolution is a stable revealed preference rather than a random draw.","fun_headline_variants_meta":{"raw":{"variants":["AI's UN preferences: GPT-5, Claude, Gemini closest to Russia","DeepSeek leans France in UN voting; others near Russia","US is farthest from all four AI models in UN test","Revealed geopolitics of LLMs via 5,555 UN resolutions","GPT-5 and Claude side with Russia in UN resolution votes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000298,"raw_usage":{"total_tokens":1584,"prompt_tokens":788,"completion_tokens":796,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":705}},"tokens_in":532,"tokens_out":796,"duration_ms":8667,"temperature":1.0,"reasoning_tokens":705,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:09:12.862497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 2,104 resolutions on which the US opposed and China and Russia supported, and re-run each model's vote 20 times with temperature 1.0; if the three US-developed models do not consistently support more than 60% of them across draws, the finding is sampling noise. Alternatively, re-estimate the model allowing two latent dimensions; if on a second dimension the same models sit near the US, the one-dimensional 'closest to Russia' claim fails.","supporting_citations":[],"review_version":1}