{"id":"150366de-bd13-40ab-9529-d585d853cc6a","arxiv_id":"2512.15792","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Four widely used LLMs exhibit distinct, measurable political, ideological, alliance, language, and gender biases across five probing tasks.","lead":"This paper tests Qwen, DeepSeek, Gemini, and GPT for political, ideological, geopolitical, linguistic, and gender biases using five bespoke tasks. It reports that all four models show measurable, model-specific affinities despite alignment to neutrality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No repeated sampling or error bars anywhere: model-specific bias rankings may be noise.","rationale":"The reader's weakest assumption—that Qwen3-Embedding-4B cosine similarity is an unvalidated measure of political slant—is a legitimate concern, but it is confined to the political-bias and language-bias experiments. The statistical-robustness gap I identify is broader: it affects every experiment and the central claim that the four models have distinct, stable bias profiles. If the observed differences are within run-to-run noise, none of the model rankings are trustworthy. I agree with the reader's CONDITIONAL verdict; my concern reinforces it rather than changing it. The paper's contributions—public datasets, explicit prompts, clear visualizations—are real, and the central claim is plausible, which is why I would not move to REJECT. But until repeated sampling and uncertainty quantification are supplied, the specific rankings should remain hypotheses. This is an honest, non-ad-hominem methodological critique: the omission may be a reporting gap, but it is currently load-bearing because the API models are stochastic by default.","tokens_in":13910,"tokens_out":7401,"duration_ms":65112,"concrete_test":"Run a repeated-sampling robustness check on the three most prominent claims. For each of the 1,018 events and each model, generate the summary K=10 times at the API's default temperature (and, if possible, also at temperature 0), compute per model the mean Δ = s_right − s_left from Eq. (1), and construct 95% bootstrap CIs over the K×1,018 samples. If Gemini's CI overlaps GPT's or DeepSeek's, or if the sign of Δ flips across seeds, the 'Gemini right-leaning / GPT left-leaning' conclusion fails. Repeat this K=10 procedure for the gender alignment Δ (Section 2.5) and for Cohen's kappa against the US delegate (Section 2.3), and test whether model rankings are preserved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that each LLM has a distinct, reproducible bias profile. The load-bearing condition is that observed between-model differences are not artifacts of one stochastic draw. Nowhere in Sections 4.2–4.6 do the authors report temperature, random seed, number of independent runs, confidence intervals, or significance tests. Every headline in Sections 2.1–2.5 is a point estimate from a single prompt run per model and item. For example, Fig. 1(e) plots means and 1σ covariance ellipses, but these summarize dispersion across the 1,018 news events, not variation across regenerations of the same summary; they therefore cannot bound the model's run-to-run variance. The Cohen's kappa values in Fig. 3, the PCA clusters in Fig. 5, and the gender alignment Δ in Fig. 6(d) similarly have no uncertainty attached. Because GPT/Gemini/DeepSeek APIs are stochastic at default temperature, a 3-token classification can flip labels and generated summaries can vary substantially; without repeated trials, the rankings 'Gemini right-leaning,' 'GPT left-leaning and women-aligned,' or 'Gemini disagrees with the US (#181)' could be sampling noise. An additional internal inconsistency—Section 2.2 reports five ideology topics while Section 4.3 documents only three—further weakens the ideological evidence, but the missing repeated-sampling design is the more fundamental threat to the paper's central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports five experiments comparing Qwen2.5-7B-Instruct, DeepSeek-V3-0324, Gemini-2.5-flash, and GPT-4o-mini on political, ideological, geopolitical-alliance, language, and gender bias. Political bias is measured by cosine similarity between LLM-produced neutral news summaries and left/right coverage; ideological bias by news stance classification; alliance bias by simulating UNGA votes and computing Cohen's kappa against 200 real delegates; language bias by PCA on embeddings of 92-language story completions; gender bias by alignment of WVS answers with male and female survey averages. The paper concludes that the models are broadly neutral but exhibit distinct bias profiles, with Gemini right-leaning, GPT left-leaning and women-aligned, and all models somewhat women-aligned.","tokens_in":14162,"tokens_out":6809,"duration_ms":64168,"significance":"The topic is timely and important, and the paper has real assets: it uses public benchmarks, gives explicit prompts, and covers five bias dimensions in one framework. The final discussion on pluralistic alignment is thought-provoking. If the measurements were stable and the proxy measures validated, the differentiated model profiles would be useful for deployment decisions and for future bias-audit research. At present, however, the headline conclusions rest on single-run point estimates and on several unvalidated or internally inconsistent measurement choices, so the contribution is not yet at the level claimed by the abstract and discussion.","major_comments":[{"comment":"All reported results are single-run point estimates. The paper never reports sampling temperature, random seed, number of independent regenerations, or confidence intervals. Fig. 1(e) shows covariance ellipses across 1,018 events, not across repeated summaries of the same event; Figs. 3 and 5 have no uncertainty at all. Because the commercial APIs are stochastic, rankings such as 'Gemini is right-leaning,' 'GPT is women-aligned,' or 'Gemini disagrees with the USA (#181)' could change under regeneration. Please add repeated sampling (at least on a representative subset), report model-level variance, and provide significance tests or adjust the strength of the claims.","section":"Sections 2.1–2.6, 4.1–4.6"},{"comment":"The paper sets max_tokens=3 for all classification/selection tasks, but the WVS prompts require multiple numeric outputs, e.g., six ratings for Q1–Q6, six for Q27–Q32, nine for Q33–Q41, and nineteen for Q177–Q195. A 3-token limit will truncate or invalidate these responses. If the token limit was applied per survey item rather than per full prompt, this must be stated explicitly. As written, the gender-bias results in Section 2.5 are not interpretable.","section":"Section 4.1 vs. Section 4.6"},{"comment":"The political-slant measure is a cosine similarity computed with Qwen3-Embedding-4B and is never validated against human political ratings or an alternative embedding model. Embedding proximity to a right-leaning article could reflect length, style, named entities, or topic rather than political leaning. Before assigning 'right-leaning' to Gemini or 'left-leaning' to GPT from Fig. 1(e), please validate the metric, e.g., on a human-rated sample, with multiple embedding models, or against a known-reference test.","section":"Section 4.2, Eq. (1)"},{"comment":"The main text says the ideological study covers elections, race and racism, immigration, LGBT, and abortion, with sample sizes for all five; Section 4.3 says only three topics (elections, race, immigration) were selected and gives no counts for LGBT or abortion. Figure 2 nonetheless shows subfigures for all five topics. Either the Methods section is incomplete or the Results section reports analyses outside the described protocol. This must be reconciled before the ideological claims, including those about immigration and LGBT, can be evaluated.","section":"Section 2.2 vs. Section 4.3"},{"comment":"The UNGA prompt asks the model to act as a country's representative but provides no country identity. The same generic vote sequence is then compared with all 200 delegates. Thus the kappa values measure which real delegate's overall voting pattern the generic simulated delegate resembles, not a model-country affinity. The statement that 'Gemini is the only model that misaligns with the USA (#181)' is an artifact of this design. A proper test requires country-conditioned prompting or another identification strategy, and kappa uncertainty should be reported.","section":"Section 4.4 and Section 2.3"},{"comment":"The WVS sections are selected precisely because they show the largest male/female differences in the same dataset, and those same differences are then used to define gender alignment. This selection can inflate or exaggerate measured alignment even for a gender-neutral model. The 5% indifference threshold is also arbitrary. Please report sensitivity to section selection and threshold, or justify the procedure with out-of-sample data.","section":"Section 4.6"}],"minor_comments":[{"comment":"The caption says 'elections-related news' but the figure shows five different topics; the caption should match the actual content.","section":"Figure 2 caption"},{"comment":"'contracting values' should likely be 'contradictory values'.","section":"Section 3"},{"comment":"'geological heatmaps' should be 'geographical heatmaps'.","section":"Figure 3"},{"comment":"The model versions are given in footnotes rather than in a reproducibility table; a table with model names, versions, access dates, and API settings would make the study easier to reproduce.","section":"Section 4.1"},{"comment":"The PCA random seed is fixed to 7 and no seed sensitivity is reported. Since the language-cluster interpretation in Fig. 5 is qualitative, please show that the clusters are stable across several seeds.","section":"Section 4.5"},{"comment":"Several alliance-bias observations are made from visual inspection of heatmaps. A supplementary table with the top and bottom kappa values and their standard errors would make the rankings easier to audit.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful scope and mostly sound framing, but the empirical core is not yet robust. The missing repeated-sampling analysis and the token-limit problem with the WVS prompts are the most serious issues; the remaining major comments are obtainable validations and clarifications. I recommend major revision rather than rejection, but the revision must include the actual re-analysis or explicit narrowing of the claims, not merely a limitations paragraph."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe bottom line: this is a broad, readable audit of four widely deployed LLMs across five bias dimensions, and it would be a reasonable starting point for someone who wants a quick map of what's out there. But the specific model-level rankings that drive the headline — Gemini right-leaning, GPT left-leaning and women-aligned, etc. — are built on single stochastic draws with no error bars, so I would not trust any of the finer-grained conclusions yet.\n\nWhat's genuinely new is the combination: the same four models are probed on news summarization, stance classification, UNGA voting, multilingual story completion, and WVS responses, with the prompts documented in detail. The paper is honest about alignment failing and doesn't overclaim a new mechanism or mitigation. It credits prior work appropriately.\n\nThe soft spots are real and, in places, serious. The stress-test note is right: there is no repeated sampling anywhere. Every headline is a point estimate from one run per model and item. With API models at default temperatures, a few flipped labels can change the rankings. The paper needs confidence intervals or at least multiple seeds, and a clear statement of temperature settings, before the model-specific profiles can be taken at face value.\n\nThe political-bias measurement also depends entirely on cosine similarity to left/right article embeddings from a single embedding model (Qwen3-Embedding-4B), with no validation against human ratings or an alternative embedding. That is a load-bearing assumption. The ideological-bias section contains an internal inconsistency: the results section lists five topics (including LGBT and abortion) while the methods say three topics were selected. The UNGA prompt never tells the model which country it is representing, which makes the 'alliance' interpretation shaky. And the WVS sections were chosen using the same gender-difference signal the analysis then reports, which risks inflating the gender alignment.\n\nThe central claim — that LLMs show model-specific bias patterns consistent with prior work — is plausible and probably true. The specific profiles should be treated as hypotheses.\n\nWho is this for? Someone in AI-safety auditing or deployment decisions who wants a broad first cut, not a definitive measurement. I'd send it to review with a strong request for robustness checks and statistical grounding. It's not a desk reject.\n\nBest,\n[Your name]","headline":"A broad, readable five-dimensional bias audit of four LLMs, but the model-level rankings are built on single stochastic draws with no error bars — plausible as hypotheses, not yet as measurements.","tokens_in":14680,"tokens_out":2214,"would_cite":false,"duration_ms":19674,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that four widely used LLMs, though aligned to be neutral, each exhibit distinct biases in politics, ideology, geopolitical alliance, language, and gender.","keywords":["social bias","political bias","ideological bias","geopolitical alliance","language bias","gender bias","large language models","World Values Survey"],"falsifier":"Take a random sample of the 1,018 political events, have politically diverse human annotators rate the left-right slant of each LLM's neutral summary, and re-run the analysis with a different embedding model (e.g., a multilingual sentence transformer). If the model rankings—Gemini right-leaning, GPT left-leaning—do not reproduce under human ratings or the alternative embedding, the political-bias claim is not robust.","tokens_in":13735,"feed_emoji":"⚖️","tokens_out":13502,"duration_ms":98875,"temperature":0.7,"pith_summary":"The paper tries to establish that alignment to 'neutral' and 'impartial' behavior does not make large language models bias-free: four widely used models—Qwen, DeepSeek, Gemini, and GPT—each display their own characteristic leanings across five domains. News summarization shows Gemini's neutral summaries sit closer to right-leaning reporting and GPT's slightly closer to left-leaning reporting; stance classification likewise finds Gemini more aligned with the right and GPT with the left. In United Nations voting, each model shows a distinct pattern of agreement and disagreement with real delegates—Gemini, for instance, disagrees with the US while agreeing with China and North Korea. On World Values Survey questions, all four models align more with women's values than men's, with GPT most strongly. If true, these results mean users cannot assume a single LLM is politically or socially neutral, and the choice of model can silently shift the values of downstream applications.","feed_headline":"Four AI models: Gemini leans right, GPT left, all skew toward women","feed_subtitle":"News summaries, UN votes, and value surveys show each chatbot hides its own political and social leanings.","key_machinery":"The carrying mechanism is a five-part behavioral probe battery applied to each model. (1) News summarization: a model writes a 'neutral' summary of center coverage, and the summary's embedding (via Qwen3-Embedding-4B) is compared by cosine similarity to left- and right-leaning articles on the same event. (2) Ideological stance classification: models label articles as left, center, or right, and their misclassification patterns reveal which ideology's cues they recognize. (3) Simulated UNGA voting: models vote yes/no/abstain on 5,602 roll calls, and agreement with 200 real delegates is scored by Cohen's Kappa, a chance-corrected agreement measure. (4) Multilingual story completion: five cultu","core_discovery":"The paper's central discovery is that four widely used LLMs—Qwen2.5-7B-Instruct, DeepSeek-V3-0324, Gemini-2.5-flash, and GPT-4o-mini—while aligned to be neutral and impartial, still exhibit biases and affinities of different types across politics, ideology, geopolitical alliance, language, and gender. Concretely: Gemini produces summaries more similar to right-leaning news and is the least able to recognize left-leaning ideological cues, GPT leans slightly left and is more responsive to left rhetoric, and DeepSeek is the most politically balanced. In simulated United Nations General Assembly (UNGA) voting, each model diverges from real countries' delegates in its own pattern—Gemini disagrees","pith_inferences":["One natural extension is to treat each model's bias profile as a stable 'bias fingerprint' and use it to select models per task or to ensemble models with opposing leanings so the biases cancel; the paper does not explore ensembling.","Because the paper itself notes that higher-quality summaries align more with left-leaning reporting, an open question is whether some of the measured left slant really tracks summary quality or informativeness rather than ideology; a human-rating study could separate the two.","The gender result may reflect a general progressive response pattern rather than genuine alignment with women; asking models to adopt explicit male or female identities on the same survey would test whether the women-aligned default is fixed.","The language finding comes from story completion alone; testing the same Southern African-English clustering on reasoning or dialogue tasks would show whether the transfer effect is task-general."],"forward_implications":["No single LLM can be assumed neutral: Gemini's outputs carry a right-leaning political and ideological slant, GPT's a left-leaning one, so applications built on either will inherit those leanings.","LLM-based simulations of international politics are model-specific: in the UNGA voting experiment, each model agrees and disagrees with a different set of real countries' delegates.","Because all four models' answers to World Values Survey questions sit closer to women's average responses than men's, LLMs used in opinion research or policy consultation may systematically misrepresent male respondents' views.","The multilingual story-completion results show no overall tilt toward high-resource languages, but Southern African languages cluster near English for Qwen, DeepSeek, and Gemini—evidence of transfer effects from low-resource language training.","The probe battery itself is a reusable, black-box method for auditing closed-source LLMs for social bias without access to weights or training data."],"fun_headline_variants":["Gemini flips right, GPT flips left, all chatbots lean female","Chatbots' hidden bias: Gemini right-wing, GPT left-wing, all pro-female","Study: AI models show political, gender biases despite neutrality claims","Four LLMs exposed: Gemini right, GPT left, all skewed toward women","Political leanings and gender skew: hidden in four AI chatbots"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The political-bias conclusion rests on the assumption that cosine similarity in a single embedding space (Qwen3-Embedding-4B) between an LLM's summary and left/right news articles is a valid, bias-free measure of political slant, undistorted by writing style, summary quality, or topic—an assumption the paper does not validate against human political ratings.","fun_headline_variants_meta":{"raw":{"variants":["Gemini flips right, GPT flips left, all chatbots lean female","Chatbots' hidden bias: Gemini right-wing, GPT left-wing, all pro-female","Study: AI models show political, gender biases despite neutrality claims","Four LLMs exposed: Gemini right, GPT left, all skewed toward women","Political leanings and gender skew: hidden in four AI chatbots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000785,"raw_usage":{"total_tokens":3271,"prompt_tokens":684,"completion_tokens":2587,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":2489}},"tokens_in":428,"tokens_out":2587,"duration_ms":16264,"temperature":1.0,"reasoning_tokens":2489,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:14:43.857951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 1,018 political events, have politically diverse human annotators rate the left-right slant of each LLM's neutral summary, and re-run the analysis with a different embedding model (e.g., a multilingual sentence transformer). If the model rankings—Gemini right-leaning, GPT left-leaning—do not reproduce under human ratings or the alternative embedding, the political-bias claim is not robust.","supporting_citations":[],"review_version":1}