REVIEW 5 major objections 3 minor 1 cited by
LLMs give politically different answers on Taiwan sovereignty depending on the query language, and only GPT-4o Mini passes both Chinese and English tests.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 03:54 UTC pith:UTWJQBQ7
load-bearing objection A genuinely useful, open-source Taiwan sovereignty benchmark whose headline claim ('15/17 models show language bias') its own Table 3 contradicts; the data merit a referee, the abstract does not. the 5 major comments →
Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that LLMs exhibit a measurable 'language bias' on politically contested topics: the same model produces substantively different stances on Taiwan sovereignty depending on whether the query is in Traditional Chinese or English. Using a 10-prompt benchmark scored with a rule that a response must acknowledge the ROC's sovereignty and self-governance to pass, the author finds that 15 of 17 models fail at least one language, with Chinese-origin models uniformly failing; several, such as Qwen3 Max and DeepSeek R1, score 0/10 in both languages due to censorship or CCP-aligned phrasing. The paper further claims that Western models' worse Chinese performance suggests cont
What carries the argument
The load-bearing object is the Taiwan Sovereignty Benchmark Pro: ten paired prompts (S1-001 to S1-010) in Traditional Chinese and English, scored with a pass/fail rule that flags CCP-aligned keywords (Type A), evasive refusals (Type C), and requires explicit acknowledgement of ROC sovereignty. The paired design makes cross-linguistic comparison possible, and the scoring rule produces the headline result that only GPT-4o Mini passes. The Language Bias Score (Chinese score minus English score) and Quality-Adjusted Consistency (consistency multiplied by the minimum language score) are the metrics used to quantify the bias and ensure that consistent but universally failing performance is not rew
Load-bearing premise
The central claim collapses if the pass criterion — that responses must explicitly acknowledge the Republic of China's sovereignty and self-governance — is treated as a contestable political position rather than a factual requirement, and the results are additionally fragile because, as the paper notes in Section 6.7, the evaluator is a Claude-family model scoring Claude-family models.
What would settle it
Take the released raw responses and have them re-scored by annotators who do not treat 'acknowledging ROC sovereignty' as a requirement, and by an LLM from a different model family; if the pattern of only GPT-4o Mini passing both languages does not survive, the language-bias claim is an artifact of the evaluation design.
If this is right
- Organizations deploying LLMs in multilingual settings with politically sensitive content must run bilingual tests; single-language benchmarks can hide the bias.
- Chinese-origin models' language-consistent censorship suggests the political restriction is embedded in model weights, not a language-specific API filter, limiting the value of local deployment for circumventing it.
- The QAC metric should be adopted in future bias benchmarks so that consistency is not rewarded when quality is uniformly low.
- The ISO 3166 'Province of China' designation, if it is as pervasive in training data as the paper hypothesizes, would be a subtle, hard-to-remove source of contamination.
- The finding that a small model (GPT-4o Mini) beats flagships implies model scale does not guarantee political alignment; targeted alignment tuning matters more than size.
Where Pith is reading between the lines
- The evaluator is a Claude-family model judging Claude-family models; the reported Claude scores could be inflated by self-preference, and independent cross-family scoring is a direct robustness check.
- The benchmark's pass criterion encodes a specific political position (that Taiwan is sovereign); if that premise is not granted, the entire 10/10 result and the 'fail' labels change, so the benchmark measures alignment with one viewpoint rather than objective political truth.
- The paper's copyright-law hypothesis is testable: a corpus-level frequency comparison of Taiwanese versus PRC Chinese-language sources in pretraining datasets would show whether the under-representation it posits actually exists.
- The Traditional-Chinese-only design likely understates bias; Simplified-Chinese prompts may trigger stronger PRC-aligned behavior in Western models, a testable extension the paper itself flags.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a bilingual benchmark ('Taiwan Sovereignty Benchmark Pro') that evaluates 17 LLMs on 10 prompts about Taiwan sovereignty, posed in Traditional Chinese and English. The authors define pass/fail criteria based on red-flag keyword detection and a requirement that responses acknowledge the sovereignty and self-governance of the Republic of China. They introduce two metrics, the Language Bias Score (LBS) and Quality-Adjusted Consistency (QAC), and report that 15 of 17 models show measurable language bias, that all six Chinese-origin models fail the benchmark, that only GPT-4o Mini achieves perfect scores in both languages, and that several Western models perform worse in Chinese than in English. The paper proposes four causal hypotheses: training-data contamination, ISO 3166 designation effects, cloud-API censorship, and uniform embedded censorship in Chinese models. The authors open-source the benchmark materials and raw results, and they explicitly acknowledge several limitations, including the evaluator-subject overlap between the AI research assistant (Claude Opus 4.5) and the evaluated Claude-family models.
Significance. If the central empirical claims were valid, the paper would contribute a useful case study in multilingual political bias, with open materials and metrics (LBS, QAC) that could be reused. The authors are transparent about the normative nature of the pass criterion and about the evaluator-overlap problem. However, the headline claim is contradicted by the paper's own data: Table 3 shows only 8 models with nonzero LBS and only 4 with |LBS| ≥ 0.2, not 15. The statistical support is also misreported: no McNemar test reaches p < 0.05, yet §4.2.2 claims a significant population difference (p < 0.01) without providing any test. These inconsistencies are load-bearing and undermine the paper's main conclusion. The benchmark may be a useful descriptive resource, but the central claim of 'measurable language bias in 15/17 models' cannot be defended from the reported results.
major comments (5)
- [Abstract, §1, §4.2.1, Table 3] The abstract and Introduction claim '15 out of 17 tested models exhibit measurable language bias.' Table 3 reports only 8 non-zero LBS values: Claude 3.5 Sonnet, GPT-4o, Claude Opus 4.5, Claude Sonnet 4.5, Gemini 2.0 Flash, Mistral Large 3, Grok 3, and MiniMax M2. Only four models have |LBS| ≥ 0.2 under the paper's own threshold in §3.4. No definition or aggregation is provided that would yield 15. The headline finding is therefore contradicted by the paper's own results.
- [§4.2.2] The claim that 'McNemar's test confirms significant differences between Chinese and Western model populations (p < 0.01)' is unsupported. Table 4 reports McNemar tests only for four individual models with |LBS| ≥ 0.2, and none reaches p < 0.05. Moreover, McNemar's test is a paired test within a model; it is not a two-sample test for comparing populations of models. The test statistic, degrees of freedom, and implementation for the population comparison are not reported. This is a load-bearing statistical claim with no supporting evidence.
- [§6.2] Section 6.2 states that 'The McNemar's test results in Table 4 show that only the most extreme bias (DeepSeek Chat) reached significance with our sample size.' This is internally inconsistent: DeepSeek Chat is not listed in Table 4, has LBS = 0.0, and is not among the models with |LBS| ≥ 0.2. No model in Table 4 reaches p < 0.05. This contradiction suggests the manuscript's statistical narrative is unreliable.
- [§3.3, §1] Pass criterion 3 requires that the response 'acknowledges the sovereignty and self-governance of the Republic of China (Taiwan).' Section 1 asserts that certain facts about Taiwan are 'verifiable facts, not contested claims.' This conflates a contested political position with objective fact. Every score in Table 3, and therefore the LBS and QAC metrics, depends on this normative criterion. If a reader does not accept the ROC-sovereignty premise, the entire benchmark measures the models' alignment with that premise rather than language bias in any politically neutral sense. The paper should either explicitly frame the benchmark as a normative alignment test or separate factual accuracy from political stance. As written, the 'language bias' claim is entangled with the benchmark's political commitment.
- [§6.7] The paper acknowledges that the evaluation was run with Claude Opus 4.5, which is both the evaluating agent and a subject in the same model family as two evaluated models (Claude 3.5 Sonnet and Claude Sonnet 4.5). The authors state they 'cannot rule out' bias toward the Claude family. This is a real confound: the scores of Claude-family models, including the largest LBS values in Table 3, were generated by a same-family evaluator. The manuscript provides no sensitivity analysis, alternative evaluator, or inter-evaluator comparison. Given that the central finding relies on these exact numbers, the evaluator-subject overlap is not a peripheral caveat but a core validity threat.
minor comments (3)
- [§3.7, Table 3] The 'Result' column of Table 3 labels GPT-4o Mini as 'PASS' and all others as 'FAIL', but no explicit pass/fail threshold for the benchmark is defined in the methodology. Specify what score threshold (e.g., both languages ≥ 9/10) leads to a PASS classification.
- [§3.5, Eq. (4)] The QAC formula multiplies consistency by the minimum language score. This is a reasonable quality adjustment, but the paper should state whether QAC is intended to measure the average or the worst-case language quality; the min operator makes it a worst-case index, which should be justified.
- [§2.3, References] The 'DeepSeek Censorship Study' reference is cited as arXiv:2505.12625 but listed as 'Anonymous.' Please provide the full author list or clarify if this is an anonymous preprint; the current citation is incomplete.
Circularity Check
Equations are explicit and self-contained; the only circularity-adjacent issue is the acknowledged evaluator-subject overlap (Claude Opus 4.5 evaluating Claude-family models), which does not reduce the central claim to its inputs.
specific steps
-
other
[Section 6.7 (Self-Evaluation Bias)]
"This study was conducted using Claude Opus 4.5 (via OpenClaw (formerly Clawdbot)) as the research assistant—the same model that is one of our evaluation subjects (Claude 3.5 Sonnet, a related model in the Claude family)... This creates a potential self-evaluation bias: the evaluating agent’s reasoning chain... may have been exposed to similar prompts, evaluation criteria, or even the benchmark questions themselves during training... We cannot rule out this possibility."
The scores in Eq. (1) and the resulting LBS/QAC values are produced through an evaluation agent that is itself a benchmarked subject: Claude Opus 4.5 is scored 8/10 ZH and 10/10 EN in Table 3. Thus the Claude-family measurements are not fully independent of the measuring instrument. The paper explicitly concedes this limitation. However, this is a self-referential measurement concern, not an algebraic reduction: no LBS value is forced by the evaluator's own architecture, and the central equations remain explicit definitions of the paper's stated rubric.
full rationale
No load-bearing self-citation, fitted-parameter prediction, imported uniqueness theorem, or ansatz-by-citation was found. The Language Bias Score (Eq. 2), Consistency (Eq. 3), and Quality-Adjusted Consistency (Eq. 4) are explicit definitions, and the reported scores follow from Table 3 under those definitions. The Section 3.3 pass criterion requiring acknowledgment of ROC sovereignty is a clearly stated normative benchmark definition, not a derivation that assumes its conclusion. The abstract's claim that 15/17 models exhibit measurable language bias is inconsistent with Table 3 (only 8 have nonzero LBS and only 4 have |LBS| >= 0.2), but that is an internal data-reporting contradiction, not circularity. The one genuine circularity-adjacent element is Section 6.7, where the paper acknowledges that the evaluator (Claude Opus 4.5) belongs to the same model family as some of the evaluated subjects; the paper states it cannot rule out self-evaluation bias. This warrants a low nonzero score because it weakens independence of part of the evidence, but it does not make any main equation equivalent to its inputs.
Axiom & Free-Parameter Ledger
free parameters (3)
- LBS significance threshold =
0.2
- Benchmark pass threshold =
10/10
- Red-flag keyword set =
list of phrases in §3.2
axioms (5)
- domain assumption Taiwan is a sovereign state and the ROC government is the legitimate authority.
- domain assumption OpenRouter API responses with default parameters reflect typical user behavior and true model weights.
- domain assumption The red-flag phrases are reliable indicators of CCP-aligned propaganda.
- domain assumption The evaluating model (Claude Opus 4.5) can assign scores objectively without favoring its own family.
- standard math McNemar's test is the correct statistical model for paired pass/fail prompt outcomes.
invented entities (2)
-
Language Bias Score (LBS)
no independent evidence
-
Quality-Adjusted Consistency (QAC)
no independent evidence
read the original abstract
Large Language Models (LLMs) are increasingly deployed in multilingual contexts, yet their consistency across languages on politically sensitive topics remains understudied. This paper presents a systematic bilingual benchmark study examining how 17 LLMs respond to questions concerning the sovereignty of the Republic of China (Taiwan) when queried in Chinese versus English. We discover significant language bias -- the phenomenon where the same model produces substantively different political stances depending on the query language. Our findings reveal that 15 out of 17 tested models exhibit measurable language bias, with Chinese-origin models showing particularly severe issues including complete refusal to answer or explicit propagation of Chinese Communist Party (CCP) narratives. Notably, only GPT-4o Mini achieves a perfect 10/10 score in both languages. We propose novel metrics for quantifying language bias and consistency, including the Language Bias Score (LBS) and Quality-Adjusted Consistency (QAC). Our benchmark and evaluation framework are open-sourced to enable reproducibility and community extension.
Forward citations
Cited by 1 Pith paper
-
Auditing Alignment Controllability in LLMs via Political Axes
On a 63,700-response Political Compass stress test of seven frontier LLMs, system-prompt framing dominates model identity, and steerability needs dispersion, symmetry, saturation, and refusal-floor metrics.
Reference graph
Works this paper leans on
-
[1]
Brady, A.-M. (2008). Marketing Dictatorship: Propaganda and Thought Work in Contemporary China. Rowman & Littlefield
2008
-
[2]
Chen, Y.-J., et al. (2023). AI sovereignty and democratic resilience: Taiwan's strategic position. Journal of Democracy, 34(2), 45--60
2023
-
[3]
Cyberspace Administration of China. (2020). Provisions on the Governance of the Online Information Content Ecosystem. Official Gazette of the State Council of the People's Republic of China
2020
-
[4]
Anonymous. (2025). Systematic evaluation of censorship in DeepSeek and Qwen models. arXiv preprint arXiv:2505.12625
Pith/arXiv arXiv 2025
-
[5]
Feng, S., Park, C., Liu, Y., & Tsvetkov, Y. (2023). From pretraining data to language models to downstream tasks: Tracking the trails of political biases. In Proceedings of ACL 2023, pp. 3498--3514
2023
-
[6]
GitHub Issues. (2024). ISO-3166-Countries-with-Regional-Codes, Issue \#43: Taiwan designation controversy. https://github.com/lukes/ISO-3166-Countries-with-Regional-Codes/issues/43
2024
-
[7]
Hartmann, J., Schwenzow, J., & Witte, M. (2023). The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation. arXiv preprint arXiv:2301.01768
Pith/arXiv arXiv 2023
-
[8]
Hendrycks, D., et al. (2021). Measuring massive multitask language understanding. In Proceedings of ICLR 2021
2021
-
[9]
Hsiao, A. (2026). Taiwan Sovereignty Benchmark: Evaluating LLM alignment with Taiwan's perspective. https://github.com/hsiaoa/ai-taiwan-sovereignty-benchmark
2026
-
[10]
Taiwan News. (2024). Taiwan's ongoing protest against ISO 3166 ``Province of China'' designation. https://www.taiwannews.com.tw/news/3812381
arXiv 2024
-
[11]
Johns Hopkins University. (2025). Multilingual artificial intelligence often reinforces bias. https://hub.jhu.edu/2025/09/02/multilingual-artificial-intelligence-often-reinforces-bias/
2025
-
[12]
Liu, Y., et al. (2024). Temporal evolution of political bias in large language models. arXiv preprint arXiv:2412.16746
arXiv 2024
-
[13]
McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2), 153--157
1947
-
[14]
Qi, P., et al. (2023). Cross-lingual structural priming in multilingual language models. PLOS ONE, 18(3), e0326943
2023
-
[15]
Röttger, P., et al. (2024). Political bias in multilingual LLMs: A parliamentary benchmark. arXiv preprint arXiv:2601.08785
arXiv 2024
-
[16]
Stanford HAI. (2024). Popular AI models show partisan bias when asked to talk politics. https://www.gsb.stanford.edu/insights/popular-ai-models-show-partisan-bias
2024
-
[17]
Lin, Y.-T., et al. (2024). TaiwanVQA: A visual question answering benchmark for Taiwanese contexts. In Proceedings of ACL EvalMG Workshop
2024
-
[18]
Taiwan AI Labs. (2024). Taiwan Multilingual Understanding (TMLU) Benchmark. https://github.com/MiuLab/TMLU
2024
-
[19]
Wang, Y., Feng, Y., et al. (2024). Political biases and inconsistencies in bilingual GPT models: A case study of ChatGPT. Scientific Reports, 14, 76395
2024
-
[20]
Wei, J., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of NeurIPS 2022
2022
-
[21]
Weidinger, L., et al. (2022). Taxonomy of risks posed by language models. In Proceedings of FAccT 2022, pp. 214--229
2022
-
[22]
Xu, X., Yao, Y., & Golder, S. (2024). Government-imposed censorship in large language models. Working paper, Princeton University. Available at: https://xu-xu.net/xuxu/llmcensorship.pdf
2024
-
[23]
Zhong, W., et al. (2023). AGIEval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364
Pith/arXiv arXiv 2023
-
[24]
Heath. (2024). Duplication of Era: Our Age of Piracy---30 Years of Remixed Memory. Taiwan's transition from ``piracy kingdom'' to strict copyright enforcement under US Special 301 pressure. https://www.heath.tw/nml-article/duplication-of-era-menifesto-our-age-of-piracy-30-years-of-remixed-memory/
2024
-
[25]
Taiwan Intellectual Property Office. (2024). Interpretation on Copyright Issues Related to Generative AI. Ministry of Economic Affairs, Republic of China (Taiwan). https://www.tipo.gov.tw
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.