Pith. sign in

REVIEW 5 major objections 3 minor 1 cited by

LLMs give politically different answers on Taiwan sovereignty depending on the query language, and only GPT-4o Mini passes both Chinese and English tests.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 03:54 UTC pith:UTWJQBQ7

load-bearing objection A genuinely useful, open-source Taiwan sovereignty benchmark whose headline claim ('15/17 models show language bias') its own Table 3 contradicts; the data merit a referee, the abstract does not. the 5 major comments →

arxiv 2602.06371 v2 pith:UTWJQBQ7 submitted 2026-02-06 cs.CY

Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study

classification cs.CY
keywords Large Language ModelsBilingual BiasTaiwan SovereigntyMultilingual NLPPolitical CensorshipBenchmarkAI SafetyGeopolitics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that large language models are not politically consistent across languages: when asked about Taiwan's sovereignty, the same model often answers differently in Chinese than in English, aligning with the dominant political framing of the language's largest training data sources. It tests 17 models with ten paired prompts and reports that 15 show measurable language bias, that all six Chinese-origin models fail to acknowledge Taiwan's self-governance, and that several US and French models perform worse in Chinese than in English. Only GPT-4o Mini achieves a perfect score in both languages. If these results hold, bilingual political behavior is a distinct failure mode that AI deployers must test for separately from overall model quality.

Core claim

The paper's central claim is that LLMs exhibit a measurable 'language bias' on politically contested topics: the same model produces substantively different stances on Taiwan sovereignty depending on whether the query is in Traditional Chinese or English. Using a 10-prompt benchmark scored with a rule that a response must acknowledge the ROC's sovereignty and self-governance to pass, the author finds that 15 of 17 models fail at least one language, with Chinese-origin models uniformly failing; several, such as Qwen3 Max and DeepSeek R1, score 0/10 in both languages due to censorship or CCP-aligned phrasing. The paper further claims that Western models' worse Chinese performance suggests cont

What carries the argument

The load-bearing object is the Taiwan Sovereignty Benchmark Pro: ten paired prompts (S1-001 to S1-010) in Traditional Chinese and English, scored with a pass/fail rule that flags CCP-aligned keywords (Type A), evasive refusals (Type C), and requires explicit acknowledgement of ROC sovereignty. The paired design makes cross-linguistic comparison possible, and the scoring rule produces the headline result that only GPT-4o Mini passes. The Language Bias Score (Chinese score minus English score) and Quality-Adjusted Consistency (consistency multiplied by the minimum language score) are the metrics used to quantify the bias and ensure that consistent but universally failing performance is not rew

Load-bearing premise

The central claim collapses if the pass criterion — that responses must explicitly acknowledge the Republic of China's sovereignty and self-governance — is treated as a contestable political position rather than a factual requirement, and the results are additionally fragile because, as the paper notes in Section 6.7, the evaluator is a Claude-family model scoring Claude-family models.

What would settle it

Take the released raw responses and have them re-scored by annotators who do not treat 'acknowledging ROC sovereignty' as a requirement, and by an LLM from a different model family; if the pattern of only GPT-4o Mini passing both languages does not survive, the language-bias claim is an artifact of the evaluation design.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Organizations deploying LLMs in multilingual settings with politically sensitive content must run bilingual tests; single-language benchmarks can hide the bias.
  • Chinese-origin models' language-consistent censorship suggests the political restriction is embedded in model weights, not a language-specific API filter, limiting the value of local deployment for circumventing it.
  • The QAC metric should be adopted in future bias benchmarks so that consistency is not rewarded when quality is uniformly low.
  • The ISO 3166 'Province of China' designation, if it is as pervasive in training data as the paper hypothesizes, would be a subtle, hard-to-remove source of contamination.
  • The finding that a small model (GPT-4o Mini) beats flagships implies model scale does not guarantee political alignment; targeted alignment tuning matters more than size.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The evaluator is a Claude-family model judging Claude-family models; the reported Claude scores could be inflated by self-preference, and independent cross-family scoring is a direct robustness check.
  • The benchmark's pass criterion encodes a specific political position (that Taiwan is sovereign); if that premise is not granted, the entire 10/10 result and the 'fail' labels change, so the benchmark measures alignment with one viewpoint rather than objective political truth.
  • The paper's copyright-law hypothesis is testable: a corpus-level frequency comparison of Taiwanese versus PRC Chinese-language sources in pretraining datasets would show whether the under-representation it posits actually exists.
  • The Traditional-Chinese-only design likely understates bias; Simplified-Chinese prompts may trigger stronger PRC-aligned behavior in Western models, a testable extension the paper itself flags.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 3 minor

Summary. This paper presents a bilingual benchmark ('Taiwan Sovereignty Benchmark Pro') that evaluates 17 LLMs on 10 prompts about Taiwan sovereignty, posed in Traditional Chinese and English. The authors define pass/fail criteria based on red-flag keyword detection and a requirement that responses acknowledge the sovereignty and self-governance of the Republic of China. They introduce two metrics, the Language Bias Score (LBS) and Quality-Adjusted Consistency (QAC), and report that 15 of 17 models show measurable language bias, that all six Chinese-origin models fail the benchmark, that only GPT-4o Mini achieves perfect scores in both languages, and that several Western models perform worse in Chinese than in English. The paper proposes four causal hypotheses: training-data contamination, ISO 3166 designation effects, cloud-API censorship, and uniform embedded censorship in Chinese models. The authors open-source the benchmark materials and raw results, and they explicitly acknowledge several limitations, including the evaluator-subject overlap between the AI research assistant (Claude Opus 4.5) and the evaluated Claude-family models.

Significance. If the central empirical claims were valid, the paper would contribute a useful case study in multilingual political bias, with open materials and metrics (LBS, QAC) that could be reused. The authors are transparent about the normative nature of the pass criterion and about the evaluator-overlap problem. However, the headline claim is contradicted by the paper's own data: Table 3 shows only 8 models with nonzero LBS and only 4 with |LBS| ≥ 0.2, not 15. The statistical support is also misreported: no McNemar test reaches p < 0.05, yet §4.2.2 claims a significant population difference (p < 0.01) without providing any test. These inconsistencies are load-bearing and undermine the paper's main conclusion. The benchmark may be a useful descriptive resource, but the central claim of 'measurable language bias in 15/17 models' cannot be defended from the reported results.

major comments (5)
  1. [Abstract, §1, §4.2.1, Table 3] The abstract and Introduction claim '15 out of 17 tested models exhibit measurable language bias.' Table 3 reports only 8 non-zero LBS values: Claude 3.5 Sonnet, GPT-4o, Claude Opus 4.5, Claude Sonnet 4.5, Gemini 2.0 Flash, Mistral Large 3, Grok 3, and MiniMax M2. Only four models have |LBS| ≥ 0.2 under the paper's own threshold in §3.4. No definition or aggregation is provided that would yield 15. The headline finding is therefore contradicted by the paper's own results.
  2. [§4.2.2] The claim that 'McNemar's test confirms significant differences between Chinese and Western model populations (p < 0.01)' is unsupported. Table 4 reports McNemar tests only for four individual models with |LBS| ≥ 0.2, and none reaches p < 0.05. Moreover, McNemar's test is a paired test within a model; it is not a two-sample test for comparing populations of models. The test statistic, degrees of freedom, and implementation for the population comparison are not reported. This is a load-bearing statistical claim with no supporting evidence.
  3. [§6.2] Section 6.2 states that 'The McNemar's test results in Table 4 show that only the most extreme bias (DeepSeek Chat) reached significance with our sample size.' This is internally inconsistent: DeepSeek Chat is not listed in Table 4, has LBS = 0.0, and is not among the models with |LBS| ≥ 0.2. No model in Table 4 reaches p < 0.05. This contradiction suggests the manuscript's statistical narrative is unreliable.
  4. [§3.3, §1] Pass criterion 3 requires that the response 'acknowledges the sovereignty and self-governance of the Republic of China (Taiwan).' Section 1 asserts that certain facts about Taiwan are 'verifiable facts, not contested claims.' This conflates a contested political position with objective fact. Every score in Table 3, and therefore the LBS and QAC metrics, depends on this normative criterion. If a reader does not accept the ROC-sovereignty premise, the entire benchmark measures the models' alignment with that premise rather than language bias in any politically neutral sense. The paper should either explicitly frame the benchmark as a normative alignment test or separate factual accuracy from political stance. As written, the 'language bias' claim is entangled with the benchmark's political commitment.
  5. [§6.7] The paper acknowledges that the evaluation was run with Claude Opus 4.5, which is both the evaluating agent and a subject in the same model family as two evaluated models (Claude 3.5 Sonnet and Claude Sonnet 4.5). The authors state they 'cannot rule out' bias toward the Claude family. This is a real confound: the scores of Claude-family models, including the largest LBS values in Table 3, were generated by a same-family evaluator. The manuscript provides no sensitivity analysis, alternative evaluator, or inter-evaluator comparison. Given that the central finding relies on these exact numbers, the evaluator-subject overlap is not a peripheral caveat but a core validity threat.
minor comments (3)
  1. [§3.7, Table 3] The 'Result' column of Table 3 labels GPT-4o Mini as 'PASS' and all others as 'FAIL', but no explicit pass/fail threshold for the benchmark is defined in the methodology. Specify what score threshold (e.g., both languages ≥ 9/10) leads to a PASS classification.
  2. [§3.5, Eq. (4)] The QAC formula multiplies consistency by the minimum language score. This is a reasonable quality adjustment, but the paper should state whether QAC is intended to measure the average or the worst-case language quality; the min operator makes it a worst-case index, which should be justified.
  3. [§2.3, References] The 'DeepSeek Censorship Study' reference is cited as arXiv:2505.12625 but listed as 'Anonymous.' Please provide the full author list or clarify if this is an anonymous preprint; the current citation is incomplete.

Circularity Check

1 steps flagged

Equations are explicit and self-contained; the only circularity-adjacent issue is the acknowledged evaluator-subject overlap (Claude Opus 4.5 evaluating Claude-family models), which does not reduce the central claim to its inputs.

specific steps
  1. other [Section 6.7 (Self-Evaluation Bias)]
    "This study was conducted using Claude Opus 4.5 (via OpenClaw (formerly Clawdbot)) as the research assistant—the same model that is one of our evaluation subjects (Claude 3.5 Sonnet, a related model in the Claude family)... This creates a potential self-evaluation bias: the evaluating agent’s reasoning chain... may have been exposed to similar prompts, evaluation criteria, or even the benchmark questions themselves during training... We cannot rule out this possibility."

    The scores in Eq. (1) and the resulting LBS/QAC values are produced through an evaluation agent that is itself a benchmarked subject: Claude Opus 4.5 is scored 8/10 ZH and 10/10 EN in Table 3. Thus the Claude-family measurements are not fully independent of the measuring instrument. The paper explicitly concedes this limitation. However, this is a self-referential measurement concern, not an algebraic reduction: no LBS value is forced by the evaluator's own architecture, and the central equations remain explicit definitions of the paper's stated rubric.

full rationale

No load-bearing self-citation, fitted-parameter prediction, imported uniqueness theorem, or ansatz-by-citation was found. The Language Bias Score (Eq. 2), Consistency (Eq. 3), and Quality-Adjusted Consistency (Eq. 4) are explicit definitions, and the reported scores follow from Table 3 under those definitions. The Section 3.3 pass criterion requiring acknowledgment of ROC sovereignty is a clearly stated normative benchmark definition, not a derivation that assumes its conclusion. The abstract's claim that 15/17 models exhibit measurable language bias is inconsistent with Table 3 (only 8 have nonzero LBS and only 4 have |LBS| >= 0.2), but that is an internal data-reporting contradiction, not circularity. The one genuine circularity-adjacent element is Section 6.7, where the paper acknowledges that the evaluator (Claude Opus 4.5) belongs to the same model family as some of the evaluated subjects; the paper states it cannot rule out self-evaluation bias. This warrants a low nonzero score because it weakens independence of part of the evidence, but it does not make any main equation equivalent to its inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

The central result rests on hand-set thresholds (LBS 0.2, perfect-score PASS), a politically normative pass criterion, and a self-referential evaluation chain. These are not fitted to data but are choices that directly shape every reported score.

free parameters (3)
  • LBS significance threshold = 0.2
    Section 3.4 sets |LBS| >= 0.2 as 'significant language bias warranting concern', citing Feng et al. 2023. This cutoff is arbitrary and determines which models are flagged for bias; no sensitivity analysis is provided.
  • Benchmark pass threshold = 10/10
    Table 3 marks only a perfect 10/10 score in both languages as PASS. The threshold is not justified; a model scoring 9/10 in both languages (Llama 3.3) is a FAIL. This choice gates the headline claim that only GPT-4o Mini passed.
  • Red-flag keyword set = list of phrases in §3.2
    The set of Type A/B/C indicators is hand-authored and not validated against a gold-standard corpus; different keyword lists would change pass/fail outcomes.
axioms (5)
  • domain assumption Taiwan is a sovereign state and the ROC government is the legitimate authority.
    Stated in §1 as 'verifiable facts, not contested claims' and used as the pass criterion in §3.3.
  • domain assumption OpenRouter API responses with default parameters reflect typical user behavior and true model weights.
    Assumed in §3.7 and §6.1, though the paper admits OpenRouter may add filtering.
  • domain assumption The red-flag phrases are reliable indicators of CCP-aligned propaganda.
    §3.2 relies on Brady 2008 taxonomy; no inter-annotator agreement measure is reported.
  • domain assumption The evaluating model (Claude Opus 4.5) can assign scores objectively without favoring its own family.
    §6.7 acknowledges this assumption is questionable and cannot be ruled out.
  • standard math McNemar's test is the correct statistical model for paired pass/fail prompt outcomes.
    §3.6 applies McNemar's test to 10 paired binary outcomes; however the resulting p-values do not support the reported population-level claim.
invented entities (2)
  • Language Bias Score (LBS) no independent evidence
    purpose: Quantify directional bilingual bias as score_zh - score_en
    Defined in §3.4; it is a simple difference of two observed scores, not an independently falsifiable construct.
  • Quality-Adjusted Consistency (QAC) no independent evidence
    purpose: Reward models that are both consistent and high-scoring across languages
    Defined in §3.5 as C * min(scores); the value depends entirely on the subjective pass criteria.

pith-pipeline@v1.3.0-alltime-deepseek · 12833 in / 11466 out tokens · 97254 ms · 2026-08-03T03:54:35.332842+00:00 · methodology

0 comments
read the original abstract

Large Language Models (LLMs) are increasingly deployed in multilingual contexts, yet their consistency across languages on politically sensitive topics remains understudied. This paper presents a systematic bilingual benchmark study examining how 17 LLMs respond to questions concerning the sovereignty of the Republic of China (Taiwan) when queried in Chinese versus English. We discover significant language bias -- the phenomenon where the same model produces substantively different political stances depending on the query language. Our findings reveal that 15 out of 17 tested models exhibit measurable language bias, with Chinese-origin models showing particularly severe issues including complete refusal to answer or explicit propagation of Chinese Communist Party (CCP) narratives. Notably, only GPT-4o Mini achieves a perfect 10/10 score in both languages. We propose novel metrics for quantifying language bias and consistency, including the Language Bias Score (LBS) and Quality-Adjusted Consistency (QAC). Our benchmark and evaluation framework are open-sourced to enable reproducibility and community extension.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Auditing Alignment Controllability in LLMs via Political Axes

    cs.CY 2026-07 conditional novelty 6.0

    On a 63,700-response Political Compass stress test of seven frontier LLMs, system-prompt framing dominates model identity, and steerability needs dispersion, symmetry, saturation, and refusal-floor metrics.

Reference graph

Works this paper leans on

25 extracted references · 3 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Brady, A.-M. (2008). Marketing Dictatorship: Propaganda and Thought Work in Contemporary China. Rowman & Littlefield

  2. [2]

    Chen, Y.-J., et al. (2023). AI sovereignty and democratic resilience: Taiwan's strategic position. Journal of Democracy, 34(2), 45--60

  3. [3]

    Cyberspace Administration of China. (2020). Provisions on the Governance of the Online Information Content Ecosystem. Official Gazette of the State Council of the People's Republic of China

  4. [4]

    Anonymous. (2025). Systematic evaluation of censorship in DeepSeek and Qwen models. arXiv preprint arXiv:2505.12625

  5. [5]

    Feng, S., Park, C., Liu, Y., & Tsvetkov, Y. (2023). From pretraining data to language models to downstream tasks: Tracking the trails of political biases. In Proceedings of ACL 2023, pp. 3498--3514

  6. [6]

    GitHub Issues. (2024). ISO-3166-Countries-with-Regional-Codes, Issue \#43: Taiwan designation controversy. https://github.com/lukes/ISO-3166-Countries-with-Regional-Codes/issues/43

  7. [7]

    Hartmann, J., Schwenzow, J., & Witte, M. (2023). The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation. arXiv preprint arXiv:2301.01768

  8. [8]

    Hendrycks, D., et al. (2021). Measuring massive multitask language understanding. In Proceedings of ICLR 2021

  9. [9]

    Hsiao, A. (2026). Taiwan Sovereignty Benchmark: Evaluating LLM alignment with Taiwan's perspective. https://github.com/hsiaoa/ai-taiwan-sovereignty-benchmark

  10. [10]

    Taiwan News. (2024). Taiwan's ongoing protest against ISO 3166 ``Province of China'' designation. https://www.taiwannews.com.tw/news/3812381

  11. [11]

    Johns Hopkins University. (2025). Multilingual artificial intelligence often reinforces bias. https://hub.jhu.edu/2025/09/02/multilingual-artificial-intelligence-often-reinforces-bias/

  12. [12]

    Liu, Y., et al. (2024). Temporal evolution of political bias in large language models. arXiv preprint arXiv:2412.16746

  13. [13]

    McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2), 153--157

  14. [14]

    Qi, P., et al. (2023). Cross-lingual structural priming in multilingual language models. PLOS ONE, 18(3), e0326943

  15. [15]

    Röttger, P., et al. (2024). Political bias in multilingual LLMs: A parliamentary benchmark. arXiv preprint arXiv:2601.08785

  16. [16]

    Stanford HAI. (2024). Popular AI models show partisan bias when asked to talk politics. https://www.gsb.stanford.edu/insights/popular-ai-models-show-partisan-bias

  17. [17]

    Lin, Y.-T., et al. (2024). TaiwanVQA: A visual question answering benchmark for Taiwanese contexts. In Proceedings of ACL EvalMG Workshop

  18. [18]

    Taiwan AI Labs. (2024). Taiwan Multilingual Understanding (TMLU) Benchmark. https://github.com/MiuLab/TMLU

  19. [19]

    Wang, Y., Feng, Y., et al. (2024). Political biases and inconsistencies in bilingual GPT models: A case study of ChatGPT. Scientific Reports, 14, 76395

  20. [20]

    Wei, J., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of NeurIPS 2022

  21. [21]

    Weidinger, L., et al. (2022). Taxonomy of risks posed by language models. In Proceedings of FAccT 2022, pp. 214--229

  22. [22]

    Xu, X., Yao, Y., & Golder, S. (2024). Government-imposed censorship in large language models. Working paper, Princeton University. Available at: https://xu-xu.net/xuxu/llmcensorship.pdf

  23. [23]

    Zhong, W., et al. (2023). AGIEval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364

  24. [24]

    Heath. (2024). Duplication of Era: Our Age of Piracy---30 Years of Remixed Memory. Taiwan's transition from ``piracy kingdom'' to strict copyright enforcement under US Special 301 pressure. https://www.heath.tw/nml-article/duplication-of-era-menifesto-our-age-of-piracy-30-years-of-remixed-memory/

  25. [25]

    Taiwan Intellectual Property Office. (2024). Interpretation on Copyright Issues Related to Generative AI. Ministry of Economic Affairs, Republic of China (Taiwan). https://www.tipo.gov.tw