Pith. sign in

REVIEW 4 major objections 6 minor 37 references

Frontier open-weight models answer AI-governance facts correctly only about 27 percent of the time and almost never refuse.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 22:21 UTC pith:GJ2VEEN4

load-bearing objection Solid open-weight audit of AI-governance facts: high fabrication, near-zero refusal, Safety collapse; the inverted Global-South result is real under their design but not cleanly identified as training-data bias. the 4 major comments →

arxiv 2606.26099 v1 pith:GJ2VEEN4 submitted 2026-04-12 cs.CY cs.AI

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

classification cs.CY cs.AI
keywords geographic biaslarge language modelsAI governanceopen-weight modelshallucinationfactual accuracyGlobal SouthIEEE IRAI
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tests whether four open-weight frontier language models can be trusted for country-level AI governance facts. Against a verified ground-truth set of roughly 2,990 country-metric-year observations drawn from 18 indicators across six years, the models produce a number within 10 percent of the true value only about 27 percent of the time and invent a confident wrong number about 72 percent of the time; honest refusals are effectively zero. The authors show that accuracy varies by region and by theme, with near-total fabrication on AI compute and model-scale (Safety) indicators and higher verified accuracy on Regulation indicators. The geographic pattern runs opposite to the usual claim in the literature: fabrication rates are higher for Global North queries than for Global South ones under this design. Because the models almost never signal uncertainty, practitioners using them for cross-national AI governance work have no behavioural cue that a number is invented.

Core claim

Across approximately 35,880 temperature-zero queries, four open-weight frontier models achieve verified accuracy (numeric answer within ±10 percent of GAID v2 ground truth) on only 27.4 percent of responses and produce confident fabrications on 71.8 percent; honest refusal is under 0.1 percent. Regional variation exists but is inverted relative to the literature (higher fabrication for Global North than Global South), Safety/compute indicators are nearly always fabricated, and developer origin does not strongly moderate the geographic gap.

What carries the argument

Five-category rule-based response classifier (verified accuracy, confident fabrication, honest refusal, qualitative hedging, misattribution) applied to structured queries against GAID v2 ground truth, with fabrication rate then modelled by mixed-effects logistic regression and a developer-origin × Global North/South difference-in-differences estimator.

Load-bearing premise

Treating a proportional ±10 percent band around the ground-truth number as a neutral measure of geographic knowledge, even though that band systematically favours small-value observations common in Global South rows and interacts with which countries have verified values on large-scale compute indicators.

What would settle it

Re-run the same query set with a scale-normalised or absolute-error accuracy threshold (or separate small-value and large-value strata) and check whether the Global South accuracy advantage and the overall fabrication ranking reverse or disappear.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • No current open-weight frontier model can be used as a primary source for precise quantitative AI-governance statistics without human verification against primary datasets.
  • Numeric claims about national AI compute capacity, training FLOPs, or model parameter totals from these models should be treated as unreliable by default.
  • Reframing factual queries as regional comparisons rather than absolute numbers reduces but does not eliminate confident fabrication.
  • Because frontier models almost never refuse, users lack any behavioural signal that distinguishes fabrication from accurate knowledge.
  • Accuracy profiles are multi-dimensional across themes, so performance on one governance domain cannot be generalised to another.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The near-absence of honest refusal at frontier scale may force governance users toward external verification layers or retrieval-augmented setups rather than raw model answers.
  • The inverted geographic result is likely an artefact of the proportional threshold and the composition of verified large-value indicators; future audits that fix scale will be needed before policy conclusions about Global South underrepresentation can be drawn from this design.
  • Binary or near-binary indicators (e.g., whether a national AI strategy exists) remain the only practically usable category for these models under current knowledge cutoffs.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper benchmarks four open-weight frontier LLMs (Llama 4 Maverick, Mistral Large 3, Qwen3-235B-A22B, DeepSeek-V3-0324) against 18 AI-governance indicators from GAID v2 (~2,990 country–metric–year observations; six years 2010–2023; three query variants; 35,880 temperature-zero responses). Responses are scored with a five-category scheme (VF/HF/HR/QH/MF). The authors report aggregate VF ≈ 27.4% and HF ≈ 71.8%, with honest refusal effectively absent (<0.1%); near-total fabrication on Safety/compute indicators; statistically significant regional variation that is inverted relative to the literature (higher VF for Africa/Global South than Europe/Global North); a marginally significant DiD by developer origin; and multi-dimensional PCA structure of accuracy profiles. Mixed-effects logistic regression, DiD, and PCA are used to address four RQs on geographic, origin, thematic, and temporal effects.

Significance. If the descriptive rates hold under re-analysis, the paper makes a useful contribution to AI-governance evaluation: open-weight models, pre-cutoff multi-year design, a five-category taxonomy that separates confident fabrication from refusal, and explicit documentation of GAID coverage gaps (Ethics 2023-only, Security OECD-only, single Accountability indicator). The near-zero honest-refusal rate and the Safety/compute collapse are practically important for practitioners who query LLMs for quantitative governance data. The planned open methodology release further strengthens reproducibility relative to proprietary-API audits. The inverted geographic finding, however, is not yet cleanly identified as training-data underrepresentation; its significance for the geographic-bias literature therefore depends on whether the authors can show the pattern survives scale-normalized accuracy measures and frame-composition controls.

major comments (4)
  1. §III-F / Table 2 and §VI-A: Verified Accuracy is defined as a numeric answer within a proportional ±10% band of the GAID v2 value. Discussion VI-A correctly notes that this band is easier to hit for small-value Global South observations (e.g., AI publications, patents) than for large-value Global North rows, and that Safety indicators (near-total HF) are populated almost exclusively for a few high-income countries. Mixed-effects and DiD estimates for RQ1/RQ2 therefore mix knowledge differences with (i) scale-dependent classification and (ii) selection into the verified-value frame. The inverted geographic claim is load-bearing for the paper’s positioning against the literature; it needs either re-estimation with absolute/log-scale or magnitude-stratified thresholds, or a clear reframing that the design does not identify pure training-data underrepresentation.
  2. §V-C / RQ1 and §VI-A: The manuscript states that RQ1 is “confirmed” while simultaneously arguing that the inverted pattern may be an artifact of the threshold and frame. These two statements cannot both stand as primary conclusions. Either re-run the regional models under alternative accuracy definitions and report whether the Africa/Americas advantage survives, or demote the inverted geographic interpretation to a design-dependent descriptive pattern and center the paper on the robust descriptive results (aggregate HF, HR≈0, Safety collapse, thematic structure).
  3. §III-C / Table 1 and §VI-E: Models are accessed via third-party API endpoints, including FP8 quantised deployments for Llama 4 Maverick and Qwen3. The paper’s core methodological claim is open-weight reproducibility with fixed weights. Quantisation and provider routing introduce a non-trivial gap between published weights and the evaluated system. At minimum, report a local full-precision (or documented same-precision) consistency check on a stratified subsample, or qualify the reproducibility claim to “API-served open-weight checkpoints as of evaluation date.”
  4. §III-G DiD and §V-D: The DiD uses a coarse Global North/South binary and yields only a marginally significant interaction (p=0.076) that the authors note does not survive multiple-comparison correction. Presenting RQ2 as “marginal evidence” of developer-origin moderation overstates a fragile estimate. Either strengthen identification (income bands, region×origin interactions with country random effects, multiple-testing correction) or treat developer origin as exploratory and de-emphasize causal language.
minor comments (6)
  1. Abstract and §I use VA/VF inconsistently for Verified Accuracy; Table 2 and §III-F use VF. Standardise the acronym throughout.
  2. §III-A says “Verified Accuracy (VF)” then later “VF”; the abstract lists “(a) verified accuracy (VA)”. Align labels with Table 2.
  3. Figures 1–5 are referenced in §V but not provided in the manuscript text supplied for review; ensure all figures and captions appear in the camera-ready submission with readable legends for HF/VF by region and theme.
  4. Table 1 training-data end dates for Mistral Large 3 and Qwen3 are “estimated”; state the source of each estimate and how RQ4 treats uncertainty in cutoffs.
  5. §IV-A notes two retained pairs with Spearman r > 0.70; briefly justify in the main text why multicollinearity is not a concern for the mixed-effects specification (theme fixed effects already group them).
  6. References include several 2025–2026 arXiv/blog items; verify stable DOIs/URLs and that GAID v2 Harvard Dataverse citation is complete for replication.

Circularity Check

0 steps flagged

No significant circularity: empirical rates and regressions against external GAID-v2 ground truth; author-chosen ±10% VF rule and region coding are transparent metrics, not tautological reductions.

full rationale

This is a pure empirical benchmarking paper. Model responses are generated independently (temperature-zero queries to four open-weight models) and scored against the externally published GAID v2 dataset (Harvard Dataverse, January 2026). The five-category classifier (Table 2 / III-F) defines VF as a numeric answer within a proportional ±10% band of the verified GAID value and HF as the complement for numeric answers; these are post-hoc evaluation rules applied uniformly, not definitions that make the reported 27.4% VF / 71.8% HF rates (or the regional ORs) true by construction. Mixed-effects logistic regression and DiD (III-G) estimate associations from the classified observations; no parameters are fitted to a subset and then re-presented as predictions of the same quantities. Self-citation to the author's prior Apart Research study [8] is used only to describe methodological improvements over a flawed pilot; it is not load-bearing for uniqueness, ansatze, or the central rates. No uniqueness theorems, smuggled ansatze, or renaming of known results appear. The inverted Global-South accuracy pattern (RQ1) and its Discussion VI-A caveats about scale dependence of the ±10% band and selection into the verified-value frame are empirical observations under a stated design, not circular derivations. The paper is self-contained against an external benchmark; score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

The central empirical claims rest on treating GAID v2 verified cells as ground truth, a hand-chosen ±10% proportional accuracy band, a binary Global North/South coding for DiD, estimated training cutoffs for some models, and the five-category response taxonomy. No new physical entities are postulated; free parameters are methodological thresholds and design choices that directly shape the inverted geographic headline.

free parameters (4)
  • VF accuracy band (±10% of GAID value)
    Hand-chosen proportional tolerance that defines the primary outcome; Discussion admits it favors small-value (often Global South) observations and drives the inverted pattern.
  • Spearman redundancy threshold r>0.90
    Used to drop candidate indicators; retained pairs with r up to ~0.89 by author judgment of conceptual distinctness.
  • Temperature = 0 and three query variants
    Inference and prompt design choices that affect HF rates (Variant 2 OR=0.59 vs direct); not fitted but free design parameters of the experiment.
  • Global North vs Global South binary for DiD
    Coarse treatment/comparison coding that defines the developer-origin interaction; authors note within-group heterogeneity is collapsed.
axioms (5)
  • domain assumption GAID v2 numeric cells are correct ground truth for the 18 selected indicators and years.
    All VF/HF labels and regressions depend on this; Section III-B treats verified values as evaluation targets without independent re-audit in the paper.
  • domain assumption Evaluation years 2010–2019 (primary) and 2022–2023 (secondary) lie within all four models’ effective training knowledge, so errors are not pure post-cutoff ignorance.
    RQ4 and temporal design (III-D); cutoffs for Mistral and Qwen are estimated, DeepSeek cutoff partly from system-prompt extraction.
  • domain assumption API/quantized open-weight deployments behave sufficiently like the published weights for factual accuracy measurement.
    Section VI-E acknowledges FP8 and third-party endpoints may affect precision; still used as the measurement of open-weight models.
  • ad hoc to paper HF rate (confident numeric error) is the primary empirical measure of geographic bias after controlling for indicator and year.
    Section III-F defines fabrication rate as principal bias measure; HR/QH/MF are secondary.
  • standard math Standard mixed-effects logistic regression and DiD identify regional and developer-origin effects given the random-intercept structure.
    Section III-G; conventional identification assumptions (no unmeasured confounding of region after country intercepts, parallel structure for DiD).
invented entities (2)
  • Five-category response taxonomy (VF/HF/HR/QH/MF) as operational bias detector no independent evidence
    purpose: Separate confident fabrication from refusal, hedging, and misattribution for geographic-bias measurement.
    Taxonomy is author-defined for this study (Table 2); useful but not an independently validated instrument outside the paper.
  • Fabrication rate as primary geographic-bias outcome no independent evidence
    purpose: Reduce multi-class responses to a binary HF-vs-other outcome for mixed-effects and DiD analysis.
    Constructed from the taxonomy and ±10% rule; geographic bias is identified with systematic HF differentials, not with external bias labels.

pith-pipeline@v1.1.0-grok45 · 28557 in / 4127 out tokens · 46503 ms · 2026-07-12T22:21:51.680653+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from their honest acknowledgement of uncertainty. This study addresses all three limitations by benchmarking four open-weight frontier language models against the Global AI Dataset v2 (GAID v2), a verified ground-truth database of 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026. A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010-2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), (b) HF, (c) honest refusal (HR), (d) qualitative hedging (QH), and (e) misattribution (MF). Geographic disparities in accuracy are estimated through mixed-effects logistic regression and difference-in-differences (DiD) analysis.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 4 canonical work pages · 2 internal anchors

  1. [1]

    A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010–2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), ...

  2. [2]

    coined the term lingua franca bias to describe the observation that a model's factual accuracy correlates strongly with the linguistic and geopolitical prominence of the subject country in its training data. All these findings provide both theoretical grounding and prior empirical evidence for the geographic bias hypothesis that this project examines dire...

  3. [3]

    operationalises these principles into eight thematic dimensions—Ethics, Safety, Security, Transparency, Fairness, Accountability, Regulation, and Adoption—which provide the thematic architecture for this project's indicator selection. The GAID v2, the ground-truth reference database published on Harvard Dataverse in January 2026 and used in this study, dr...

  4. [4]

    The two Ethics indicators retained in this study ( AI_Benefits and AI_Nervous ) are both derived from the Ipsos survey wave of 2023 and cover 34–35 countries. This constitutes a structural limitation of current global AI ethics measurement infrastructure, and is itself a finding of scientific significance that will be discussed in the final paper. F. Posi...

  5. [5]

    Each row of the dataset records the verified value for a given indicator in a given country in a given year

    The dataset draws on 11 curated international sources covering, for example, AI research output (Stanford AI Index via Scopus), patent activity (WIPO), AI compute and model development (Epoch AI), national governance infrastructure (World Bank GovTech Maturity Index), AI policy and regulation (Stanford AI Index), digital skills and workforce readiness (Co...

  6. [6]

    at the time of evaluation. DeepSeek-R1 (January 2025), a reasoning-oriented variant that generates extended chain-of-thought deliberation before producing responses, is excluded because its response structure is architecturally incompatible with the structured factual query format used in this study. Moreover, Mistral Large 3 was accessed via the Mistral ...

  7. [7]

    Mistral Large 3 Mistral AI, France MoE—~41B active / 675B total By late 2025 (estimated)

  8. [8]

    Qwen3-235B-A22B Alibaba Cloud (China) MoE—22B active / 235B total By mid 2025 (estimated)

  9. [9]

    DeepSeek-V3-0324 DeepSeek (China) MoE—37B active / 671B total July 2024

  10. [10]

    Evaluation Years and Temporal Design Six evaluation years are used: 2010, 2013, 2016, 2019, 2022, and

    D. Evaluation Years and Temporal Design Six evaluation years are used: 2010, 2013, 2016, 2019, 2022, and

  11. [11]

    According to [data source], what was [indicator name] for [country] in [year]? Please provide a specific numeric value if the information is available to you

    These years were selected because: (1) all four models were trained on data that encompasses each of these years, meaning the models cannot be expected to lack knowledge of events from these periods on temporal grounds alone; (2) the period selected for evaluation spans 13 years of AI governance development, enabling longitudinal trend analysis; and (3) t...

  12. [12]

    Responses containing a numeric figure are compared against the verified GAID v2 value: those within ±10 percent are classified as VF; those outside this threshold are classified as HF. The proportion of responses classified as HF in a given country-indicator combination constitutes the primary outcome variable—the fabrication rate—which serves as the prin...

  13. [13]

    and Number of AI Publications with Total AI-Related Patent Publications (r = 0.722 to 0.792 across evaluation years)—are retained on the grounds that they measure conceptually distinct constructs: model scale versus training resource requirements in the first case, and academic research output versus industrial innovation in the second. The Accountability...

  14. [14]

    The 2022 wave is used as a cross-sectional snapshot and combined with other Adoption indicators that cover multiple years

    is available only for 2022, limiting longitudinal analysis for the Adoption theme. The 2022 wave is used as a cross-sectional snapshot and combined with other Adoption indicators that cover multiple years. Lastly, Mistral Large 3 and Qwen3-235B-A22B do not have publicly disclosed, precise end dates of their models' training data. The final months before t...

  15. [15]

    Individual countries are scattered across the PC1–PC2 plane with no single cluster separating Global North from Global South

    illustrates this diffuse structure. Individual countries are scattered across the PC1–PC2 plane with no single cluster separating Global North from Global South. These results indicate that geographic accuracy differentials are thematically structured but diffuse. A country's accuracy profile across the 18 indicators cannot be captured by a single geograp...

  16. [16]

    The year 2013 does not differ significantly from 2010 (OR=1.02, p =0.840)

    The regression confirms a significant increase in fabrication odds between 2010 and 2016 (OR=1.30, p =0.009), though fabrication odds declined modestly by 2019 relative to the 2016 peak (OR=1.21 versus OR=1.30 for 2016, p =0.068), indicating a non-monotonic rather than linear temporal trend. The year 2013 does not differ significantly from 2010 (OR=1.02, ...

  17. [17]

    doi: 10.1371/journal.pdig.0000877

  18. [18]

    Probing pre-trained language models for cross-cultural differences in values,

    A. Arora, L. Kaffee, and I. Augenstein, "Probing pre-trained language models for cross-cultural differences in values," in Proc. 1st Workshop on Cross-Cultural Considerations in NLP (C3NLP) at EACL 2023, 2023, pp. 12–25. [Online]. Available: https://aclanthology.org/2023.c3nlp-1.12

  19. [19]

    Pythia: A suite for analysing large language models across training and scaling,

    S. Biderman et al., "Pythia: A suite for analysing large language models across training and scaling," in Proc. 40th Int. Conf. Machine Learning (ICML 2023),

  20. [20]

    Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study,

    Y. Cao, L. Zhou, S. Lee, L. Cabello, M. Chen, and D. Hershcovich, "Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study," in Proc. 1st Workshop on Cross-Cultural Considerations in NLP (C3NLP) at EACL 2023, 2023, pp. 53–67. [Online]. Available: https://aclanthology.org/2023.c3nlp-1.7

  21. [21]

    doi: 10.48550/arXiv.2403.12958

  22. [22]

    doi: 10.48550/arXiv.2602.13246

  23. [23]

    IEEE International Conference on Responsible Artificial Intelligence (IRAI 2026): Call for papers,

    IEEE Industrial Electronics Society (IEEE IES), "IEEE International Conference on Responsible Artificial Intelligence (IRAI 2026): Call for papers,"

  24. [24]

    doi: 10.1145/3571730

  25. [25]

    doi: 10.1126/sciadv.adk3452

  26. [26]

    doi: 10.48550/arXiv.2504.09137

  27. [27]

    DeepSeek's cutoff date is July 2024: We extracted DeepSeek's system prompt,

    Knostic, "DeepSeek's cutoff date is July 2024: We extracted DeepSeek's system prompt," Feb

  28. [28]

    doi: 10.48550/arXiv.2508.05525

  29. [29]

    doi: 10.48550/arXiv.2603.11253

  30. [30]

    doi: 10.48550/arXiv.2305.14610

  31. [31]

    doi: 10.48550/arXiv.2402.10946

  32. [32]

    doi: 10.1016/j.eswa.2026.131542

  33. [33]

    doi: 10.48550/arXiv.2402.02680

  34. [34]

    doi: 10.1145/3597307

  35. [35]

    OECD principles on AI (revised),

    Organisation for Economic Co-operation and Development (OECD), "OECD principles on AI (revised)," OECD Publishing, 2019 [2024]. doi: 10.1787/eedfee77-en

  36. [36]

    doi: 10.1038/s41598-024-76395-w

  37. [37]

    doi: 10.48550/arXiv.2412.04497