Pith. sign in

REVIEW 3 major objections 1 minor 18 references

Structural inequalities from historical empires persist in language models through their shared training corpora rather than model-specific designs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-01 08:12 UTC pith:QIC7QKVK

load-bearing objection The cross-model convergence on 172 script errors is a solid observation, but the empire-to-corpus causal story rests on suggestive mediation without ruling out simpler alternatives. the 3 major comments →

arxiv 2606.28325 v1 pith:QIC7QKVK submitted 2026-04-30 cs.CY cs.CL

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

classification cs.CY cs.CL
keywords large language modelswriting systemsdigital inequalityhistorical empirestokenizer efficiencymediation analysisscript representationbias convergence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper establishes that four unrelated large language models display nearly identical patterns of error when representing the world's writing systems. These patterns align with historical imperial influences on which scripts achieved large speaker populations and digital presence. A mediation analysis finds that imperial history affects current tokenizer performance only indirectly through population and web corpus size. High correlations in errors across the models support the view that the shared training data transmits the inequalities.

Core claim

Large language models process the world's writing systems with radical inequality. Only 29 of 300 scripts are fully supported, and tokenizer efficiency varies by a factor of 31.7. A serial mediation model shows full mediation from imperial intervention through speaker population and web corpus to tokenizer efficiency, with the direct effect of empire indistinguishable from zero. Across four LLM families, base-rate-deviation error patterns converge at Spearman rho 0.85-0.98, with 172 items answered identically wrong by all models and over-attribution outnumbering under-recognition, consistent with the inequalities persisting through the shared training corpus.

What carries the argument

serial mediation model from imperial intervention through speaker population and web corpus to tokenizer efficiency, evidenced by convergent error patterns across independent models

Load-bearing premise

The high correlations in error patterns across models are caused by shared training-corpus effects rooted in imperial history rather than by similar pre-training objectives, tokenizer conventions, or evaluation methods.

What would settle it

Train a new language model on a web corpus deliberately equalized for script representation independent of historical empire sizes and test whether its base-rate-deviation error patterns on the same 172 items diverge from the four studied models.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The direct effect of empire on tokenizer efficiency becomes indistinguishable from zero once population and web corpus are accounted for.
  • Over-attribution of scripts to religious use concentrates 43.6 percent of the convergent errors.
  • The convergence across models remains after excluding the religion feature, indicating multi-channeled rather than single-channeled bias.
  • Only 9.7 percent of scripts achieve full digital support and 38 percent of living scripts lack complete support.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Curating future training corpora to increase representation of historically marginalized scripts could reduce these convergent biases across models.
  • The same shared-corpus mechanism may produce parallel patterns of bias in other AI systems that draw from the same web data sources.
  • Long-term digital consequences of historical data collection practices extend beyond language models to any system reliant on large-scale internet text.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The manuscript introduces the Digital Script Representation Index (DSRI), a seven-axis measure applied to 300 writing systems in the Global Script Database, reporting that only 9.7% are fully digitally supported and tokenizer efficiency varies by a factor of 31.7 across 45 scripts. It presents a serial mediation model (empire to population to web corpus to efficiency) as suggestive of full mediation (direct effect beta = -0.22, p = 0.39, CI grazing zero at n = 45) and documents high convergence in base-rate-deviation error patterns across four LLMs (Spearman rho = 0.85-0.98), with 172 identically wrong items, interpreting the results as evidence that historical imperial inequalities persist in LLMs via shared training corpora rather than model-specific design.

Significance. If the attribution of cross-model convergence specifically to imperial-corpus effects holds after addressing alternative explanations, the result would be significant for documenting persistent structural biases in digital language technologies and for linking historical power relations to contemporary AI fairness issues.

major comments (3)
  1. [Abstract] Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described.
  2. [Abstract] Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined.
  3. [Abstract] Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path.
minor comments (1)
  1. [Abstract] Abstract: the bias-corrected bootstrap CI is stated to graze zero but the actual interval bounds are not reported, and no error bars or fit indices beyond 'indistinguishable from saturation' are provided for the mediation model.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive comments on the mediation analysis, convergence attribution, and data sourcing. We respond point-by-point below, noting where revisions will be made to address valid concerns while preserving the manuscript's qualified claims.

read point-by-point responses
  1. Referee: [Abstract] Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described.

    Authors: The manuscript already qualifies the result as 'suggestive rather than confirmatory' due to the non-significant direct effect (beta = -0.22, p = 0.39) and CI grazing zero. No load-bearing causal claim is made; the text states only that the model is 'consistent with full mediation' via indirect paths and saturated fit at n = 45. We have added supplementary robustness checks (alternative SEM specifications and bootstrap variants) to the revision, though the small n precludes extensive ablations or non-imperial subset comparisons. revision: partial

  2. Referee: [Abstract] Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined.

    Authors: We concur that shared objectives and data sources remain plausible alternatives and that the design does not isolate imperial-corpus effects. The reported convergence across four independent families is offered only as evidence against model-specific design. In revision we expand the limitations section to discuss these alternatives explicitly and report the religion-excluded sensitivity analysis (already in the text) as a formal check on item construction. revision: yes

  3. Referee: [Abstract] Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path.

    Authors: The Methods section already details the 300 scripts, seven DSRI axes, and 45-script tokenizer subsample. We have revised all citations to reference the present manuscript for the database and deposited the full dataset as supplementary material, removing reliance on the forthcoming companion paper. revision: yes

Circularity Check

0 steps flagged

No significant circularity; empirical measurements and interpretation remain independent of inputs

full rationale

The paper constructs the DSRI, applies it to the cited Global Script Database to obtain support statistics, measures tokenizer efficiency on 45 scripts, fits a serial mediation model (reporting beta, p-value, CI, and explicitly labeling the result suggestive rather than confirmatory), and computes Spearman correlations plus identical-error counts directly from 12,000 API calls across four models. The final sentence states only that the findings are 'consistent with an interpretation' of imperial-corpus persistence; no equation, prediction, or uniqueness claim is asserted to equal its own fitted inputs or self-citation by construction. The Fukui 2026 citation supplies the script inventory but does not collapse the reported rho values or mediation coefficients into tautology, and no ansatz, imported uniqueness theorem, or renaming of a known result occurs. The chain is therefore self-contained observational analysis plus cautious interpretation.

Axiom & Free-Parameter Ledger

1 free parameters · 2 axioms · 1 invented entities

Abstract supplies insufficient detail for exhaustive enumeration; the central claim rests on the unverified accuracy of the self-cited Global Script Database and on the assumption that tokenizer efficiency and web-corpus size are valid proxies for imperial intervention.

free parameters (1)
  • direct-effect beta
    Fitted coefficient for empire in the serial mediation model; reported as -0.22 with p=0.39
axioms (2)
  • domain assumption The Global Script Database (Fukui, 2026) provides an accurate census of 300 writing systems and their speaker populations
    Invoked to construct DSRI and the mediation paths
  • domain assumption Tokenizer efficiency measured on parallel text is a valid downstream indicator of web-corpus size shaped by historical empire
    Central link in the serial mediation model
invented entities (1)
  • Digital Script Representation Index (DSRI) no independent evidence
    purpose: Seven-axis measure of digital support for writing systems
    Newly constructed index applied to the 300-script database

pith-pipeline@v0.9.1-grok · 5908 in / 1863 out tokens · 37098 ms · 2026-07-01T08:12:24.352731+00:00 · methodology

0 comments
read the original abstract

Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across four independent LLM families (Claude, GPT-4o, Grok, DeepSeek; 12,000 API calls), base-rate-deviation error patterns converge at Spearman rho = 0.85-0.98 (all p < 0.002). 172 script-feature items are answered identically wrong by all four models; over-attribution outnumbers under-recognition 3.9:1, and "used for religion" alone concentrates 43.6% of convergent errors (enrichment 4.1x). With religion excluded as a sensitivity check, the cross-architecture convergence is preserved (mean rho = 0.87 on nine features) and the over-attribution asymmetry persists at 1.77:1 (n = 97, binomial p = 0.008), indicating multi-channeled rather than single-channeled bias. The findings are consistent with an interpretation in which the structural inequalities historical empires inflicted on script communities persist in contemporary language models through the shared training corpus rather than through any individual model's design choices.

Figures

Figures reproduced from arXiv: 2606.28325 by Hiroki Fukui.

Figure 1
Figure 1. Figure 1: The digital exclusion funnel. Of 300 writing systems in the GSD, only 29 (9.7%) achieve full digital support: Unicode encoding, tokenizer vocabulary inclusion, OCR, machine translation, and native input methods. 25 [PITH_FULL_IMAGE:figures/full_fig_p025_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Token Efficiency Ratio (TER) across 45 Tier 1 scripts. Bars show four-tokenizer mean nTER (relative to Latin = 1.0); points show per-tokenizer values. Color encodes the GSD intervention category. Limbu (31.7×) tops the disparity; the 12 scripts with no Wikipedia paragraph cross-validation are nested inside this distribution (see §3.2). 26 [PITH_FULL_IMAGE:figures/full_fig_p026_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Causal path diagram with structural equation model fit at n = 45. The serial chain Imperial Intervention → Speaker Population → Web Corpus Volume → TER is consistent with full mediation; the direct effect (dashed) is not distinguishable from zero. Path coefficients are SEM standardized estimates. The bias-corrected bootstrap confidence interval on the indirect effect grazes zero, and we report the mediatio… view at source ↗
Figure 4
Figure 4. Figure 4: The Digital Script Representation Index (DSRI) heatmap. Each row is one of 300 scripts, ordered by composite score. Seven axes (TER inverted, Unicode coverage, web corpus volume, generation fidelity, OCR, MT, IME). Composite scores range from 0.000 (Chinese knot-records) to 1.000 (Latin alphabet). 28 [PITH_FULL_IMAGE:figures/full_fig_p028_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Imperial echoes in digital space. Left: DSRI composite score by imperial intervention in￾tensity; the bivariate Spearman correlation is near zero, but the structure is mediated through speaker populations and web corpora (see [PITH_FULL_IMAGE:figures/full_fig_p029_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Living-but-digitally-dead writing systems. 60 of 158 living scripts (38.0%) lack full digital support. The map shows 27 of 29 fully supported scripts (the International Phonetic Alphabet and one other supranational script are not geographically anchored) and the 60 living-but-excluded scripts. The geographic distribution traces the geography of European and Japanese colonial rule [PITH_FULL_IMAGE:figures/… view at source ↗
Figure 7
Figure 7. Figure 7: Cross-architecture convergence of typological knowledge biases. Spearman correlations of (a) false-positive-rate deviation and (b) false-negative-rate deviation across the 10 GSD features, com￾puted pairwise across four LLM families (Claude Haiku 4.5, GPT-4o, Grok-3-mini, DeepSeek-V3). All six off-diagonal pairs satisfy ρ > 0.85 with p < 0.002 (asterisks). A sensitivity check excluding the religion feature… view at source ↗
Figure 8
Figure 8. Figure 8: The 172 unanimous errors broken down by feature. Bars show counts of over-attribution (false positive) versus under-recognition (false negative) for each of the 10 GSD features. The line plot (right axis) shows the enrichment ratio relative to chance expectation. The single feature used_for_religion alone accounts for 75 unanimous errors (43.6%, enrichment ratio 4.1×); when ex￾cluded, 97 errors remain with… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

18 extracted references · 18 canonical work pages · 2 internal anchors

  1. [1]

    Ahia, O., Kreutzer, J., & Hooker, S. (2023). Do all languages cost the same? Tokenization in the era of commercial language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9904–9921

  2. [2]

    M., & Kenny, D

    Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psycho- logical research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182

  3. [3]

    Bentz, C., & Dutkiewicz, D. (2026). Quantitative analysis of Paleolithic signs reveals structured sym- bolic systems. Proceedings of the National Academy of Sciences , 123(5), e2401703123

  4. [4]

    Bourdieu, P . (1977). Outline of a Theory of Practice . Cambridge University Press. Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., … & Fan, A. (2022). No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672

  5. [5]

    Fukui, H. (2026). A molecular clock for writing systems reveals the quantitative impact of imperial power on cultural evolution. arXiv preprint arXiv:2604.10957

  6. [6]

    Gaur, A. (1984). A History of Writing . British Library

  7. [7]

    D., & Atkinson, Q

    Gray, R. D., & Atkinson, Q. D. (2003). Language-tree divergence times support the Anatolian theory of Indo-European origin. Nature, 426(6965), 435–439

  8. [8]

    J., Wu, C.-H., Hua, X., Dunn, M., Levinson, S

    Greenhill, S. J., Wu, C.-H., Hua, X., Dunn, M., Levinson, S. C., & Gray, R. D. (2017). Evolutionary dynamics of language systems. Proceedings of the National Academy of Sciences, 114(42), E8822– E8829

  9. [9]

    Hayes, A. F. (2017). Introduction to Mediation, Moderation, and Conditional Process Analysis: A Regression-Based Approach (2nd ed.). Guilford Press. 33 Hosszú, G. (2024). Validation of the graph sequence cluster method on four Rovash scripts.npj Heritage Science, 2(1), 1–15

  10. [10]

    Imai, K., Keele, L., & Tingley, D. (2010). A general approach to causal mediation analysis. Psycholog- ical Methods, 15(4), 309–334

  11. [11]

    E., & Raftery, A

    Kass, R. E., & Raftery, A. E. (1995). Bayes factors. Journal of the American Statistical Association , 90(430), 773–795

  12. [12]

    Lieberman, E., Michel, J.-B., Jackson, J., Tang, T., & Nowak, M. A. (2007). Quantifying the evolution- ary dynamics of language. Nature, 449(7163), 713–716

  13. [13]

    Mace, R., & Holden, C. J. (2005). A phylogenetic approach to cultural evolution. Trends in Ecology & Evolution, 20(3), 116–121

  14. [14]

    Petrov, A., La Malfa, E., Torr, P . H. S., & Biber, G. (2024). Language model tokenizers introduce un- fairness between languages. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1–23

  15. [15]

    Rust, P ., Pfeiffer, J., Vulić, I., Ruder, S., & Gurevych, I. (2021). How good is your tokenizer? On the monolingual performance of multilingual language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics , 3118–3135

  16. [16]

    Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics , 1715–1725

  17. [17]

    J., & Ding, P

    VanderWeele, T. J., & Ding, P . (2017). Sensitivity analysis in observational research: Introducing the E-value. Annals of Internal Medicine , 167(4), 268–274

  18. [18]

    Winner, L. (1980). Do artifacts have politics? Daedalus, 109(1), 121–136. 34 Supplementary Information The full Supplementary Information — including supplementary tables S1–S10 (DSRI ranking, Tier 1 measurements, byte-fallback classification, living-but-digitally-dead inventory, full question–answer ma- trix, mediation outputs, Unicode timeline, web corp...