REVIEW 3 major objections 1 minor 18 references
Structural inequalities from historical empires persist in language models through their shared training corpora rather than model-specific designs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-01 08:12 UTC pith:QIC7QKVK
load-bearing objection The cross-model convergence on 172 script errors is a solid observation, but the empire-to-corpus causal story rests on suggestive mediation without ruling out simpler alternatives. the 3 major comments →
The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Large language models process the world's writing systems with radical inequality. Only 29 of 300 scripts are fully supported, and tokenizer efficiency varies by a factor of 31.7. A serial mediation model shows full mediation from imperial intervention through speaker population and web corpus to tokenizer efficiency, with the direct effect of empire indistinguishable from zero. Across four LLM families, base-rate-deviation error patterns converge at Spearman rho 0.85-0.98, with 172 items answered identically wrong by all models and over-attribution outnumbering under-recognition, consistent with the inequalities persisting through the shared training corpus.
What carries the argument
serial mediation model from imperial intervention through speaker population and web corpus to tokenizer efficiency, evidenced by convergent error patterns across independent models
Load-bearing premise
The high correlations in error patterns across models are caused by shared training-corpus effects rooted in imperial history rather than by similar pre-training objectives, tokenizer conventions, or evaluation methods.
What would settle it
Train a new language model on a web corpus deliberately equalized for script representation independent of historical empire sizes and test whether its base-rate-deviation error patterns on the same 172 items diverge from the four studied models.
If this is right
- The direct effect of empire on tokenizer efficiency becomes indistinguishable from zero once population and web corpus are accounted for.
- Over-attribution of scripts to religious use concentrates 43.6 percent of the convergent errors.
- The convergence across models remains after excluding the religion feature, indicating multi-channeled rather than single-channeled bias.
- Only 9.7 percent of scripts achieve full digital support and 38 percent of living scripts lack complete support.
Where Pith is reading between the lines
- Curating future training corpora to increase representation of historically marginalized scripts could reduce these convergent biases across models.
- The same shared-corpus mechanism may produce parallel patterns of bias in other AI systems that draw from the same web data sources.
- Long-term digital consequences of historical data collection practices extend beyond language models to any system reliant on large-scale internet text.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Digital Script Representation Index (DSRI), a seven-axis measure applied to 300 writing systems in the Global Script Database, reporting that only 9.7% are fully digitally supported and tokenizer efficiency varies by a factor of 31.7 across 45 scripts. It presents a serial mediation model (empire to population to web corpus to efficiency) as suggestive of full mediation (direct effect beta = -0.22, p = 0.39, CI grazing zero at n = 45) and documents high convergence in base-rate-deviation error patterns across four LLMs (Spearman rho = 0.85-0.98), with 172 identically wrong items, interpreting the results as evidence that historical imperial inequalities persist in LLMs via shared training corpora rather than model-specific design.
Significance. If the attribution of cross-model convergence specifically to imperial-corpus effects holds after addressing alternative explanations, the result would be significant for documenting persistent structural biases in digital language technologies and for linking historical power relations to contemporary AI fairness issues.
major comments (3)
- [Abstract] Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described.
- [Abstract] Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined.
- [Abstract] Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path.
minor comments (1)
- [Abstract] Abstract: the bias-corrected bootstrap CI is stated to graze zero but the actual interval bounds are not reported, and no error bars or fit indices beyond 'indistinguishable from saturation' are provided for the mediation model.
Simulated Author's Rebuttal
We thank the referee for the constructive comments on the mediation analysis, convergence attribution, and data sourcing. We respond point-by-point below, noting where revisions will be made to address valid concerns while preserving the manuscript's qualified claims.
read point-by-point responses
-
Referee: [Abstract] Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described.
Authors: The manuscript already qualifies the result as 'suggestive rather than confirmatory' due to the non-significant direct effect (beta = -0.22, p = 0.39) and CI grazing zero. No load-bearing causal claim is made; the text states only that the model is 'consistent with full mediation' via indirect paths and saturated fit at n = 45. We have added supplementary robustness checks (alternative SEM specifications and bootstrap variants) to the revision, though the small n precludes extensive ablations or non-imperial subset comparisons. revision: partial
-
Referee: [Abstract] Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined.
Authors: We concur that shared objectives and data sources remain plausible alternatives and that the design does not isolate imperial-corpus effects. The reported convergence across four independent families is offered only as evidence against model-specific design. In revision we expand the limitations section to discuss these alternatives explicitly and report the religion-excluded sensitivity analysis (already in the text) as a formal check on item construction. revision: yes
-
Referee: [Abstract] Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path.
Authors: The Methods section already details the 300 scripts, seven DSRI axes, and 45-script tokenizer subsample. We have revised all citations to reference the present manuscript for the database and deposited the full dataset as supplementary material, removing reliance on the forthcoming companion paper. revision: yes
Circularity Check
No significant circularity; empirical measurements and interpretation remain independent of inputs
full rationale
The paper constructs the DSRI, applies it to the cited Global Script Database to obtain support statistics, measures tokenizer efficiency on 45 scripts, fits a serial mediation model (reporting beta, p-value, CI, and explicitly labeling the result suggestive rather than confirmatory), and computes Spearman correlations plus identical-error counts directly from 12,000 API calls across four models. The final sentence states only that the findings are 'consistent with an interpretation' of imperial-corpus persistence; no equation, prediction, or uniqueness claim is asserted to equal its own fitted inputs or self-citation by construction. The Fukui 2026 citation supplies the script inventory but does not collapse the reported rho values or mediation coefficients into tautology, and no ansatz, imported uniqueness theorem, or renaming of a known result occurs. The chain is therefore self-contained observational analysis plus cautious interpretation.
Axiom & Free-Parameter Ledger
free parameters (1)
- direct-effect beta
axioms (2)
- domain assumption The Global Script Database (Fukui, 2026) provides an accurate census of 300 writing systems and their speaker populations
- domain assumption Tokenizer efficiency measured on parallel text is a valid downstream indicator of web-corpus size shaped by historical empire
invented entities (1)
-
Digital Script Representation Index (DSRI)
no independent evidence
read the original abstract
Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across four independent LLM families (Claude, GPT-4o, Grok, DeepSeek; 12,000 API calls), base-rate-deviation error patterns converge at Spearman rho = 0.85-0.98 (all p < 0.002). 172 script-feature items are answered identically wrong by all four models; over-attribution outnumbers under-recognition 3.9:1, and "used for religion" alone concentrates 43.6% of convergent errors (enrichment 4.1x). With religion excluded as a sensitivity check, the cross-architecture convergence is preserved (mean rho = 0.87 on nine features) and the over-attribution asymmetry persists at 1.77:1 (n = 97, binomial p = 0.008), indicating multi-channeled rather than single-channeled bias. The findings are consistent with an interpretation in which the structural inequalities historical empires inflicted on script communities persist in contemporary language models through the shared training corpus rather than through any individual model's design choices.
Figures
Reference graph
Works this paper leans on
-
[1]
Ahia, O., Kreutzer, J., & Hooker, S. (2023). Do all languages cost the same? Tokenization in the era of commercial language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9904–9921
work page 2023
-
[2]
Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psycho- logical research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182
work page 1986
-
[3]
Bentz, C., & Dutkiewicz, D. (2026). Quantitative analysis of Paleolithic signs reveals structured sym- bolic systems. Proceedings of the National Academy of Sciences , 123(5), e2401703123
work page 2026
-
[4]
Bourdieu, P . (1977). Outline of a Theory of Practice . Cambridge University Press. Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., … & Fan, A. (2022). No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672
work page internal anchor Pith review Pith/arXiv arXiv 1977
-
[5]
Fukui, H. (2026). A molecular clock for writing systems reveals the quantitative impact of imperial power on cultural evolution. arXiv preprint arXiv:2604.10957
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[6]
Gaur, A. (1984). A History of Writing . British Library
work page 1984
-
[7]
Gray, R. D., & Atkinson, Q. D. (2003). Language-tree divergence times support the Anatolian theory of Indo-European origin. Nature, 426(6965), 435–439
work page 2003
-
[8]
J., Wu, C.-H., Hua, X., Dunn, M., Levinson, S
Greenhill, S. J., Wu, C.-H., Hua, X., Dunn, M., Levinson, S. C., & Gray, R. D. (2017). Evolutionary dynamics of language systems. Proceedings of the National Academy of Sciences, 114(42), E8822– E8829
work page 2017
-
[9]
Hayes, A. F. (2017). Introduction to Mediation, Moderation, and Conditional Process Analysis: A Regression-Based Approach (2nd ed.). Guilford Press. 33 Hosszú, G. (2024). Validation of the graph sequence cluster method on four Rovash scripts.npj Heritage Science, 2(1), 1–15
work page 2017
-
[10]
Imai, K., Keele, L., & Tingley, D. (2010). A general approach to causal mediation analysis. Psycholog- ical Methods, 15(4), 309–334
work page 2010
-
[11]
Kass, R. E., & Raftery, A. E. (1995). Bayes factors. Journal of the American Statistical Association , 90(430), 773–795
work page 1995
-
[12]
Lieberman, E., Michel, J.-B., Jackson, J., Tang, T., & Nowak, M. A. (2007). Quantifying the evolution- ary dynamics of language. Nature, 449(7163), 713–716
work page 2007
-
[13]
Mace, R., & Holden, C. J. (2005). A phylogenetic approach to cultural evolution. Trends in Ecology & Evolution, 20(3), 116–121
work page 2005
-
[14]
Petrov, A., La Malfa, E., Torr, P . H. S., & Biber, G. (2024). Language model tokenizers introduce un- fairness between languages. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1–23
work page 2024
-
[15]
Rust, P ., Pfeiffer, J., Vulić, I., Ruder, S., & Gurevych, I. (2021). How good is your tokenizer? On the monolingual performance of multilingual language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics , 3118–3135
work page 2021
-
[16]
Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics , 1715–1725
work page 2016
-
[17]
VanderWeele, T. J., & Ding, P . (2017). Sensitivity analysis in observational research: Introducing the E-value. Annals of Internal Medicine , 167(4), 268–274
work page 2017
-
[18]
Winner, L. (1980). Do artifacts have politics? Daedalus, 109(1), 121–136. 34 Supplementary Information The full Supplementary Information — including supplementary tables S1–S10 (DSRI ranking, Tier 1 measurements, byte-fallback classification, living-but-digitally-dead inventory, full question–answer ma- trix, mediation outputs, Unicode timeline, web corp...
work page 1980
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.