{"id":"b17daf17-d9d0-47e1-a040-86df91014a49","arxiv_id":"2606.28325","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Convergent errors across Claude, GPT-4o, Grok, and DeepSeek on 172 script features are consistent with imperial-era script inequalities persisting in LLMs via training corpora rather than model-specific design.","lead":"This paper builds a seven-axis Digital Script Representation Index and applies it to 300 writing systems, finding only 9.7% fully supported digitally and large tokenizer efficiency gaps. It reports that error patterns on script features converge across four different LLMs and interprets this as evidence that historical imperial inequalities persist through shared training data.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Cross-model error convergence may arise from shared LLM objectives or item design rather than imperial corpus effects","rationale":"The reader's weakest_assumption exactly identifies the load-bearing inference gap between observed convergence and the imperial-corpus causal story. No stronger internal inconsistency (e.g., in the mediation equations or DSRI construction) is visible from the supplied material; the primary risk remains alternative common causes for the LLM results. This matches the reader's UNVERDICTED assessment without requiring adjustment.","tokens_in":1946,"tokens_out":378,"duration_ms":21934,"concrete_test":"Recompute the 172-item error matrix and Spearman correlations after replacing one model family with an open model whose training corpus is documented to exclude or down-weight the same web sources (e.g., a model trained on curated non-web or post-2023 balanced script data); if mean rho remains >0.80, the imperial-corpus mechanism is not required to explain convergence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that Spearman rho = 0.85-0.98 convergence in base-rate-deviation errors (and 172 identically wrong items) across Claude/GPT-4o/Grok/DeepSeek is specifically caused by shared training corpus reflecting imperial history, not by common next-token objectives, overlapping web-scale data sources, tokenizer conventions, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment). The serial mediation (empire → population → web corpus → efficiency) is reported only as suggestive (beta = -0.22, p=0.39; CI grazes zero; n=45), with no ablation removing imperial-corpus subsets or comparison against models with divergent objectives. This leaves the historical attribution underdetermined by the reported evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the Digital Script Representation Index (DSRI), a seven-axis measure applied to 300 writing systems in the Global Script Database, reporting that only 9.7% are fully digitally supported and tokenizer efficiency varies by a factor of 31.7 across 45 scripts. It presents a serial mediation model (empire to population to web corpus to efficiency) as suggestive of full mediation (direct effect beta = -0.22, p = 0.39, CI grazing zero at n = 45) and documents high convergence in base-rate-deviation error patterns across four LLMs (Spearman rho = 0.85-0.98), with 172 identically wrong items, interpreting the results as evidence that historical imperial inequalities persist in LLMs via shared training corpora rather than model-specific design.","tokens_in":2119,"tokens_out":610,"duration_ms":28299,"significance":"If the attribution of cross-model convergence specifically to imperial-corpus effects holds after addressing alternative explanations, the result would be significant for documenting persistent structural biases in digital language technologies and for linking historical power relations to contemporary AI fairness issues.","major_comments":[{"comment":"Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described.","section":"Abstract"},{"comment":"Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined.","section":"Abstract"},{"comment":"Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: the bias-corrected bootstrap CI is stated to graze zero but the actual interval bounds are not reported, and no error bars or fit indices beyond 'indistinguishable from saturation' are provided for the mediation model.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The self-citation pattern to a future same-author paper for the core database raises concerns about data independence and whether the manuscript's scope is appropriate without prior independent publication of the Global Script Database."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments on the mediation analysis, convergence attribution, and data sourcing. We respond point-by-point below, noting where revisions will be made to address valid concerns while preserving the manuscript's qualified claims.","responses":[{"response":"The manuscript already qualifies the result as 'suggestive rather than confirmatory' due to the non-significant direct effect (beta = -0.22, p = 0.39) and CI grazing zero. No load-bearing causal claim is made; the text states only that the model is 'consistent with full mediation' via indirect paths and saturated fit at n = 45. We have added supplementary robustness checks (alternative SEM specifications and bootstrap variants) to the revision, though the small n precludes extensive ablations or non-imperial subset comparisons.","revision_made":"partial","referee_comment":"[Abstract] Abstract (mediation model): the serial mediation reports a direct-effect beta = -0.22 with p = 0.39 and CI grazing zero at n = 45, already labeled suggestive; this undercuts the load-bearing claim that the model supports causal attribution to imperial effects, as no ablations, robustness checks, or comparisons to non-imperial corpus subsets are described."},{"response":"We concur that shared objectives and data sources remain plausible alternatives and that the design does not isolate imperial-corpus effects. The reported convergence across four independent families is offered only as evidence against model-specific design. In revision we expand the limitations section to discuss these alternatives explicitly and report the religion-excluded sensitivity analysis (already in the text) as a formal check on item construction.","revision_made":"yes","referee_comment":"[Abstract] Abstract (convergence analysis): the Spearman rho = 0.85-0.98 convergence and 172 identically wrong items are attributed to shared imperial training corpus, yet the analysis provides no controls for shared next-token objectives, overlapping web-scale sources, or uniform construction of the 172 script-feature items (e.g., religion feature at 4.1x enrichment), leaving the specific causal attribution underdetermined."},{"response":"The Methods section already details the 300 scripts, seven DSRI axes, and 45-script tokenizer subsample. We have revised all citations to reference the present manuscript for the database and deposited the full dataset as supplementary material, removing reliance on the forthcoming companion paper.","revision_made":"yes","referee_comment":"[Abstract] Abstract (data source): the Global Script Database is cited as Fukui (2026), a future publication by the same author; the mediation analysis at n = 45 relies on this database, creating circularity that weakens the grounding of the imperial-intervention path."}],"tokens_in":1667,"tokens_out":590,"duration_ms":25490,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's clearest contribution is documenting that four unrelated LLM families produce nearly identical error patterns on script questions, with Spearman correlations of 0.85-0.98 and 172 items missed by all of them. The over-attribution asymmetry and the religion-feature enrichment are also reported plainly. Those numbers are worth checking against the actual items.\n\nWhat is new is the DSRI seven-axis index applied to 300 scripts and the tokenizer-efficiency measurements across 45 parallel texts. The descriptive counts on digital support (only 29 fully supported, 60 of 158 living scripts incomplete) give a concrete baseline that earlier bias papers have not assembled at this scale.\n\nThe work is weaker on the causal step. The serial mediation uses n=45, shows a direct-effect beta of -0.22 with p=0.39 and a CI that grazes zero, and is explicitly called suggestive. No ablation removes imperial-corpus subsets or compares against models trained on deliberately different data mixtures. The high convergence could equally come from shared next-token objectives, overlapping web crawls, or how the 172 test items were constructed. The Global Script Database is cited to a 2026 self-reference, which blocks immediate verification.\n\nThis is for people who track tokenizer fairness and low-resource script access. A reader focused on measurement will get usable numbers; anyone needing a confirmed historical mechanism will not.\n\nI would send it for peer review. The convergence data and DSRI construction are concrete enough to justify referee time, provided the authors add controls for non-historical explanations and release the error-item list and database for inspection.","headline":"The cross-model convergence on 172 script errors is a solid observation, but the empire-to-corpus causal story rests on suggestive mediation without ruling out simpler alternatives.","tokens_in":2672,"tokens_out":405,"would_cite":false,"duration_ms":22108,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Structural inequalities from historical empires persist in language models through their shared training corpora rather than model-specific designs.","keywords":["large language models","writing systems","digital inequality","historical empires","tokenizer efficiency","mediation analysis","script representation","bias convergence"],"falsifier":"Train a new language model on a web corpus deliberately equalized for script representation independent of historical empire sizes and test whether its base-rate-deviation error patterns on the same 172 items diverge from the four studied models.","tokens_in":2790,"feed_emoji":"🗺️","tokens_out":670,"duration_ms":25453,"temperature":0.7,"pith_summary":"The paper establishes that four unrelated large language models display nearly identical patterns of error when representing the world's writing systems. These patterns align with historical imperial influences on which scripts achieved large speaker populations and digital presence. A mediation analysis finds that imperial history affects current tokenizer performance only indirectly through population and web corpus size. High correlations in errors across the models support the view that the shared training data transmits the inequalities.","feed_headline":"Four LLMs converge on identical imperial script biases","feed_subtitle":"Error patterns across models trace to shared training data shaped by historical empires, not to differences in model architecture.","key_machinery":"serial mediation model from imperial intervention through speaker population and web corpus to tokenizer efficiency, evidenced by convergent error patterns across independent models","core_discovery":"Large language models process the world's writing systems with radical inequality. Only 29 of 300 scripts are fully supported, and tokenizer efficiency varies by a factor of 31.7. A serial mediation model shows full mediation from imperial intervention through speaker population and web corpus to tokenizer efficiency, with the direct effect of empire indistinguishable from zero. Across four LLM families, base-rate-deviation error patterns converge at Spearman rho 0.85-0.98, with 172 items answered identically wrong by all models and over-attribution outnumbering under-recognition, consistent with the inequalities persisting through the shared training corpus.","pith_inferences":["Curating future training corpora to increase representation of historically marginalized scripts could reduce these convergent biases across models.","The same shared-corpus mechanism may produce parallel patterns of bias in other AI systems that draw from the same web data sources.","Long-term digital consequences of historical data collection practices extend beyond language models to any system reliant on large-scale internet text."],"forward_implications":["The direct effect of empire on tokenizer efficiency becomes indistinguishable from zero once population and web corpus are accounted for.","Over-attribution of scripts to religious use concentrates 43.6 percent of the convergent errors.","The convergence across models remains after excluding the religion feature, indicating multi-channeled rather than single-channeled bias.","Only 9.7 percent of scripts achieve full digital support and 38 percent of living scripts lack complete support."],"fun_headline_variants":["Empire biases persist in LLM script processing","Four LLMs share script error patterns from empires","Historical empires shape LLM tokenizer inequalities","Script biases converge across four language models"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The high correlations in error patterns across models are caused by shared training-corpus effects rooted in imperial history rather than by similar pre-training objectives, tokenizer conventions, or evaluation methods.","fun_headline_variants_meta":{"raw":{"variants":["Empire biases persist in LLM script processing","Four LLMs share script error patterns from empires","Historical empires shape LLM tokenizer inequalities","Script biases converge across four language models"]},"model":"grok-4.3","cost_usd":0.007825,"raw_usage":{"total_tokens":3676,"prompt_tokens":877,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":78249500,"prompt_tokens_details":{"text_tokens":877,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2749,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":877,"tokens_out":50,"duration_ms":28954,"temperature":1.0,"reasoning_tokens":2749,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T08:12:24.352731+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Train a new language model on a web corpus deliberately equalized for script representation independent of historical empire sizes and test whether its base-rate-deviation error patterns on the same 172 items diverge from the four studied models.","supporting_citations":[],"review_version":1}