{"id":"98da89c3-d27c-4379-ba38-69ac66fcd0be","arxiv_id":"2603.28213","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM tokenizers, training data and benchmarks reproduce standard-language hierarchies, leaving South Tyrolean and most Kurdish varieties marginalized.","lead":"This paper argues that large language models keep marginalizing people who speak non-standard dialects, using South Tyrolean German and Kurdish as case studies. It connects tokenization costs, missing benchmarks and standard-language bias, and says fixing this requires policy choices as much as technical ones.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Empirical premise of LLM marginalization of South Tyrolean and Kurdish varieties is anecdotal, not measured","rationale":"The strongest claim is an explanatory one about how the digital language divide is maintained. For that explanation to be convincing, the empirical fact of LLM deficiency on the selected varieties must be established. The paper provides only anecdotal evidence ('V ersteasch du mi?'), generic tokenization examples from other languages, and unsourced resource tables. The §6.2 survey is not even attributable to a named author. This matches the reader's weakest_assumption exactly. However, the work is explicitly a position/review essay, not a controlled experimental study. Its value is in synthesizing existing literature (Petrov et al., Ahia et al., Faisal et al.) and proposing policy directions. The absence of a controlled evaluation does not make the essay internally inconsistent; it leaves the central claim as an untested hypothesis. The reader's UNVERDICTED verdict is therefore appropriate, and my concern does not move it to accept or reject. A concrete evaluation would be the natural next step and would turn the paper's empirical premise into a testable claim.","tokens_in":17630,"tokens_out":4458,"duration_ms":47519,"concrete_test":"Build a small but representative benchmark: 100 natural sentences each in South Tyrolean dialect, Central Kurdish (Sorani), and Northern Kurdish (Kurmanji), matched to standard German/English equivalents. For several widely used LLMs (e.g., GPT-4o, Llama-3.1-70B), measure (1) token-to-word ratio and per-token cost, (2) comprehension via a multiple-choice paraphrase or follow-up instruction task, and (3) human-rated response fluency. Compare all metrics against standard-language baselines. If the tokenization difference is ≤1.2× and comprehension accuracy is statistically indistinguishable from standard baselines, the paper's marginalization premise is contradicted; if a clear gap emerges, the premise is empirically supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that the digital language divide is sustained by market forces, state policies, and data-pipeline biases (§7)—depends on the premise that LLMs actually fail on South Tyrolean and Kurdish varieties. That premise is asserted, not demonstrated. The only direct evidence is the informal comprehension check in §5: several models answered 'yes' to 'V ersteasch du mi?', which if anything suggests some working knowledge of the dialect. The tokenization illustrations (e.g., Irish 'ionchomharthú') are not for the two case-study varieties, and the resource inventories (Tables 1, 3, 4) lack per-cell sources and rely on vague figures (e.g., 'speaker counts based on the best information in Wikipedia'). The §6.2 literature survey is attributed to 'AUTHOR' and is unverifiable as presented. If a controlled evaluation showed that LLMs comprehend South Tyrolean and Kurdish reasonably well, and that tokenization costs are only marginally worse than for standard languages, the paper's empirical foundation would collapse—even though its sociopolitical argument might still stand. Thus the key assumption of actual failure is load-bearing and unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the digital language divide is sustained by market forces, historical state policies, and biases in standardized data pipelines, and that non-standard varieties such as South Tyrolean German and Kurdish are therefore marginalized in LLMs and GenAI. It combines a sociolinguistic review of standardization and language policy with a computational-linguistics discussion of tokenization, benchmark design, and resource scarcity, then applies this framework to two case studies. For South Tyrolean it notes the lack of an ISO code, sparse resources, and an anecdotal comprehension check; for Kurdish it describes the unequal institutional support and computational resources across varieties. The conclusion proposes policy measures (dialect gap reporting, CSR credits, data sovereignty, and interdisciplinary synthesis) as necessary complements to technical fixes.","tokens_in":17914,"tokens_out":4291,"duration_ms":48483,"significance":"If its empirical premises hold, the paper provides a useful interdisciplinary bridge between critical sociolinguistics and NLP, and its policy proposals (e.g., dialect gap reporting, community data sovereignty) are concrete and actionable. It draws on a broad and appropriate literature, clearly structures the technical and sociopolitical dimensions, and identifies gaps that are real in the research landscape. Its main weaknesses are empirical: the load-bearing claim that LLMs actually fail on the two case-study varieties rests on anecdote and unverified inventories rather than measurement, and several resource tables lack sources or methodology. There are no original computational evaluations or machine-checked proofs; the contribution is conceptual and policy-oriented. The paper is therefore better framed as a position paper than an empirical demonstration, and the empirical claims should be either substantiated or explicitly scoped down.","major_comments":[{"comment":"The only direct behavioral evidence about South Tyrolean is the statement that 'several Generative AI models responded affirmative when we asked \"V ersteasch du mi?\"'. This is anecdotal: no model names, versions, dates, prompt variants, number of trials, or evaluation criteria are given. More importantly, an affirmative answer to 'Do you understand me?' indicates at least some basic comprehension, which sits uneasily with the paper's premise that LLMs marginalize the dialect. The conclusion in §7 depends on that premise. I recommend a small controlled probe (multiple dialect sentences, several models, exact model/version metadata) or an explicit rephrasing of the paper's claim from 'LLMs fail on these varieties' to 'LLMs are not evaluated or optimized for these varieties.'","section":"§5, 'Versteasch du mi?' check"},{"comment":"The tokenization-tax argument is central to the charge of algorithmic discrimination, but the examples are not reproducible and are not from the case studies. The Irish word 'ionchomharthú' is attributed to 'bert-based-uncased' without a model version, input normalization, or access date, and no token counts are reported for South Tyrolean or Kurdish. The claim that 'a speaker of Kurdish or South Tyrolean literally pays more' is supported by citations to Petrov et al. (2023) and Ahia et al. (2023) for the general phenomenon, but no case-specific measurement is shown. Please provide reproducible tokenizer comparisons (model, version, date, token sequences, and token counts for representative South Tyrolean and Kurdish phrases) or soften the cost claim accordingly.","section":"§3, tokenization examples"},{"comment":"Tables 3 and 4 are empirical pillars of the Kurdish case. Table 3 reports token counts, audio hours, and model/MT support with blank cells and no per-cell sources, definitions, or access dates; Table 4 does not visibly show inclusion/exclusion marks. The 'survey of the NLP literature carried out by AUTHOR' over 'over 100 papers' is unverifiable as presented: there is no search protocol, inclusion/exclusion criteria, or inter-coder reliability. The strong statement that Southern Kurdish, Laki, Zazaki, and Hawrami are 'entirely unsupported' (Table 3) should be backed by a documented, repeatable search procedure or hedged to 'not covered by the resources we surveyed.'","section":"§6.2 and Tables 3–4"},{"comment":"The speaker counts are based on 'the best information in Wikipedia,' which is not adequate provenance for demographic claims, and the OPUS/VLO columns lack definitions and access dates. Because Table 1 is used to argue that South Tyrolean's lack of an ISO code hampers its visibility in NLP resources, the absence of a systematic check for South Tyrolean-specific entries in OPUS and VLO weakens the point. Please state how the numbers were collected and whether any South Tyrolean-specific data are found under the Bavarian code.","section":"Table 1"},{"comment":"The conclusion asserts that the digital language divide 'is maintained by a complex interplay of market forces, historical state policies, and the inherent biases of standardized data pipelines.' The historical and political parts are well supported by the sociolinguistic literature, but the causal force of 'is maintained by' is stronger than the evidence presented, which does not adjudicate among alternative explanations (e.g., pure economic incentives or technical path dependence). I suggest framing this as a hypothesis or analytical framework rather than an empirically established statement, unless the authors add evidence that directly tests the relative contributions of these factors.","section":"§7, causal conclusion"}],"minor_comments":[{"comment":"The author name is typeset as 'V erena Platzgummer' with an unwanted space; the title's dialect phrase is also inconsistently formatted across the text.","section":"Abstract and author list"},{"comment":"'does not bare any resemblance' should be 'bear any resemblance.' Also, the abbreviation 'URLs' for Under-Resourced Languages may be confusing to readers who identify URL as Uniform Resource Locator; consider spelling out the term at first use.","section":"§3"},{"comment":"The caption contains a typo: 'Virtual Language Obvservatory' should be 'Virtual Language Observatory.' Please also state the access date for OPUS and VLO data.","section":"Table 1 caption"},{"comment":"The benchmark name is written as 'V oxLect' in the text and 'Voxlect' elsewhere; please standardize the spelling (e.g., 'VoxLect').","section":"§3, VoxLect"},{"comment":"The model 'KuBERT' appears in Table 3 but is not cited or described in the text. Also, the table legend should explain whether blank cells mean 'no data' or 'no support.'","section":"§6.2, Table 3"},{"comment":"Reference entries with '1 others' (e.g., Glaznieks et al. 2018) need full author lists or at least standard 'et al.' formatting.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is best understood as a position/conceptual paper, but its title and conclusion make empirical claims about LLM failure on South Tyrolean and Kurdish varieties that the evidence does not support. The 'Versteasch du mi?' anecdote, in particular, could be read as counterevidence to the paper's own framing. I would encourage the authors to either add a modest controlled evaluation (even a small probe study) or explicitly reposition the contribution as one about non-evaluation and non-resourcing rather than demonstrated LLM failure. The placeholder 'AUTHOR' in §6.2 should be resolved before review, as it prevents verification of a stated survey result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a position/review essay, not an empirical paper. The core claims about tokenization unfairness, benchmark absence, and standardization bias come from the literature the paper cites, and the new evidence is a few handpicked examples plus two resource-inventory tables. That does not mean the argument is wrong; it is a coherent and mostly fair synthesis. But the manuscript should not be read as a demonstration that LLMs actually fail on South Tyrolean or Kurdish varieties.\n\nWhat is genuinely useful: bringing South Tyrolean and Kurdish into one frame works well. The missing ISO-639 code point for South Tyrolean, and its absorption into 'Bavarian' in NLP resources, is concrete and underappreciated. The Kurdish resource tables give a quick sense of which varieties are supported and which are invisible. The survey of recent dialect benchmarks (DialectBench, Global MMLU, IRLBench, VoxLex) is up to date and largely accurate as far as I can tell. For a reader new to this area, this is a good map.\n\nThe soft spots are real and load-bearing. The 'V ersteasch du mi?' check in §5 is anecdotal; several models answering yes is weak support and could even cut the other way. The tokenizer examples use Irish, not the two case varieties, so they do not directly demonstrate the claimed tax for South Tyrolean or Kurdish. Tables 1, 3, and 4 lack per-cell sources, collection dates, and methodology; 'speaker counts based on the best information in Wikipedia' is not an adequate basis. The §6.2 literature survey attributed to 'AUTHOR' is a placeholder and unverifiable as presented. These are not cosmetic issues.\n\nThe central argument still holds up as a synthesis. The structural claim — that market forces, historical state policies, and data-pipeline biases combine to maintain the digital language divide — is consistent with the cited literature. But because the paper's own evidence is thin, the conclusion should be framed as an agenda or hypothesis, not as an established finding. If controlled evaluations later show strong dialect comprehension despite tokenization costs, the marginalization claim would need serious qualification.\n\nWho gets value: people wanting a compact introduction to the dialect-gap debate, and policy readers looking for a map of actors and levers. It deserves a serious referee only if the venue accepts position papers and the authors are asked to fix the placeholder, source the tables, and either add minimal reproducible evaluations or explicitly narrow the empirical claims. As original empirical research, it is not ready.","headline":"A readable, useful synthesis on dialect/LLM inequity, but the empirical premise is asserted rather than demonstrated — treat it as a position paper, not as a result.","tokens_in":18343,"tokens_out":4686,"would_cite":false,"duration_ms":51199,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs systematically tax and exclude non-standard dialects, and closing the gap requires policy change, not just better models.","keywords":["tokenization tax","non-standard language","South Tyrolean dialect","Kurdish varieties","linguistic standardization","digital language divide","LLM fairness","language policy"],"falsifier":"A controlled evaluation on South Tyrolean and Southern Kurdish, comparing task accuracy and per-token cost against Standard German across several models; if dialect performance matched the standard when token counts and cost are equalized, the claim that pipelines impose a systematic penalty would weaken.","tokens_in":17548,"feed_emoji":"💬","tokens_out":3639,"duration_ms":38300,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models do not merely under-serve non-standard language varieties—they actively penalize them. Through the cases of South Tyrolean dialects and Kurdish varieties, it argues that tokenizer fragmentation, missing ISO codes, and benchmark gaps combine with market forces and historical language policy to push dialect speakers into paying more and getting worse results. The authors contend that closing the 'dialect gap' therefore requires policy change, community data sovereignty, and new evaluation frameworks, not just technical fixes. A sympathetic reader would care because it reframes an apparently technical problem as a question of linguistic justice and digital citizenship.","feed_headline":"Dialects pay a tokenization tax in LLMs, and the fix is political","feed_subtitle":"South Tyrolean and Kurdish varieties get worse results at higher cost; the paper says policy, not just tech, must close the gap.","key_machinery":"The load-bearing mechanism is the tokenization tax: because BPE tokenizers split text based on frequency in a training corpus, dialect words often break into five or six low-information byte fragments, tripling or quintupling per-token cost, shrinking the usable context window, and degrading performance. Alongside it, the ISO-639 code acts as a gatekeeper that determines whether a variety is even catalogued in NLP resources, and translation-pivot benchmarks define what counts as competence in a language.","core_discovery":"The paper's central claim is that the digital language divide is maintained by a complex interplay of market forces, historical state policies, and the inherent biases of standardized data pipelines. For South Tyrolean, the lack of an ISO-639 code makes the variety invisible to NLP resource catalogs and benchmarks, while its non-standard orthography and the pull of Standard German leave it under-sampled and fragmented by BPE tokenizers. For Kurdish, varieties beyond Central and Northern Kurdish—Southern Kurdish, Laki, Zazaki, Hawrami—are almost entirely absent from corpora, models, benchmarks, and translation services, a marginalization that mirrors and reinforces political hierarchies. The","pith_inferences":["A testable extension would be to measure dialect competence while controlling for tokenization—for instance, comparing task success against the number of tokens consumed—since the paper predicts that even equal-accuracy cases still suffer economically under per-token pricing.","The same pipeline biases likely generalize to other dialect continua beyond South Tyrolean and Kurdish, such as Arabic, Chinese, or African varieties, because the ISO-code and benchmark arguments are largely portable.","The tokenization tax could be quantified directly: tokenizing a fixed sentence in several varieties and comparing token counts and per-token prices across vendors would turn the paper's anecdotal evidence into a metric.","If the paper's policy recommendations were adopted, an observable consequence would be a shift from web-scraped, proprietary data to consent-based, open corpora with community audit—changing who controls language data."],"forward_implications":["If the analysis is right, dialect speakers pay more per prompt and get shorter conversations and poorer reasoning—a direct economic and quality penalty.","Adding dialect data or fine-tuning alone will not close the gap while tokenizers remain frequency-biased and per-token pricing persists.","Requiring 'dialect gap' reporting, as the paper urges, could turn the currently invisible performance cliff into a measurable, auditable inequality.","Benchmarks built by translating English questions actively mislead: a model can score well while failing culturally grounded tasks, so local, non-translated benchmarks are needed.","The EU AI Act's non-discrimination requirement could become a legal hook: an AI that fails a South Tyrolean citizen might count as discrimination."],"fun_headline_variants":["LLMs levy a hidden tax on dialects—policy must intervene","Tokenizers mangle non-standard speech, and tech alone won't fix it","Why LLMs shortchange South Tyrolean and Kurdish dialects","The digital language divide: it's not just algorithms, it's politics","Missing ISO codes and tokenizer bias: the real cost of non-standard dialects"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes, without a controlled measurement, that LLMs actually fail on South Tyrolean and Kurdish varieties; its evidence is an anecdotal 'yes' from a few models, a few tokenizer examples, and resource tables that cite no per-cell sources.","fun_headline_variants_meta":{"raw":{"variants":["LLMs levy a hidden tax on dialects—policy must intervene","Tokenizers mangle non-standard speech, and tech alone won't fix it","Why LLMs shortchange South Tyrolean and Kurdish dialects","The digital language divide: it's not just algorithms, it's politics","Missing ISO codes and tokenizer bias: the real cost of non-standard dialects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000876,"raw_usage":{"total_tokens":3676,"prompt_tokens":846,"completion_tokens":2830,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":2736}},"tokens_in":590,"tokens_out":2830,"duration_ms":19877,"temperature":1.0,"reasoning_tokens":2736,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:06:33.873185+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled evaluation on South Tyrolean and Southern Kurdish, comparing task accuracy and per-token cost against Standard German across several models; if dialect performance matched the standard when token counts and cost are equalized, the claim that pipelines impose a systematic penalty would weaken.","supporting_citations":[],"review_version":1}