REVIEW 2 major objections 2 minor
Complexity counts: global and local perspectives on Indo-Aryan numeral systems
T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3
Pith's one-line read Indo-Aryan languages have decisively more complex numeral systems than the world's languages as a whole, yet still follow universal efficiency pressures.
desk verdict Indo-Aryan languages show higher numeral complexity by the new metrics, with some regional links, but the metrics need checks for coding and sampling bias. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cross-linguistically applicable metrics for quantifying numeral system complexity, applied to data from multiple databases to compare Indo-Aryan languages against the world.
What would settle it
A finding that Indo-Aryan languages score no higher than the global average on the same metrics when applied to a larger, independently compiled sample of numeral systems.
Extended reading notes
Core claim
Indo-Aryan languages possess numeral systems that are markedly more complex than the global average, as quantified by newly developed cross-linguistic metrics, with individual variation among them; these systems persist due to local factors such as religion and isolation in South Asia but conform to universal pressures for communicative efficiency.
Load-bearing premise
The newly developed cross-linguistically applicable metrics accurately quantify numeral complexity without introducing bias from language-specific features or data selection choices.
Editorial extensions
If this is right
- Complex numeral systems can persist in certain linguistic communities despite general trends toward simplicity.
- Local socio-cultural factors like geographic isolation influence the maintenance of complex numeral forms.
- Numeral systems in Indo-Aryan languages balance high complexity with adherence to efficiency principles found elsewhere.
- The dimension of complexity in numeral systems deserves more attention in typological studies of language variation.
Reading between the lines
- The same metrics could be applied to other linguistic domains, such as kinship terms or color naming, to test whether complexity patterns recur.
- Persistence of complex systems may indicate that efficiency pressures interact with cultural transmission rather than acting as strict constraints.
- Future work could examine whether similar non-transparent forms appear in number-related expressions beyond basic counting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that Indo-Aryan numeral systems (e.g., Hindi ikyānve for 91) are highly non-transparent and, using newly developed cross-linguistically applicable metrics on data from multiple databases, demonstrates that Indo-Aryan languages have decisively more complex numeral systems than the world's languages overall. Individual Indo-Aryan languages vary in complexity; the paper explores factors such as religion and geographic isolation that may explain persistence of complexity in South Asia, while showing that these systems still conform to general cross-linguistic pressures toward efficient communication.
Significance. If the metrics prove robust and unbiased, the work would usefully expand numeral-system typology by documenting an under-studied locus of complexity and by linking local South-Asian patterns to broader communicative-efficiency constraints. The multi-database approach and explicit call to treat complexity as a serious dimension of variation are positive contributions.
major comments (2)
- [§3] §3 (Development of complexity metrics): The central claim that Indo-Aryan systems are 'decisively more complex' rests entirely on the new metrics. The manuscript provides no validation against independent codings, no sensitivity tests to database inclusion/exclusion rules, and no explicit inter-coder reliability or error-handling protocol. Without these, it is impossible to rule out systematic bias favoring well-documented or morphologically rich languages.
- [§4.2] §4.2 (Coding of opacity/non-transparency): The treatment of forms such as ikyānve requires a fully explicit, replicable decision tree for deciding when a numeral is 'non-decomposable.' The current description leaves open whether the same rules, applied to a balanced global sample, would still place Indo-Aryan languages as clear outliers.
minor comments (2)
- [Table 1] Table 1: column headers for the global baseline sample should explicitly state the total number of languages and the source databases used for each row.
- [Figure 3] Figure 3: the y-axis label 'Complexity Score' should be accompanied by a brief parenthetical reminding readers which metric(s) are aggregated.
Simulated Author's Rebuttal
Thank you for your constructive and detailed review. We address each major comment below and indicate revisions made to enhance methodological transparency and robustness while preserving the core findings on Indo-Aryan numeral complexity.
read point-by-point responses
-
Referee: §3 (Development of complexity metrics): The central claim that Indo-Aryan systems are 'decisively more complex' rests entirely on the new metrics. The manuscript provides no validation against independent codings, no sensitivity tests to database inclusion/exclusion rules, and no explicit inter-coder reliability or error-handling protocol. Without these, it is impossible to rule out systematic bias favoring well-documented or morphologically rich languages.
Authors: We agree that explicit validation strengthens the claims. In the revised manuscript we add sensitivity tests demonstrating that Indo-Aryan languages remain the highest-complexity group when any single database is excluded. We also document our error-handling protocol, which cross-references primary sources for ambiguous forms and flags uncertain cases. Inter-coder reliability is now reported via independent recoding of a 50-language subsample with substantial agreement. These additions address potential bias without altering the quantitative results. Full external validation by independent teams lies beyond the present scope but is noted as desirable future work. revision: yes
-
Referee: §4.2 (Coding of opacity/non-transparency): The treatment of forms such as ikyānve requires a fully explicit, replicable decision tree for deciding when a numeral is 'non-decomposable.' The current description leaves open whether the same rules, applied to a balanced global sample, would still place Indo-Aryan languages as clear outliers.
Authors: We accept that greater explicitness is required. The revised version includes a step-by-step decision tree for decomposability, specifying morphological segmentation criteria, semantic transparency thresholds, and handling of irregular forms, illustrated with ikyānve and parallel examples from non-Indo-Aryan languages. Reapplication of these rules to the full multi-database global sample confirms that Indo-Aryan languages remain clear outliers on all complexity metrics, consistent with the original figures and tables. revision: yes
Circularity Check
No significant circularity: empirical metrics derived from external databases support comparative claim without definitional reduction.
full rationale
The paper develops cross-linguistically applicable metrics from multiple external databases to quantify numeral complexity and compares Indo-Aryan languages against a global baseline. The central demonstration that Indo-Aryan systems are decisively more complex rests on this empirical application rather than any self-definitional loop, fitted parameter renamed as prediction, or load-bearing self-citation chain. No equations or derivations reduce the output to the input by construction; the approach is self-contained against external data sources and benchmarks.
Assumptions & free parameters
assumptions (2)
- domain assumption Numeral system complexity can be quantified by cross-linguistically applicable metrics that capture transparency and decomposability.
- domain assumption Database data on numeral systems is sufficiently complete and unbiased for global comparison.
Cite this review
Pith. "Pith review of Complexity counts: global and local perspectives on Indo-Aryan numeral systems." pith.science (2026). https://pith.science/paper/2505.21510
@misc{pith2026250521510,
author = {Pith},
title = {Pith review of: Complexity counts: global and local perspectives on Indo-Aryan numeral systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/2505.21510}},
note = {Machine review of arXiv:2505.21510}
}
read the original abstract
The numeral systems of Indo-Aryan languages such as Hindi, Gujarati, and Bengali are highly unusual in that unlike most numeral systems (e.g., those of English, Chinese, etc.), forms referring to 1--99 are highly non-transparent and cannot be constructed using straightforward rules for forming combinations of tens and digits. As an example, Hindi/Urdu {\it iky\=anve} `91' is not decomposable into the composite elements {\it ek} `one' and {\it nave} `ninety' in the way that its English counterpart is. This paper further clarifies the position of Indo-Aryan languages within the typology of numeral systems, and explores the linguistic and non-linguistic factors that may be responsible for the persistence of complex systems in these languages. Using data from multiple databases, we develop and employ a number of cross-linguistically applicable metrics to quantify the complexity of languages' numeral systems, and demonstrate that Indo-Aryan languages have decisively more complex numeral systems than the world's languages as a whole, though individual Indo-Aryan languages differ from each other in terms of the complexity of the patterns they display. We investigate the factors (e.g., religion, geographic isolation, etc.) that underlie complexity in numeral systems, with a focus on South Asia, in an attempt to develop an account of why complex numeral systems developed and persisted in certain Indo-Aryan languages but not elsewhere. Finally, we demonstrate that Indo-Aryan numeral systems adhere to certain general pressures toward efficient communication found cross-linguistically, despite their high complexity. We call for this somewhat overlooked dimension of complexity to be taken seriously when discussing general variation in numeral systems.
Reviewed May 22, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.