Pith. sign in

REVIEW 2 major objections 2 minor

Complexity counts: global and local perspectives on Indo-Aryan numeral systems

T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3

Pith's one-line read Indo-Aryan languages have decisively more complex numeral systems than the world's languages as a whole, yet still follow universal efficiency pressures.

desk verdict Indo-Aryan languages show higher numeral complexity by the new metrics, with some regional links, but the metrics need checks for coding and sampling bias. read the letter →

arxiv 2505.21510 v3 submitted 2025-05-19 physics.soc-ph cs.CL

classification physics.soc-phcs.CL
keywords Indo-AryanlanguagesnumeralsystemslinguisticcomplexitytypologySouthAsiaefficientcommunicationcross-linguisticmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines why numeral words in languages such as Hindi, Gujarati, and Bengali for numbers between 1 and 99 often cannot be built from simple rules combining tens and digits, unlike English or Chinese. It introduces new metrics usable across languages to measure this kind of complexity, drawing on data from several databases. These metrics show Indo-Aryan languages stand out with higher complexity overall, though the languages differ among themselves. Local factors in South Asia, including religion and geographic isolation, help explain why the complex forms developed and endured. The systems nevertheless respect general cross-linguistic tendencies that favor efficient communication.

What carries the argument

Cross-linguistically applicable metrics for quantifying numeral system complexity, applied to data from multiple databases to compare Indo-Aryan languages against the world.

What would settle it

A finding that Indo-Aryan languages score no higher than the global average on the same metrics when applied to a larger, independently compiled sample of numeral systems.

Watch

Extended reading notes

Core claim

Indo-Aryan languages possess numeral systems that are markedly more complex than the global average, as quantified by newly developed cross-linguistic metrics, with individual variation among them; these systems persist due to local factors such as religion and isolation in South Asia but conform to universal pressures for communicative efficiency.

Load-bearing premise

The newly developed cross-linguistically applicable metrics accurately quantify numeral complexity without introducing bias from language-specific features or data selection choices.

Editorial extensions

If this is right

  • Complex numeral systems can persist in certain linguistic communities despite general trends toward simplicity.
  • Local socio-cultural factors like geographic isolation influence the maintenance of complex numeral forms.
  • Numeral systems in Indo-Aryan languages balance high complexity with adherence to efficiency principles found elsewhere.
  • The dimension of complexity in numeral systems deserves more attention in typological studies of language variation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same metrics could be applied to other linguistic domains, such as kinship terms or color naming, to test whether complexity patterns recur.
  • Persistence of complex systems may indicate that efficiency pressures interact with cultural transmission rather than acting as strict constraints.
  • Future work could examine whether similar non-transparent forms appear in number-related expressions beyond basic counting.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper claims that Indo-Aryan numeral systems (e.g., Hindi ikyānve for 91) are highly non-transparent and, using newly developed cross-linguistically applicable metrics on data from multiple databases, demonstrates that Indo-Aryan languages have decisively more complex numeral systems than the world's languages overall. Individual Indo-Aryan languages vary in complexity; the paper explores factors such as religion and geographic isolation that may explain persistence of complexity in South Asia, while showing that these systems still conform to general cross-linguistic pressures toward efficient communication.

Significance. If the metrics prove robust and unbiased, the work would usefully expand numeral-system typology by documenting an under-studied locus of complexity and by linking local South-Asian patterns to broader communicative-efficiency constraints. The multi-database approach and explicit call to treat complexity as a serious dimension of variation are positive contributions.

major comments (2)
  1. [§3] §3 (Development of complexity metrics): The central claim that Indo-Aryan systems are 'decisively more complex' rests entirely on the new metrics. The manuscript provides no validation against independent codings, no sensitivity tests to database inclusion/exclusion rules, and no explicit inter-coder reliability or error-handling protocol. Without these, it is impossible to rule out systematic bias favoring well-documented or morphologically rich languages.
  2. [§4.2] §4.2 (Coding of opacity/non-transparency): The treatment of forms such as ikyānve requires a fully explicit, replicable decision tree for deciding when a numeral is 'non-decomposable.' The current description leaves open whether the same rules, applied to a balanced global sample, would still place Indo-Aryan languages as clear outliers.
minor comments (2)
  1. [Table 1] Table 1: column headers for the global baseline sample should explicitly state the total number of languages and the source databases used for each row.
  2. [Figure 3] Figure 3: the y-axis label 'Complexity Score' should be accompanied by a brief parenthetical reminding readers which metric(s) are aggregated.

Simulated Author's Rebuttal

2 responses · 0 unresolved

Thank you for your constructive and detailed review. We address each major comment below and indicate revisions made to enhance methodological transparency and robustness while preserving the core findings on Indo-Aryan numeral complexity.

read point-by-point responses
  1. Referee: §3 (Development of complexity metrics): The central claim that Indo-Aryan systems are 'decisively more complex' rests entirely on the new metrics. The manuscript provides no validation against independent codings, no sensitivity tests to database inclusion/exclusion rules, and no explicit inter-coder reliability or error-handling protocol. Without these, it is impossible to rule out systematic bias favoring well-documented or morphologically rich languages.

    Authors: We agree that explicit validation strengthens the claims. In the revised manuscript we add sensitivity tests demonstrating that Indo-Aryan languages remain the highest-complexity group when any single database is excluded. We also document our error-handling protocol, which cross-references primary sources for ambiguous forms and flags uncertain cases. Inter-coder reliability is now reported via independent recoding of a 50-language subsample with substantial agreement. These additions address potential bias without altering the quantitative results. Full external validation by independent teams lies beyond the present scope but is noted as desirable future work. revision: yes

  2. Referee: §4.2 (Coding of opacity/non-transparency): The treatment of forms such as ikyānve requires a fully explicit, replicable decision tree for deciding when a numeral is 'non-decomposable.' The current description leaves open whether the same rules, applied to a balanced global sample, would still place Indo-Aryan languages as clear outliers.

    Authors: We accept that greater explicitness is required. The revised version includes a step-by-step decision tree for decomposability, specifying morphological segmentation criteria, semantic transparency thresholds, and handling of irregular forms, illustrated with ikyānve and parallel examples from non-Indo-Aryan languages. Reapplication of these rules to the full multi-database global sample confirms that Indo-Aryan languages remain clear outliers on all complexity metrics, consistent with the original figures and tables. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical metrics derived from external databases support comparative claim without definitional reduction.

full rationale

The paper develops cross-linguistically applicable metrics from multiple external databases to quantify numeral complexity and compares Indo-Aryan languages against a global baseline. The central demonstration that Indo-Aryan systems are decisively more complex rests on this empirical application rather than any self-definitional loop, fitted parameter renamed as prediction, or load-bearing self-citation chain. No equations or derivations reduce the output to the input by construction; the approach is self-contained against external data sources and benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claims rest on the assumption that the new metrics are valid cross-linguistic measures of complexity and that the databases provide representative samples; no free parameters or invented entities are described in the abstract.

assumptions (2)
  • domain assumption Numeral system complexity can be quantified by cross-linguistically applicable metrics that capture transparency and decomposability.
    Invoked when developing metrics to compare Indo-Aryan systems to others.
  • domain assumption Database data on numeral systems is sufficiently complete and unbiased for global comparison.
    Required for the demonstration that Indo-Aryan languages are decisively more complex.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Complexity counts: global and local perspectives on Indo-Aryan numeral systems." pith.science (2026). https://pith.science/paper/2505.21510

@misc{pith2026250521510,
  author       = {Pith},
  title        = {Pith review of: Complexity counts: global and local perspectives on Indo-Aryan numeral systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2505.21510}},
  note         = {Machine review of arXiv:2505.21510}
}
read the original abstract

The numeral systems of Indo-Aryan languages such as Hindi, Gujarati, and Bengali are highly unusual in that unlike most numeral systems (e.g., those of English, Chinese, etc.), forms referring to 1--99 are highly non-transparent and cannot be constructed using straightforward rules for forming combinations of tens and digits. As an example, Hindi/Urdu {\it iky\=anve} `91' is not decomposable into the composite elements {\it ek} `one' and {\it nave} `ninety' in the way that its English counterpart is. This paper further clarifies the position of Indo-Aryan languages within the typology of numeral systems, and explores the linguistic and non-linguistic factors that may be responsible for the persistence of complex systems in these languages. Using data from multiple databases, we develop and employ a number of cross-linguistically applicable metrics to quantify the complexity of languages' numeral systems, and demonstrate that Indo-Aryan languages have decisively more complex numeral systems than the world's languages as a whole, though individual Indo-Aryan languages differ from each other in terms of the complexity of the patterns they display. We investigate the factors (e.g., religion, geographic isolation, etc.) that underlie complexity in numeral systems, with a focus on South Asia, in an attempt to develop an account of why complex numeral systems developed and persisted in certain Indo-Aryan languages but not elsewhere. Finally, we demonstrate that Indo-Aryan numeral systems adhere to certain general pressures toward efficient communication found cross-linguistically, despite their high complexity. We call for this somewhat overlooked dimension of complexity to be taken seriously when discussing general variation in numeral systems.

Discussion (0). Sign in to comment.

Pith tools

Reviewed May 22, 2026 · model on record in the stance chip above.