REVIEW 2 major objections 5 minor 46 references
On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?
T0 review · 2 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Short primate sequences make random dependency parsing accurate enough that evaluation without gold trees becomes feasible.
desk verdict Clean combinatorial lower bounds show random free-tree parsers already hit high edge accuracy on short primate sequences, making evaluation without gold standard feasible there but hard for humans. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The London–Pluhár expectation that two uniform random labeled free trees on n vertices share 2(n−1)/n edges; averaging this quantity over the empirical or geometric length distribution of a corpus yields the lower bounds E[Q] and E[P_e^c].
What would settle it
Collect a new primate corpus whose length distribution is no longer short-tailed (or whose true trees are known by independent means) and check whether a random free-tree parser still recovers more than half the edges; if accuracy collapses or the true trees systematically deviate from the uniform model, the lower-bound argument fails.
Extended reading notes
Core claim
Because non-human primate sequence lengths decay rapidly (often geometrically), the expected fraction of correct undirected edges recovered by a random free-tree parser is high—above 50 % for geladas and above 79 % for chimpanzees—while the same quantity for unrestricted human sentences is only about 13 %. Therefore a lower bound on parser accuracy can be stated without any gold standard for those species, rendering evaluation feasible where it is hard for human language.
Load-bearing premise
Every sequence is assumed to possess exactly one correct free tree and all labeled free trees of the same size are equally likely a priori.
Editorial extensions
If this is right
- Any statistically informed unsupervised parser trained on gelada or chimpanzee sequences is guaranteed a high lower bound on undirected edge accuracy even without gold data.
- The same guarantee does not hold for ordinary human sentences, so unsupervised evaluation remains hard for human languages.
- Curriculum-style training that begins with short sequences is automatically favored by the natural length distribution of primate data.
- The same network-science calculation can be applied to any other species once its sequence-length histogram is known.
Reading between the lines
- The argument supplies a practical checklist—fast length decay plus some statistical structure—for deciding whether unsupervised parsing is worth attempting on a new animal communication system.
- If the uniform-tree prior is replaced by a more concentrated distribution over trees, the lower bounds would only improve, so the feasibility claim is conservative.
- The same length-driven accuracy effect may apply to other short-sequence domains (e.g., short social-media posts or early child language) that have been treated as hard for unsupervised parsing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that unsupervised free-tree dependency parsing of non-human primate vocal/gestural sequences can be evaluated without a gold standard. Because sequence lengths are short and decay rapidly (geometric-like), even a random parser that samples uniformly from Cayley’s labeled trees already recovers a high expected fraction of correct edges (E[Q] or E[P^e_c]). Any statistically informed (“good-enough”) parser therefore inherits a high lower bound on accuracy. Analytic expressions for E[P^t_c], E[Q] and E[P^e_c] are derived under uniform, geometric and empirical length distributions (Appendices B–E, using the London–Pluhár intersection formula). Empirical estimates reach 51 % (geladas) and >79 % (chimpanzees) for n_min=2, versus ~13–28 % on human PUD sentences; the same pattern holds for a 31-species ensemble when only n_max is known. Hence evaluation is feasible for primates but hard for human languages.
Significance. If the result holds, it removes a long-standing methodological barrier to applying dependency-parsing tools outside human language and supplies an immediately usable, parameter-free lower-bound protocol. Strengths include fully analytic derivations controlled to 10^{-8} error, transparent use of published multi-species length data, parallel human baselines (PUD/PUD10), and an explicit end-to-end methodology (§5.3). The work therefore has clear value for both computational linguistics and comparative communication research.
major comments (2)
- [abstract, §2.1, Appendices B–C, §5.1] The lower-bound argument (abstract, §2.3–2.6, §5.1) is sound under the stated premises, yet the abstract and conclusion phrase the result as an unconditional “must be high.” The claim is conditional on every sequence possessing a unique correct free tree (explicitly assumed in §2.1 and used to invoke Cayley + London–Pluhár in Appendices B–C). A brief paragraph acknowledging that the numerical guarantees collapse if animal sequences lack a unique tree (or are better modeled as forests/DAGs) would prevent over-reading.
- [§4.2, Table 4] Table 4 and the 31-species analysis rely on the uniform distribution as a lower bound for any non-increasing length distribution. While this is mathematically correct, the paper never verifies that the length distributions of the 29 species for which only n_max is known are in fact non-increasing. A single sentence noting that the bound is conservative only under that additional empirical premise would tighten the claim.
minor comments (5)
- [§1] Introduction contains multiple missing spaces (“asyntactic”, “atreebank”, “agold standard”, “Unsupervised parserslearn”).
- [front matter] Placeholder text remains: “Action editor: {action editor name}”, “Submission received: DD Month YYYY”.
- [Figure 4] Figure 4 caption refers to a “2-parameter geometric distribution” while the plotted curve uses the MVU estimator of Appendix A; a one-line clarification would help.
- [Tables 5–8, Appendix G] Tables 5–8 and G.1–G.2 are dense; adding a short note that all geometric expectations use ε=10^{-8} (Appendix E) would improve reproducibility.
- [§2.3] Eq. (19) defines E[Q]^* but is never used numerically; either drop it or show the comparison for the empirical cases.
Circularity Check
No significant circularity; random-parser edge-accuracy bounds follow by direct application of Cayley's formula and the London–Pluhár intersection expectation to given length distributions.
full rationale
The central derivation (Sections 2–4, Appendices B–C) starts from two explicit modeling premises—each sequence has a unique free tree and all labeled free trees of size n are equiprobable—then invokes the classical count n^{n-2} (Cayley) and the external network-science result E[m_c|n]=2(n-1)/n (London & Pluhár 2023) to obtain E[P^e_c|n]=2/n. Averaging over an empirical, uniform or geometric length distribution yields the reported lower bounds. These steps are pure calculation; no parameter is fitted to any accuracy figure, and the target claim (high random accuracy for short primate sequences) is not presupposed. Geometric q is estimated only for illustration; the load-bearing lower bounds use the uniform distribution (non-increasing) or raw empirical frequencies. Self-citations supply descriptive length data or prior observations of sequential structure, which serve as ordinary empirical inputs rather than load-bearing uniqueness theorems or ansätze. The derivation is therefore self-contained once the stated premises and external combinatorial facts are granted.
Assumptions & free parameters
assumptions (5)
- domain assumption Every sequence of length n possesses exactly one correct free (undirected) tree.
- domain assumption All labeled free trees on n vertices are equiprobable a priori.
- domain assumption Edge direction may be ignored, reducing the problem to free trees.
- domain assumption Sequence-length distributions of non-human primates are non-increasing (or geometric).
- ad hoc to paper A parser always returns exactly one free tree (ties broken uniformly at random).
Cite this review
Pith. "Pith review of On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?." pith.science (2026). https://pith.science/paper/ACCDVW3U
@misc{pith2026260706542,
author = {Pith},
title = {Pith review of: On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?},
year = {2026},
howpublished = {\url{https://pith.science/paper/ACCDVW3U}},
note = {Machine review of arXiv:2607.06542}
}
read the original abstract
Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, dependency parsing is unfeasible in other species. However, here we apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce due to the fast decay of the sequence length distribution. In contrast, human language sequences lack that property. Therefore, evaluation without a gold standard is feasible in non-human primates but a hard problem in humans.
Reference graph
Works this paper leans on
-
[1]
Alemany-Puig , Luís and Ramon Ferrer-i-Cancho . 2024. The expected sum of edge lengths in planar linearizations of trees. Journal of Language Modelling, 12(1):1--42
work page 2024
-
[2]
Altmann, Stuart A. 1965. Sociobiology of rhesus monkeys. II: Stochastics of social communication. J. Theor. Biol., 8:490--522
work page 1965
-
[3]
Blasi, Damián E., Joseph Henrich, Evangelia Adamou, David Kemmerer, and Asifa Majid. 2022. Over-reliance on English hinders cognitive science. Trends in Cognitive Sciences, 26(12):1153--1170
work page 2022
-
[4]
Campos, Daniel, Frederic Bartumeus, Vicenç Méndez, José S. Andrade, and Xavier Espadaler. 2016. Variability in individual activity bursts improves ant foraging success. Journal of The Royal Society Interface, 13(125):20160856
work page 2016
-
[5]
Carroll, John. 2014. Parsing. In The Oxford Handbook of Computational Linguistics. Oxford University Press
work page 2014
-
[6]
Cayley, Arthur. 1889. A theorem on trees. Quart. J. Math, 23:376--378
-
[7]
Ferrer-i-Cancho , Ramon, Carlos G\'omez-Rodr\'iguez , Juan Luis Esteban, and Lluís Alemany-Puig . 2022. Optimality of syntactic dependency distances. Physical Review E, 105(1):014308
work page 2022
-
[8]
Ferrer-i-Cancho , Ramon and David Lusseau. 2006. Long-term correlations in the surface behavior of dolphins. Europhysics Letters, 74(6):1095--1101
work page 2006
Show all 46 references
-
[9]
Ferrer-i-Cancho , Ramon and B. McCowan. 2012. The span of dependencies in dolphin whistle sequences. Journal of Statistical Mechanics, page P06002
2012
-
[10]
and Morten H
Frank, Stefan L. and Morten H. Christiansen. 2018. Hierarchical and sequential processing of language. Language, Cognition and Neuroscience, 33(9):1213--1218
2018
-
[11]
Furuhashi, Sho and Yoshinori Hayakawa. 2012. Lognormality of the distribution of Japanese sentence lengths. Journal of the Physical Society of Japan, 81(3):034004
2012
-
[12]
Gerdes, Kim, Bruno Guillaume, Sylvain Kahane, and Guy Perrier. 2018. SUD or surface-syntactic universal dependencies: An annotation scheme near-isomorphic to UD . In Proceedings of the Second Workshop on Universal Dependencies ( UDW 2018) , pages 66--74, Association for Comput...
2018
-
[13]
Friederici, Roman M
Girard-Buttoz, Cédric, Emiliano Zaccarella, Tatiana Bortolato, Angela D. Friederici, Roman M. Wittig, and Catherine Crockford. 2022. Chimpanzees produce diverse vocal sequences with ordered and recombinatorial properties. Communications Biology, 5(1)
2022
-
[14]
Graham, Alexandra Safryghin, and Catherine Hobaiter
Grund, Charlotte, Gal Badihi, Kirsty E. Graham, Alexandra Safryghin, and Catherine Hobaiter. 2023. Gesturalorigins: A bottom-up framework for establishing systematic gesture data across ape species. Behavior Research Methods
2023
-
[15]
Robbins, and Catherine Hobaiter
Grund, Charlotte, Martha M. Robbins, and Catherine Hobaiter. 2025. The gestural repertoire of Bwindi mountain gorillas ( Gorilla beringei beringei): gesture form and frequency of use. Animal Cognition, 28(1)
2025
-
[16]
Gustison, Morgan. 2017. The phylogeny and function of vocal complexity in geladas. Phd thesis, University of Michigan, Michigan, USA
2017
-
[17]
Gustison, Morgan L., Stuart Semple, Ramon Ferrer-i-Cancho , and Thore Bergman. 2016. Gelada vocal sequences follow Menzerath 's linguistic law. Proceedings of the National Academy of Sciences USA, 13(19):E2750--E2758
2016
-
[18]
Han, Wenjuan, Yong Jiang, Hwee Tou Ng, and Kewei Tu. 2020. A survey of unsupervised dependency parsing. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2522--2533, International Committee on Computational Linguistics, Barcelona, Spain (Online)
2020
-
[19]
Hobaiter, C., R. W. Byrne, and K. Zuberb\" u hler. 2017. Wild chimpanzees’ use of single and combined vocal and gestural signals. Behavioral Ecology and Sociobiology, 71(6)
2017
-
[20]
Hobaiter, Catherine and Richard W. Byrne. 2011. Serial gesturing by wild chimpanzees: its nature and function for communication. Animal Cognition, 14(6):827–838
2011
-
[21]
Blumstein, Marie A
Kershenbaum, Arik, Daniel T. Blumstein, Marie A. Roch, C a g lar Ak c ay, Gregory Backus, Mark A. Bee, Kirsten Bohn, Yan Cao, Gerald Carter, Cristiane C\" a sar, Michael Coen, Stacy L. DeRuiter, Laurance Doyle, Shimon Edelman, Ramon Ferrer-i Cancho, Todd M. Freeberg, Ellen C. ...
2016
-
[22]
Bowles, Todd M
Kershenbaum, Arik, Ann E. Bowles, Todd M. Freeberg, Dezhe Z. Jin, Adriano R. Lameira, and Kirsten Bohn. 2014. Animal vocal sequences: not the markov chains we thought they were. Proceedings of the Royal Society B: Biological Sciences, 281(1792):20141370
2014
-
[23]
Klein, Dan and Christopher Manning. 2004. Corpus-based induction of syntactic structure: Models of dependency and constituency. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics ( ACL -04) , pages 478--485, Barcelona, Spain
2004
-
[24]
Liebal, Katja, Josep Call, and Michael Tomasello. 2004. Use of gesture sequences in chimpanzees. American Journal of Primatology, 64(4):377--396
2004
-
[25]
London, András and András Pluhár. 2023. Intersection of random spanning trees in complex networks. Applied Network Science, 8(1)
2023
-
[26]
Marecek, David. 2012. Unsupervised Dependency Parsing. Ph.D. thesis, Charles University in Prague
2012
-
[27]
Marecek, David. 2016. Twelve years of unsupervised dependency parsing. In Proceedings of the 16th ITAT Conference Information Technologies - Applications and Theory, Tatransk \' e Matliare, Slovakia, September 15-19, 2016 , volume 1649 of CEUR Workshop Proceedings , pages 56--...
2016
-
[28]
Mart \'i n Rodr \'i guez, Lorena, Tatiana Merzhevich, Wellington Silva, Tiago Tresoldi, Carolina Aragon, and Fabr \'i cio F. Gerardi. 2022. Tup \'i an language ressources: Data, tools, analyses. In Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group o...
2022
-
[29]
Hanser, and Laurance R
McCowan, Brenda, Sean F. Hanser, and Laurance R. Doyle. 1999. Quantitative tools for comparing animal communication systems: information theory applied to bottlenose dolphin whistle repertoires. Animal Behaviour, 57:409--419
1999
-
[30]
McDonald, Ryan, Fernando Pereira, Kiril Ribarov, and Jan Haji c . 2005. Non-projective dependency parsing using spanning tree algorithms. In Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Processing, HLT '05, pages 523--530...
2005
-
[31]
Mel' c uk, Igor. 1988. Dependency syntax: theory and practice. State of New York University Press, Albany
1988
-
[32]
Graham, Chie Hashimoto, Joseph G
Mielke, Alexander, Gal Badihi, Ed Donnellan, Kirsty E. Graham, Chie Hashimoto, Joseph G. Mine, Alex K. Piel, Alexandra Safryghin, Katie E. Slocombe, Adrian Soldati, Fiona A. Stewart, Simon W. Townsend, Claudia Wilke, Klaus Zuberb \"u hler, Chiara Zulberti, and Catherine Hobait...
2024
-
[33]
Graham, Charlotte Grund, Chie Hashimoto, Alex K
Mielke, Alexander, Gal Badihi, Kirsty E. Graham, Charlotte Grund, Chie Hashimoto, Alex K. Piel, Alexandra Safryghin, Katie E. Slocombe, Fiona Stewart, Claudia Wilke, Klaus Zuberb\" u hler, and Catherine Hobaiter. 2024 b . Many morphs: Parsing gesture signals from the noise. Be...
2024
-
[34]
Bosshard, Sabine Stoll, Zarin P
Mine, Joseph G., Claudia Wilke, Chiara Zulberti, Melika Behjati, Alexandra B. Bosshard, Sabine Stoll, Zarin P. Machanda, Andri Manser, Katie E. Slocombe, and Simon W. Townsend. 2024. Vocal-visual combinations in wild chimpanzees. Behavioral Ecology and Sociobiology, 78(10)
2024
-
[35]
Newman, Mark E. J. 2010. Networks. An introduction . Oxford University Press, Oxford
2010
-
[36]
Park, Chanseok and Min Wang. 2023. A study on the g and h control charts. Communications in Statistics - Theory and Methods, 52(20):7334--7349
2023
-
[37]
Petrini, Sonia and Ramon Ferrer-i-Cancho . 2025. The distribution of syntactic dependency distances. Glottometrics, 58:35--94
2025
-
[38]
Scheinerman, Edward R. 2012. Mathematics: A Discrete Introduction, 3rd ed. edition. Cengage Learning
2012
-
[39]
Sigurd, Bengt, Mats Eeg-Olofsson, and Joost van Weijer . 2004. Word length, sentence length and frequency - Zipf revisited. Studia Linguistica, 58(1):37--52
2004
-
[40]
S gaard, Anders. 2011. From ranked words to dependency trees: two-stage unsupervised non-projective dependency parsing. In Proceedings of T ext G raphs-6: Graph-based Methods for Natural Language Processing , pages 60--68, Association for Computational Linguistics, Portland, Oregon
2011
-
[41]
Spitkovsky, Valentin I., Hiyan Alshawi, and Daniel Jurafsky. 2010. From baby steps to leapfrog: How ``less is more'' in unsupervised dependency parsing. In Human Language Technologies: The 2010 Annual Conference of the North A merican Chapter of the Association for Computation...
2010
-
[42]
Tu, Kewei and Vasant Honavar. 2011. On the utility of curricula in unsupervised learning of probabilistic grammars. In Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence - Volume Volume Two, IJCAI'11, page 1523–1528, AAAI Press
2011
-
[43]
Vasquez, Alonso, Renzo Ego Aguirre, Candy Angulo, John Miller, Claudia Villanueva, Z eljko Agi \'c , Roberto Zariquiey, and Arturo Oncevay. 2018. Toward U niversal D ependencies for S hipibo-konibo. In Proceedings of the Second Workshop on Universal Dependencies ( UDW 2018) , ...
2018
-
[44]
Yuret, D. 1998 a . Discovery of linguistic relations using lexical attraction. Ph.D. thesis, Massachusets Institute of Technology, USA
1998
-
[45]
Yuret, Deniz. 1998 b . Discovery of Linguistic Relations Using Lexical Attraction. Ph.D. thesis, MIT
1998
-
[46]
Zeman, Daniel, Joakim Nivre, Mitchell Abrams, Elia Ackermann, No \"e mi Aepli, Z eljko Agi \'c , Lars Ahrenberg, Chika Kennedy Ajede, Gabriel \.e Aleksandravi c i \=u t \.e , Lene Antonsen, et al. 2020. Universal dependencies 2.6. LINDAT / CLARIAH - CZ digital library at the I...
2020
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.