Pith. sign in

REVIEW 2 major objections 5 minor 46 references

On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?

T0 review · 2 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read Short primate sequences make random dependency parsing accurate enough that evaluation without gold trees becomes feasible.

desk verdict Clean combinatorial lower bounds show random free-tree parsers already hit high edge accuracy on short primate sequences, making evaluation without gold standard feasible there but hard for humans. read the letter →

arxiv 2607.06542 v2 pith:ACCDVW3U submitted 2026-07-07 cs.CL

classification cs.CL
keywords dependencyparsingunsupervisedevaluationwithoutgoldstandardsequencelengthdistributionnon-humanprimatecommunicationfreetreesnetworkscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether one can evaluate an unsupervised dependency parser on animal vocal or gestural sequences when no gold-standard trees exist. It shows that a random free-tree parser already recovers a large fraction of correct undirected edges simply because those sequences are short and their length distribution decays exponentially. For gelada and chimpanzee data the expected accuracy sits between 50 % and 80 %; for human sentences of unrestricted length it falls to roughly 13 %. Consequently any parser that beats chance is guaranteed a high lower bound on accuracy in the animal case, while the same guarantee is unavailable for ordinary human text. The result turns the usual pessimism about non-human syntax on its head: evaluation without gold data is actually easier for other primates than for us.

What carries the argument

The London–Pluhár expectation that two uniform random labeled free trees on n vertices share 2(n−1)/n edges; averaging this quantity over the empirical or geometric length distribution of a corpus yields the lower bounds E[Q] and E[P_e^c].

What would settle it

Collect a new primate corpus whose length distribution is no longer short-tailed (or whose true trees are known by independent means) and check whether a random free-tree parser still recovers more than half the edges; if accuracy collapses or the true trees systematically deviate from the uniform model, the lower-bound argument fails.

Watch

Extended reading notes

Core claim

Because non-human primate sequence lengths decay rapidly (often geometrically), the expected fraction of correct undirected edges recovered by a random free-tree parser is high—above 50 % for geladas and above 79 % for chimpanzees—while the same quantity for unrestricted human sentences is only about 13 %. Therefore a lower bound on parser accuracy can be stated without any gold standard for those species, rendering evaluation feasible where it is hard for human language.

Load-bearing premise

Every sequence is assumed to possess exactly one correct free tree and all labeled free trees of the same size are equally likely a priori.

Editorial extensions

If this is right

  • Any statistically informed unsupervised parser trained on gelada or chimpanzee sequences is guaranteed a high lower bound on undirected edge accuracy even without gold data.
  • The same guarantee does not hold for ordinary human sentences, so unsupervised evaluation remains hard for human languages.
  • Curriculum-style training that begins with short sequences is automatically favored by the natural length distribution of primate data.
  • The same network-science calculation can be applied to any other species once its sequence-length histogram is known.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The argument supplies a practical checklist—fast length decay plus some statistical structure—for deciding whether unsupervised parsing is worth attempting on a new animal communication system.
  • If the uniform-tree prior is replaced by a more concentrated distribution over trees, the lower bounds would only improve, so the feasibility claim is conservative.
  • The same length-driven accuracy effect may apply to other short-sequence domains (e.g., short social-media posts or early child language) that have been treated as hard for unsupervised parsing.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper claims that unsupervised free-tree dependency parsing of non-human primate vocal/gestural sequences can be evaluated without a gold standard. Because sequence lengths are short and decay rapidly (geometric-like), even a random parser that samples uniformly from Cayley’s labeled trees already recovers a high expected fraction of correct edges (E[Q] or E[P^e_c]). Any statistically informed (“good-enough”) parser therefore inherits a high lower bound on accuracy. Analytic expressions for E[P^t_c], E[Q] and E[P^e_c] are derived under uniform, geometric and empirical length distributions (Appendices B–E, using the London–Pluhár intersection formula). Empirical estimates reach 51 % (geladas) and >79 % (chimpanzees) for n_min=2, versus ~13–28 % on human PUD sentences; the same pattern holds for a 31-species ensemble when only n_max is known. Hence evaluation is feasible for primates but hard for human languages.

Significance. If the result holds, it removes a long-standing methodological barrier to applying dependency-parsing tools outside human language and supplies an immediately usable, parameter-free lower-bound protocol. Strengths include fully analytic derivations controlled to 10^{-8} error, transparent use of published multi-species length data, parallel human baselines (PUD/PUD10), and an explicit end-to-end methodology (§5.3). The work therefore has clear value for both computational linguistics and comparative communication research.

major comments (2)
  1. [abstract, §2.1, Appendices B–C, §5.1] The lower-bound argument (abstract, §2.3–2.6, §5.1) is sound under the stated premises, yet the abstract and conclusion phrase the result as an unconditional “must be high.” The claim is conditional on every sequence possessing a unique correct free tree (explicitly assumed in §2.1 and used to invoke Cayley + London–Pluhár in Appendices B–C). A brief paragraph acknowledging that the numerical guarantees collapse if animal sequences lack a unique tree (or are better modeled as forests/DAGs) would prevent over-reading.
  2. [§4.2, Table 4] Table 4 and the 31-species analysis rely on the uniform distribution as a lower bound for any non-increasing length distribution. While this is mathematically correct, the paper never verifies that the length distributions of the 29 species for which only n_max is known are in fact non-increasing. A single sentence noting that the bound is conservative only under that additional empirical premise would tighten the claim.
minor comments (5)
  1. [§1] Introduction contains multiple missing spaces (“asyntactic”, “atreebank”, “agold standard”, “Unsupervised parserslearn”).
  2. [front matter] Placeholder text remains: “Action editor: {action editor name}”, “Submission received: DD Month YYYY”.
  3. [Figure 4] Figure 4 caption refers to a “2-parameter geometric distribution” while the plotted curve uses the MVU estimator of Appendix A; a one-line clarification would help.
  4. [Tables 5–8, Appendix G] Tables 5–8 and G.1–G.2 are dense; adding a short note that all geometric expectations use ε=10^{-8} (Appendix E) would improve reproducibility.
  5. [§2.3] Eq. (19) defines E[Q]^* but is never used numerically; either drop it or show the comparison for the empirical cases.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; random-parser edge-accuracy bounds follow by direct application of Cayley's formula and the London–Pluhár intersection expectation to given length distributions.

full rationale

The central derivation (Sections 2–4, Appendices B–C) starts from two explicit modeling premises—each sequence has a unique free tree and all labeled free trees of size n are equiprobable—then invokes the classical count n^{n-2} (Cayley) and the external network-science result E[m_c|n]=2(n-1)/n (London & Pluhár 2023) to obtain E[P^e_c|n]=2/n. Averaging over an empirical, uniform or geometric length distribution yields the reported lower bounds. These steps are pure calculation; no parameter is fitted to any accuracy figure, and the target claim (high random accuracy for short primate sequences) is not presupposed. Geometric q is estimated only for illustration; the load-bearing lower bounds use the uniform distribution (non-increasing) or raw empirical frequencies. Self-citations supply descriptive length data or prior observations of sequential structure, which serve as ordinary empirical inputs rather than load-bearing uniqueness theorems or ansätze. The derivation is therefore self-contained once the stated premises and external combinatorial facts are granted.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on three domain assumptions (unique free tree per sequence, equiprobable labeled trees, free rather than rooted trees) plus the empirical fact of fast-decaying length distributions. No free parameters are fitted to force the result; geometric q is estimated only for secondary illustration. No new entities are postulated.

assumptions (5)
  • domain assumption Every sequence of length n possesses exactly one correct free (undirected) tree.
    Stated in §2.1; required for the notions of 'correct tree' and 'correct edges' to be well-defined for non-human sequences.
  • domain assumption All labeled free trees on n vertices are equiprobable a priori.
    Explicitly adopted in §2.3 to remain 'maximally agnostic'; yields p(correct tree) = n^{2-n} via Cayley's formula.
  • domain assumption Edge direction may be ignored, reducing the problem to free trees.
    Justified in §2 as common practice in unsupervised parsing evaluation and as avoiding a priori hierarchy assumptions for other species.
  • domain assumption Sequence-length distributions of non-human primates are non-increasing (or geometric).
    Supported by the empirical plots in §4.1 and used to justify the uniform distribution as a lower-bound model for the 31-species table.
  • ad hoc to paper A parser always returns exactly one free tree (ties broken uniformly at random).
    Introduced in §2.1 to equate precision and recall and simplify the expectation calculations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?." pith.science (2026). https://pith.science/paper/ACCDVW3U

@misc{pith2026260706542,
  author       = {Pith},
  title        = {Pith review of: On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ACCDVW3U}},
  note         = {Machine review of arXiv:2607.06542}
}
read the original abstract

Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, dependency parsing is unfeasible in other species. However, here we apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce due to the fast decay of the sequence length distribution. In contrast, human language sequences lack that property. Therefore, evaluation without a gold standard is feasible in non-human primates but a hard problem in humans.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 46 canonical work pages

  1. [1]

    Alemany-Puig , Luís and Ramon Ferrer-i-Cancho . 2024. The expected sum of edge lengths in planar linearizations of trees. Journal of Language Modelling, 12(1):1--42

  2. [2]

    Altmann, Stuart A. 1965. Sociobiology of rhesus monkeys. II: Stochastics of social communication. J. Theor. Biol., 8:490--522

  3. [3]

    Blasi, Damián E., Joseph Henrich, Evangelia Adamou, David Kemmerer, and Asifa Majid. 2022. Over-reliance on English hinders cognitive science. Trends in Cognitive Sciences, 26(12):1153--1170

  4. [4]

    Andrade, and Xavier Espadaler

    Campos, Daniel, Frederic Bartumeus, Vicenç Méndez, José S. Andrade, and Xavier Espadaler. 2016. Variability in individual activity bursts improves ant foraging success. Journal of The Royal Society Interface, 13(125):20160856

  5. [5]

    Carroll, John. 2014. Parsing. In The Oxford Handbook of Computational Linguistics. Oxford University Press

  6. [6]

    Cayley, Arthur. 1889. A theorem on trees. Quart. J. Math, 23:376--378

  7. [7]

    Ferrer-i-Cancho , Ramon, Carlos G\'omez-Rodr\'iguez , Juan Luis Esteban, and Lluís Alemany-Puig . 2022. Optimality of syntactic dependency distances. Physical Review E, 105(1):014308

  8. [8]

    Ferrer-i-Cancho , Ramon and David Lusseau. 2006. Long-term correlations in the surface behavior of dolphins. Europhysics Letters, 74(6):1095--1101

Show all 46 references
  1. [9]

    Ferrer-i-Cancho , Ramon and B. McCowan. 2012. The span of dependencies in dolphin whistle sequences. Journal of Statistical Mechanics, page P06002

  2. [10]

    and Morten H

    Frank, Stefan L. and Morten H. Christiansen. 2018. Hierarchical and sequential processing of language. Language, Cognition and Neuroscience, 33(9):1213--1218

  3. [11]

    Furuhashi, Sho and Yoshinori Hayakawa. 2012. Lognormality of the distribution of Japanese sentence lengths. Journal of the Physical Society of Japan, 81(3):034004

  4. [12]

    Gerdes, Kim, Bruno Guillaume, Sylvain Kahane, and Guy Perrier. 2018. SUD or surface-syntactic universal dependencies: An annotation scheme near-isomorphic to UD . In Proceedings of the Second Workshop on Universal Dependencies ( UDW 2018) , pages 66--74, Association for Comput...

  5. [13]

    Friederici, Roman M

    Girard-Buttoz, Cédric, Emiliano Zaccarella, Tatiana Bortolato, Angela D. Friederici, Roman M. Wittig, and Catherine Crockford. 2022. Chimpanzees produce diverse vocal sequences with ordered and recombinatorial properties. Communications Biology, 5(1)

  6. [14]

    Graham, Alexandra Safryghin, and Catherine Hobaiter

    Grund, Charlotte, Gal Badihi, Kirsty E. Graham, Alexandra Safryghin, and Catherine Hobaiter. 2023. Gesturalorigins: A bottom-up framework for establishing systematic gesture data across ape species. Behavior Research Methods

  7. [15]

    Robbins, and Catherine Hobaiter

    Grund, Charlotte, Martha M. Robbins, and Catherine Hobaiter. 2025. The gestural repertoire of Bwindi mountain gorillas ( Gorilla beringei beringei): gesture form and frequency of use. Animal Cognition, 28(1)

  8. [16]

    Gustison, Morgan. 2017. The phylogeny and function of vocal complexity in geladas. Phd thesis, University of Michigan, Michigan, USA

  9. [17]

    Gustison, Morgan L., Stuart Semple, Ramon Ferrer-i-Cancho , and Thore Bergman. 2016. Gelada vocal sequences follow Menzerath 's linguistic law. Proceedings of the National Academy of Sciences USA, 13(19):E2750--E2758

  10. [18]

    Han, Wenjuan, Yong Jiang, Hwee Tou Ng, and Kewei Tu. 2020. A survey of unsupervised dependency parsing. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2522--2533, International Committee on Computational Linguistics, Barcelona, Spain (Online)

  11. [19]

    Hobaiter, C., R. W. Byrne, and K. Zuberb\" u hler. 2017. Wild chimpanzees’ use of single and combined vocal and gestural signals. Behavioral Ecology and Sociobiology, 71(6)

  12. [20]

    Hobaiter, Catherine and Richard W. Byrne. 2011. Serial gesturing by wild chimpanzees: its nature and function for communication. Animal Cognition, 14(6):827–838

  13. [21]

    Blumstein, Marie A

    Kershenbaum, Arik, Daniel T. Blumstein, Marie A. Roch, C a g lar Ak c ay, Gregory Backus, Mark A. Bee, Kirsten Bohn, Yan Cao, Gerald Carter, Cristiane C\" a sar, Michael Coen, Stacy L. DeRuiter, Laurance Doyle, Shimon Edelman, Ramon Ferrer-i Cancho, Todd M. Freeberg, Ellen C. ...

  14. [22]

    Bowles, Todd M

    Kershenbaum, Arik, Ann E. Bowles, Todd M. Freeberg, Dezhe Z. Jin, Adriano R. Lameira, and Kirsten Bohn. 2014. Animal vocal sequences: not the markov chains we thought they were. Proceedings of the Royal Society B: Biological Sciences, 281(1792):20141370

  15. [23]

    Klein, Dan and Christopher Manning. 2004. Corpus-based induction of syntactic structure: Models of dependency and constituency. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics ( ACL -04) , pages 478--485, Barcelona, Spain

  16. [24]

    Liebal, Katja, Josep Call, and Michael Tomasello. 2004. Use of gesture sequences in chimpanzees. American Journal of Primatology, 64(4):377--396

  17. [25]

    London, András and András Pluhár. 2023. Intersection of random spanning trees in complex networks. Applied Network Science, 8(1)

  18. [26]

    Marecek, David. 2012. Unsupervised Dependency Parsing. Ph.D. thesis, Charles University in Prague

  19. [27]

    Marecek, David. 2016. Twelve years of unsupervised dependency parsing. In Proceedings of the 16th ITAT Conference Information Technologies - Applications and Theory, Tatransk \' e Matliare, Slovakia, September 15-19, 2016 , volume 1649 of CEUR Workshop Proceedings , pages 56--...

  20. [28]

    Mart \'i n Rodr \'i guez, Lorena, Tatiana Merzhevich, Wellington Silva, Tiago Tresoldi, Carolina Aragon, and Fabr \'i cio F. Gerardi. 2022. Tup \'i an language ressources: Data, tools, analyses. In Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group o...

  21. [29]

    Hanser, and Laurance R

    McCowan, Brenda, Sean F. Hanser, and Laurance R. Doyle. 1999. Quantitative tools for comparing animal communication systems: information theory applied to bottlenose dolphin whistle repertoires. Animal Behaviour, 57:409--419

  22. [30]

    McDonald, Ryan, Fernando Pereira, Kiril Ribarov, and Jan Haji c . 2005. Non-projective dependency parsing using spanning tree algorithms. In Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Processing, HLT '05, pages 523--530...

  23. [31]

    Mel' c uk, Igor. 1988. Dependency syntax: theory and practice. State of New York University Press, Albany

  24. [32]

    Graham, Chie Hashimoto, Joseph G

    Mielke, Alexander, Gal Badihi, Ed Donnellan, Kirsty E. Graham, Chie Hashimoto, Joseph G. Mine, Alex K. Piel, Alexandra Safryghin, Katie E. Slocombe, Adrian Soldati, Fiona A. Stewart, Simon W. Townsend, Claudia Wilke, Klaus Zuberb \"u hler, Chiara Zulberti, and Catherine Hobait...

  25. [33]

    Graham, Charlotte Grund, Chie Hashimoto, Alex K

    Mielke, Alexander, Gal Badihi, Kirsty E. Graham, Charlotte Grund, Chie Hashimoto, Alex K. Piel, Alexandra Safryghin, Katie E. Slocombe, Fiona Stewart, Claudia Wilke, Klaus Zuberb\" u hler, and Catherine Hobaiter. 2024 b . Many morphs: Parsing gesture signals from the noise. Be...

  26. [34]

    Bosshard, Sabine Stoll, Zarin P

    Mine, Joseph G., Claudia Wilke, Chiara Zulberti, Melika Behjati, Alexandra B. Bosshard, Sabine Stoll, Zarin P. Machanda, Andri Manser, Katie E. Slocombe, and Simon W. Townsend. 2024. Vocal-visual combinations in wild chimpanzees. Behavioral Ecology and Sociobiology, 78(10)

  27. [35]

    Newman, Mark E. J. 2010. Networks. An introduction . Oxford University Press, Oxford

  28. [36]

    Park, Chanseok and Min Wang. 2023. A study on the g and h control charts. Communications in Statistics - Theory and Methods, 52(20):7334--7349

  29. [37]

    Petrini, Sonia and Ramon Ferrer-i-Cancho . 2025. The distribution of syntactic dependency distances. Glottometrics, 58:35--94

  30. [38]

    Scheinerman, Edward R. 2012. Mathematics: A Discrete Introduction, 3rd ed. edition. Cengage Learning

  31. [39]

    Sigurd, Bengt, Mats Eeg-Olofsson, and Joost van Weijer . 2004. Word length, sentence length and frequency - Zipf revisited. Studia Linguistica, 58(1):37--52

  32. [40]

    S gaard, Anders. 2011. From ranked words to dependency trees: two-stage unsupervised non-projective dependency parsing. In Proceedings of T ext G raphs-6: Graph-based Methods for Natural Language Processing , pages 60--68, Association for Computational Linguistics, Portland, Oregon

  33. [41]

    Spitkovsky, Valentin I., Hiyan Alshawi, and Daniel Jurafsky. 2010. From baby steps to leapfrog: How ``less is more'' in unsupervised dependency parsing. In Human Language Technologies: The 2010 Annual Conference of the North A merican Chapter of the Association for Computation...

  34. [42]

    Tu, Kewei and Vasant Honavar. 2011. On the utility of curricula in unsupervised learning of probabilistic grammars. In Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence - Volume Volume Two, IJCAI'11, page 1523–1528, AAAI Press

  35. [43]

    Vasquez, Alonso, Renzo Ego Aguirre, Candy Angulo, John Miller, Claudia Villanueva, Z eljko Agi \'c , Roberto Zariquiey, and Arturo Oncevay. 2018. Toward U niversal D ependencies for S hipibo-konibo. In Proceedings of the Second Workshop on Universal Dependencies ( UDW 2018) , ...

  36. [44]

    Yuret, D. 1998 a . Discovery of linguistic relations using lexical attraction. Ph.D. thesis, Massachusets Institute of Technology, USA

  37. [45]

    Yuret, Deniz. 1998 b . Discovery of Linguistic Relations Using Lexical Attraction. Ph.D. thesis, MIT

  38. [46]

    Zeman, Daniel, Joakim Nivre, Mitchell Abrams, Elia Ackermann, No \"e mi Aepli, Z eljko Agi \'c , Lars Ahrenberg, Chika Kennedy Ajede, Gabriel \.e Aleksandravi c i \=u t \.e , Lene Antonsen, et al. 2020. Universal dependencies 2.6. LINDAT / CLARIAH - CZ digital library at the I...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.