Pith. sign in

REVIEW 1 cited by

An efficient automated data analytics approach to large scale computational comparative linguistics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.11899 v1 pith:HP7VBSAE submitted 2020-01-31 cs.CL

classification cs.CL
keywords setstechniqueswordscuratedlargenumbersanalyticsautomated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This research project aimed to overcome the challenge of analysing human language relationships, facilitate the grouping of languages and formation of genealogical relationship between them by developing automated comparison techniques. Techniques were based on the phonetic representation of certain key words and concept. Example word sets included numbers 1-10 (curated), large database of numbers 1-10 and sheep counting numbers 1-10 (other sources), colours (curated), basic words (curated). To enable comparison within the sets the measure of Edit distance was calculated based on Levenshtein distance metric. This metric between two strings is the minimum number of single-character edits, operations including: insertions, deletions or substitutions. To explore which words exhibit more or less variation, which words are more preserved and examine how languages could be grouped based on linguistic distances within sets, several data analytics techniques were involved. Those included density evaluation, hierarchical clustering, silhouette, mean, standard deviation and Bhattacharya coefficient calculations. These techniques lead to the development of a workflow which was later implemented by combining Unix shell scripts, a developed R package and SWI Prolog. This proved to be computationally efficient and permitted the fast exploration of large language sets and their analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disparity between multipartite entangling and disentangling powers of unitaries: Even vs Odd

    quant-ph 2025-05 conditional novelty 6.0 of 10

    Non-diagonal unitary gates can have unequal multipartite entangling and disentangling powers, with the asymmetry appearing for even versus odd numbers of qubits.

Pith tools