Pith. sign in

REVIEW 1 cited by

Phonetic Word Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.14796 v1 pith:TZURPOD6 submitted 2021-09-30 cs.CL

classification cs.CL
keywords phoneticsimilarityembeddingmethodologynovelpresentedspacetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work presents a novel methodology for calculating the phonetic similarity between words taking motivation from the human perception of sounds. This metric is employed to learn a continuous vector embedding space that groups similar sounding words together and can be used for various downstream computational phonology tasks. The efficacy of the method is presented for two different languages (English, Hindi) and performance gains over previous reported works are discussed on established tests for predicting phonetic similarity. To address limited benchmarking mechanisms in this field, we also introduce a heterographic pun dataset based evaluation methodology to compare the effectiveness of acoustic similarity algorithms. Further, a visualization of the embedding space is presented with a discussion on the various possible use-cases of this novel algorithm. An open-source implementation is also shared to aid reproducibility and enable adoption in related tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Cognate Data Bottleneck in Language Phylogenetics

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Automatic extraction of cognate matrices from BabelNet yields sparse, noisy data with GQ distances above 0.39 from the Glottolog gold standard, confirming a data bottleneck for computational historical linguistics.

Pith tools