Pith. sign in

REVIEW 2 cited by

How Familiar Does That Sound? Cross-Lingual Representational Similarity Analysis of Acoustic Word Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.10179 v1 pith:QQW7JLIF submitted 2021-09-21 cs.CL

classification cs.CL
keywords similaritylanguagerepresentationalspeechtypologicalacousticanalysisawes
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

How do neural networks "perceive" speech sounds from unknown languages? Does the typological similarity between the model's training language (L1) and an unknown language (L2) have an impact on the model representations of L2 speech signals? To answer these questions, we present a novel experimental design based on representational similarity analysis (RSA) to analyze acoustic word embeddings (AWEs) -- vector representations of variable-duration spoken-word segments. First, we train monolingual AWE models on seven Indo-European languages with various degrees of typological similarity. We then employ RSA to quantify the cross-lingual similarity by simulating native and non-native spoken-word processing using AWEs. Our experiments show that typological similarity indeed affects the representational similarity of the models in our study. We further discuss the implications of our work on modeling speech processing and language similarity with neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new probe that bypasses the fully-connected layer shows convolutional layers of a small-bottleneck WaveGAN can reflect a phonotactic restriction learned from lexical training data.

  2. Exploring the encoding of linguistic representations in the Fully-Connected Layer of generative CNNs for Speech

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Vowel-like columns of the fully connected layer of a speech GAN cluster by phonetic similarity across different words, and can be recombined into speech-like segments.

Pith tools