REVIEW 2 cited by
How Familiar Does That Sound? Cross-Lingual Representational Similarity Analysis of Acoustic Word Embeddings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
How do neural networks "perceive" speech sounds from unknown languages? Does the typological similarity between the model's training language (L1) and an unknown language (L2) have an impact on the model representations of L2 speech signals? To answer these questions, we present a novel experimental design based on representational similarity analysis (RSA) to analyze acoustic word embeddings (AWEs) -- vector representations of variable-duration spoken-word segments. First, we train monolingual AWE models on seven Indo-European languages with various degrees of typological similarity. We then employ RSA to quantify the cross-lingual similarity by simulating native and non-native spoken-word processing using AWEs. Our experiments show that typological similarity indeed affects the representational similarity of the models in our study. We further discuss the implications of our work on modeling speech processing and language similarity with neural networks.
Forward citations
Cited by 2 Pith papers
-
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
A new probe that bypasses the fully-connected layer shows convolutional layers of a small-bottleneck WaveGAN can reflect a phonotactic restriction learned from lexical training data.
-
Exploring the encoding of linguistic representations in the Fully-Connected Layer of generative CNNs for Speech
Vowel-like columns of the fully connected layer of a speech GAN cluster by phonetic similarity across different words, and can be recombined into speech-like segments.
Discussion (0). Continue with ORCID to comment.