REVIEW 2 cited by
Acoustic Neighbor Embeddings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper proposes a novel acoustic word embedding called Acoustic Neighbor Embeddings where speech or text of arbitrary length are mapped to a vector space of fixed, reduced dimensions by adapting stochastic neighbor embedding (SNE) to sequential inputs. The Euclidean distance between coordinates in the embedding space reflects the phonetic confusability between their corresponding sequences. Two encoder neural networks are trained: an acoustic encoder that accepts speech signals in the form of frame-wise subword posterior probabilities obtained from an acoustic model and a text encoder that accepts text in the form of subword transcriptions. Compared to a triplet loss criterion, the proposed method is shown to have more effective gradients for neural network training. Experimentally, it also gives more accurate results with low-dimensional embeddings when the two encoder networks are used in tandem in a word (name) recognition task, and when the text encoder network is used standalone in an approximate phonetic matching task. In particular, in an isolated name recognition task depending solely on Euclidean nearest-neighbor search between the proposed embedding vectors, the recognition accuracy is identical to that of conventional finite state transducer(FST)-based decoding using test data with up to 1 million names in the vocabulary and 40 dimensions in the embeddings.
Forward citations
Cited by 2 Pith papers
-
A Theoretical Framework for Acoustic Neighbor Embeddings
Euclidean distances between acoustic neighbor embeddings are interpreted as phonetic similarity through a Bayes-error and Gaussian-isotropy approximation, with four validation experiments.
-
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
CLASP aligns speech and text embeddings with contrastive learning, reporting strong retrieval scores on a new multilingual benchmark while only partially beating ASR-based retrieval baselines.
Discussion (0). Continue with ORCID to comment.