Pith. sign in

REVIEW 1 cited by

Matching Latent Encoding for Audio-Text based Keyword Spotting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.05245 v1 pith:ZW7QPK3O submitted 2023-06-08 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords embeddingstextaudiosequencearchitecturekeywordresultsspotting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using audio and text embeddings jointly for Keyword Spotting (KWS) has shown high-quality results, but the key challenge of how to semantically align two embeddings for multi-word keywords of different sequence lengths remains largely unsolved. In this paper, we propose an audio-text-based end-to-end model architecture for flexible keyword spotting (KWS), which builds upon learned audio and text embeddings. Our architecture uses a novel dynamic programming-based algorithm, Dynamic Sequence Partitioning (DSP), to optimally partition the audio sequence into the same length as the word-based text sequence using the monotonic alignment of spoken content. Our proposed model consists of an encoder block to get audio and text embeddings, a projector block to project individual embeddings to a common latent space, and an audio-text aligner containing a novel DSP algorithm, which aligns the audio and text embeddings to determine if the spoken content is the same as the text. Experimental results show that our DSP is more effective than other partitioning schemes, and the proposed architecture outperformed the state-of-the-art results on the public dataset in terms of Area Under the ROC Curve (AUC) and Equal-Error-Rate (EER) by 14.4 % and 28.9%, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Noise-Agnostic Multitask Whisper Training for Reducing False Alarm Errors in Call-for-Help Detection

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Adding a noise classification head to Whisper during fine-tuning improved call-for-help detection accuracy from 65% to 88% on real-world recordings, though out-of-domain noise accuracy remained low.

Pith tools