Pith. sign in

REVIEW

Disentangled Training with Adversarial Examples For Robust Small-footprint Keyword Spotting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.13355 v1 pith:FDYW63LK submitted 2024-08-23 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords adversarialexamplesdatasetdisentangledfalsekeywordlearningmismatch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

A keyword spotting (KWS) engine that is continuously running on device is exposed to various speech signals that are usually unseen before. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we explore how to effectively apply adversarial examples to improve KWS robustness. We propose datasource-aware disentangled learning with adversarial examples to reduce the mismatch between the original and adversarial data as well as the mismatch across original training datasources. The KWS model architecture is based on depth-wise separable convolution and a simple attention module. Experimental results demonstrate that the proposed learning strategy improves false reject rate by $40.31%$ at $1%$ false accept rate on the internal dataset, compared to the strongest baseline without using adversarial examples. Our best-performing system achieves $98.06%$ accuracy on the Google Speech Commands V1 dataset.

Discussion (0). Sign in to comment.

Pith tools