Pith. sign in

REVIEW 1 cited by

To Wake-up or Not to Wake-up: Reducing Keyword False Alarm by Successive Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03416 v1 pith:53XWJFTH submitted 2023-04-06 eess.SP cs.LGcs.SDeess.AS

classification eess.SPcs.LGcs.SDeess.AS
keywords keywordrefinementspottingsuccessivealarmaudioclassifiesdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Keyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hello Afrika: Speech Commands in Kinyarwanda

    eess.AS 2025-06 conditional novelty 4.0 of 10

    Hello Afrika compiles a Kinyarwanda speech command dataset and trains LSTM classifiers, reaching 78.1% validation accuracy on MSWC data but only 36.8% on local recordings.

Pith tools