Pith. sign in

REVIEW 1 cited by

Speech Recognition: Keyword Spotting Through Image Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.03759 v2 pith:5NLKSOGK submitted 2018-03-10 stat.ML cs.LG

classification stat.MLcs.LG
keywords recognitionproblemspeechaudiodomainimageneuralparticular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The problem of identifying voice commands has always been a challenge due to the presence of noise and variability in speed, pitch, etc. We will compare the efficacies of several neural network architectures for the speech recognition problem. In particular, we will build a model to determine whether a one second audio clip contains a particular word (out of a set of 10), an unknown word, or silence. The models to be implemented are a CNN recommended by the Tensorflow Speech Recognition tutorial, a low-latency CNN, and an adversarially trained CNN. The result is a demonstration of how to convert a problem in audio recognition to the better-studied domain of image classification, where the powerful techniques of convolutional neural networks are fully developed. Additionally, we demonstrate the applicability of the technique of Virtual Adversarial Training (VAT) to this problem domain, functioning as a powerful regularizer with promising potential future applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Words: Interjection Classification for Improved Human-Computer Interaction

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Interjection classification over five speakers improves when pitch, tempo, and background-noise augmentation is added, but absolute accuracy stays below 60% on unseen speakers.

Pith tools