Pith. sign in

REVIEW 60 cited by

Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.03209 v1 pith:7AD7MQAP submitted 2018-04-09 cs.CL cs.HC

classification cs.CLcs.HC
keywords datasetspeechdescribesrecognitiontaskaccuracyaudioautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional datasets used for automatic speech recognition of full sentences. Suggests a methodology for reproducible and comparable accuracy metrics for this task. Describes how the data was collected and verified, what it contains, previous versions and properties. Concludes by reporting baseline results of models trained on this dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 60 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mamba: Linear-Time Sequence Modeling with Selective State Spaces

    cs.LG 2023-12 unverdicted novelty 8.0 of 10

    Mamba is a linear-time sequence model using input-dependent selective SSMs that achieves SOTA results across modalities and matches twice-larger Transformers on language modeling with 5x higher inference throughput.

  2. Efficiently Modeling Long Sequences with Structured State Spaces

    cs.LG 2021-10 unverdicted novelty 8.0 of 10

    S4 is an efficient state space sequence model that captures long-range dependencies via structured parameterization of the SSM, achieving state-of-the-art results on the Long Range Arena and other benchmarks while bei...

  3. DiffWave: A Versatile Diffusion Model for Audio Synthesis

    eess.AS 2020-09 unverdicted novelty 8.0 of 10

    DiffWave is a non-autoregressive diffusion model that generates high-fidelity audio waveforms from noise in constant steps, matching WaveNet vocoder quality while being orders of magnitude faster and outperforming pri...

  4. LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    LongSpike integrates fractional-order state-space modeling into spiking neural networks, enabling better long-sequence performance than prior SNNs on LRA, WikiText-103, and Speech Commands benchmarks while retaining s...

  5. Attention by Synchronization in Coupled Oscillator Networks

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Kuramoto synchronization dynamics implement a provably unique and globally attractive attention mechanism that replaces softmax for physical substrates and shows competitive empirical performance.

  6. Totoro$^+$: An Adaptive and Scalable Edge Federated Learning System

    cs.DC 2026-05 unverdicted novelty 7.0 of 10

    Totoro+ is a DHT-based fully decentralized FL system with locality-aware multi-ring P2P structure, pub/sub forest, and game-theoretic path planning that claims O(log N) hops and 1.2-14x speedup for many concurrent app...

  7. FiTS: Interpretable Spiking Neurons via Frequency Selectivity and Temporal Shaping

    cs.NE 2026-05 unverdicted novelty 7.0 of 10

    FiTS spiking neurons improve auditory task performance over LIF baselines by factorizing computation into frequency selectivity and group-delay-based temporal shaping, yielding interpretable per-neuron parameters.

  8. Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Cumulative state updates in CMRU restore gradient flow through time in quantized bistable RNNs, yielding more stable convergence and competitive or superior performance versus LRUs and minGRUs on long-range sequence tasks.

  9. End-to-End Keyword Spotting on FPGA Using Graph Neural Networks with a Neuromorphic Auditory Sensor

    cs.LG 2026-05 conditional novelty 7.0 of 10

    An FPGA implementation of a neuromorphic auditory sensor plus graph neural network achieves 87.43% accuracy on Google Speech Commands v2 with sub-35 µs latency and 1.12 W power.

  10. MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    MMEB-V3 benchmark shows omni-modality embedding models fail to enforce instruction-specified modality constraints and exhibit asymmetric, query-biased retrieval.

  11. Quantum Dimension Reduction of Hidden Markov Models

    quant-ph 2026-01 conditional novelty 7.0 of 10

    Any finite ergodic HMM can be made deterministic by labelling transitions, yielding a normal iMPS that can be variationally compressed into a smaller quantum model.

  12. Covariance Estimation for Matrix-variate Data via Fixed-rank Core Covariance Geometry

    math.DG 2025-11 unverdicted novelty 7.0 of 10

    The space of rank-r core covariances forms a smooth manifold except on a measure-zero set, enabling a partial-isotropy shrinkage estimator for matrix-variate data.

  13. DASB - Discrete Audio and Speech Benchmark

    cs.SD 2024-06 unverdicted novelty 7.0 of 10

    DASB is a new benchmark for discrete audio tokens showing semantic tokens outperform acoustic ones but discrete representations remain less robust than continuous features across domains.

  14. Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications

    cs.CV 2024-05 unverdicted novelty 7.0 of 10

    Wake Vision pipeline produces a 6M-image person detection dataset for TinyML with 2.2% label error, improving model accuracy up to 6.6% over prior VWW benchmark across architectures and subsets.

  15. TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild

    cs.CL 2026-03 accept novelty 6.5 of 10

    A real-world elderly Taiwanese Hokkien intent dataset of 3k utterances is released, and two mining strategies from drama video are shown to suffer severe domain mismatch on it.

  16. NABEATs: Noise-Aware Audio Representation Learning

    eess.AS 2026-07 conditional novelty 6.0 of 10

    NABEATs estimates clean BEATs audio representations from noisy input using a reference noise signal, improving downstream tasks under seen and unseen noise.

  17. TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

    cs.LG 2026-07 accept novelty 6.0 of 10

    TabPFN as a zero-gradient head on frozen multimodal embeddings ranks best on NLL and ECE across 22 820 episodes while matching accuracy in mid-shot, mid-dimension regimes and also fixes miscalibration after fine-tuning.

  18. FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs

    cs.AR 2026-07 conditional novelty 6.0 of 10

    A full-stack LUT-as-neuron FPGA framework reports up to 205× lower latency than BNN accelerators and higher LUT efficiency than prior differentiable LUT networks at competitive binary accuracy.

  19. End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    An end-to-end SLU architecture with frozen SSL acoustic encoder, LSTM classification head, and cross-modal distillation achieves 93% accuracy on simple commands and 82% on spontaneous speech at 7 ms latency on the new...

  20. Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    A conditional generator operating in neural audio codec latent space produces targeted adversarial audio examples in one forward pass, reaching up to 99% success rate at sub-7 ms inference.

  21. Adaptive Speech-to-Spike Encoding for Spiking Neural Networks

    cs.NE 2026-06 unverdicted novelty 6.0 of 10

    A learnable residual speech-to-spike encoder jointly trained with an R-LIF SNN achieves up to 94.97% accuracy on GSC-v2 with a 35k-parameter model and supports DFA credit assignment at 91.5%.

  22. NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    NeuralMUSIC combines neural covariance estimation with the MUSIC pipeline, frequency attention fusion, and self-supervised learning to improve direction-of-arrival estimation for robotic sound source localization.

  23. Representation Matters in Randomized Smoothing for Audio Classification

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    Randomized smoothing in audio classification requires explicit specification of the certified representation and preprocessing because different choices produce different certified accuracies and effective perturbatio...

  24. What changes after deployment? A survey on On-device Learning in TinyML

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A survey of on-device learning in TinyML organized by distribution change regimes, highlighting influences on applications, hardware, and solutions plus a gap between benchmarks and deployments.

  25. Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Plug-in losses approximate EDL training objectives at the Dirichlet mean with decaying error as evidence grows, including softmax under a specific mapping, and match classical EDL performance on Google Speech Commands.

  26. AudioMosaic: Contrastive Masked Audio Representation Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    AudioMosaic learns general-purpose audio representations through contrastive pre-training with structured spectrogram masking, reaching state-of-the-art results on standard benchmarks and improving audio-language tasks.

  27. Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations

    cs.AR 2026-05 unverdicted novelty 6.0 of 10

    BMRUs enable a direct one-to-one mapping from learned parameters to current-mode analog circuit elements, with discrete hysteretic outputs suppressing noise by at least 20x and supporting sub-microwatt RNN inference i...

  28. EdgeSpike: Spiking Neural Networks for Low-Power Autonomous Sensing in Edge IoT Architectures

    cs.NE 2026-04 unverdicted novelty 6.0 of 10

    EdgeSpike delivers 91.4% mean accuracy on five sensing tasks with 31x lower energy on neuromorphic hardware and 6.3x longer battery life in a seven-month field deployment compared to conventional CNNs.

  29. Three factor delay learning rules for spiking neural networks

    cs.NE 2026-01 conditional novelty 6.0 of 10

    A three-factor eligibility-trace rule lets LIF spiking networks learn synaptic and axonal delays online, matching offline backpropagation accuracy on SHD while cutting model size.

  30. ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples

    cs.CR 2025-12 unverdicted novelty 6.0 of 10

    ComMark embeds covert watermarks in models using frequency-domain compressed samples and simulated attacks, claiming state-of-the-art covertness and robustness across image, speech, text, and video tasks.

  31. AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers

    cs.SD 2025-12 unverdicted novelty 6.0 of 10

    AaSP learns aliasing-stable audio representations by augmenting patch tokens with adaptive subband features from alias-prone bands and using teacher-student masked modeling plus multi-mask contrastive regularization, ...

  32. DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

    cs.SD 2025-11 conditional novelty 6.0 of 10

    DHAuDS is a new audio benchmark that corrupts four existing datasets with dynamically varying and diverse acoustic noise, and evaluates three classifiers under test-time adaptation.

  33. Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A binarized prototypical probing method that pools per-class evidence from patch tokens substantially outperforms [cls]-token and attentive probes on multi-label audio classification benchmarks.

  34. MambaLite-Micro: Memory-Optimized Mamba Inference on MCUs

    cs.LG 2025-09 conditional novelty 6.0 of 10

    MambaLite-Micro deploys Mamba inference in C on ESP32S3 and STM32H7 with 83.0% lower peak memory than an unfused baseline and identical output labels to PyTorch.

  35. AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition

    cs.SD 2025-09 conditional novelty 6.0 of 10

    AudioRWKV, an RWKV7-based audio backbone with 2D convolution and bidirectional WKV, beats AST and AuM baselines on five benchmarks at linear complexity.

  36. Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Replacing square spectrogram patches with full-frequency temporal patches plus patch-aligned masking improves audio classification accuracy and reduces compute for Transformer and Mamba models.

  37. Label Smoothing++: Enhanced Label Regularization for Training Neural Networks

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Label Smoothing++ learns a class-wise non-target probability distribution to replace the uniform distribution in label smoothing, yielding accuracy gains across many benchmarks.

  38. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  39. ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ViRN combines variational inference with Wasserstein-distance neighbor fusion to estimate tail-class distributions in continual learning, reporting gains on six long-tailed benchmarks.

  40. Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...

  41. SiLIF: Structured State Space Model Dynamics and Parametrization for Spiking Neural Networks

    cs.NE 2025-06 unverdicted novelty 6.0 of 10

    SiLIF models apply SSM dynamics and parametrization to spiking neurons for stable training, reaching new SOTA on event-based and raw-audio speech datasets while using half the compute of SSMs via synaptic delays.

  42. Simplified State Space Layers for Sequence Modeling

    cs.LG 2022-08 accept novelty 6.0 of 10

    S5 uses a single MIMO state space model with S4-derived initialization to match S4 efficiency and reach 87.4% average accuracy on the Long Range Arena benchmark.

  43. Sub-band Convolutional Neural Networks for Small-footprint Spoken Term Classification

    eess.AS 2019-07 unverdicted novelty 6.0 of 10

    Sub-band CNN applies distinct kernels per frequency sub-band to reduce computation 39.7-49.3% versus full-band CNN on Speech Commands dataset while maintaining accuracy.

  44. Federated Learning with Non-IID Data

    cs.LG 2018-06 conditional novelty 6.0 of 10

    Non-IID data causes up to 55% accuracy loss in federated learning due to weight divergence measured by earth mover's distance; 5% globally shared data recovers 30% accuracy on CIFAR-10.

  45. The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

    cs.ET 2026-07 accept novelty 5.5 of 10

    SpiNNaker2 delivers a measured many-core platform combining ARM cores, ML accelerators, and event routing that runs SNNs, DNNs, and hybrid event-based models on one scalable chip.

  46. NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

    cs.SD 2026-06 conditional novelty 5.5 of 10

    A neural covariance estimator plus MUSIC, frequency attention fusion, and self-supervised channel masking yields lower angular error and better cross-domain DOA estimates than classical and deep baselines on robot aud...

  47. From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Task-specific digit-recognition attacks show that temporal smoothing, resampling, and shredding protect spoken digits less—and differently—than read-speech word-error rates suggest.

  48. Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting

    cs.SD 2026-07 conditional novelty 5.0 of 10

    Finite-window unitary phase-transport layers whose recurrences collapse to cumulative sums match compact CNN baselines on Speech Commands and run with lower latency than a custom scan.

  49. Scalable Keyword Spotting via Modular Network Expansion

    cs.SD 2026-07 conditional novelty 5.0 of 10

    Freezing a deployed keyword-spotting model and training only a small attached branch with a separate head adds new keywords with no change to old outputs, cutting new-keyword false reject rate from 6.46% to 4.37% vers...

  50. Natural Backdoor Attacks on Speech Recognition Models

    cs.CR 2026-07 conditional novelty 5.0 of 10

    Backdoors in speech classifiers can be planted with natural ambient sounds as triggers, achieving high attack success at 5% poisoning while preserving benign accuracy.

  51. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

    cs.SD 2026-07 reject novelty 5.0 of 10

    SpeechGuard combines an SNR-adapted STRIP detector with an autoencoder that learns time-frequency masks to suppress backdoor triggers in speech, but purifier training needs oracle poisoned/clean pairs.

  52. TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

    cs.LG 2026-07 conditional novelty 5.0 of 10

    TabPFN as a training-free head on PCA-reduced frozen multimodal embeddings broadly improves calibration (NLL, ECE) over classical heads, with an accuracy edge only for k≥50 shots and d≤32 features.

  53. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  54. Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

    cs.CR 2026-07 unverdicted novelty 5.0 of 10

    Pmeta-TLA combines a frame-level timbre leakage trigger with meta-learning and PCGrad to inject multiple backdoors into speech models in one training run, claiming better attack success, stealth, and lower cost than b...

  55. Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

    eess.AS 2026-06 unverdicted novelty 5.0 of 10

    ZP-KWS combines a phoneme-supervised audio encoder with a 0.9M-parameter GE2E speaker encoder and multiplicative late fusion to cut target-only FRR at 1% FAR by up to 60% on LibriPhrase, Google Speech Commands, and Qu...

  56. Learning task-specific subspaces via interventional post-training of speech foundation models

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Interventional contrastive learning applied after pre-training disentangles speaker and content information in speech foundation model embeddings, yielding better out-of-domain speaker verification.

  57. Representation Matters in Randomized Smoothing for Audio Classification

    eess.AS 2026-06 conditional novelty 5.0 of 10

    Randomized smoothing robustness certification in audio is under-specified without explicit choice of certified object, preprocessing policy, and perturbation model, as different representations yield different certifi...

  58. Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

    eess.AS 2026-05 unverdicted novelty 5.0 of 10

    A continuous-token model with shared Haar wavelet coefficients reports 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR on three datasets and shows energy-based selection outperforms uniform selection by roughly 16 dB.

  59. Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation

    eess.AS 2026-05 unverdicted novelty 5.0 of 10

    DMA-KWS achieves 97.85% AUC and 6.13% EER on LibriPhrase Hard via dual-stage CTC/QbyT matching, multi-modal enrollment, and lightweight continual adaptation with 187k parameters.

  60. Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations

    cs.AR 2026-05 unverdicted novelty 5.0 of 10

    BMRUs enable analog recurrent neural network hardware via discrete outputs that suppress noise 20-fold, with one-to-one parameter-to-circuit mapping and linear power scaling for recurrence.

Pith tools