Pith. sign in

REVIEW 1 cited by

Conformer-Based Speech Recognition On Extreme Edge-Computing Devices

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.10359 v3 pith:PMK4GCT5 submitted 2023-12-16 cs.LG cs.PF

classification cs.LGcs.PF
keywords devicesrecognitionsmartspeechaccuracyotherresource-constrainedwearables
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With increasingly more powerful compute capabilities and resources in today's devices, traditionally compute-intensive automatic speech recognition (ASR) has been moving from the cloud to devices to better protect user privacy. However, it is still challenging to implement on-device ASR on resource-constrained devices, such as smartphones, smart wearables, and other smart home automation devices. In this paper, we propose a series of model architecture adaptions, neural network graph transformations, and numerical optimizations to fit an advanced Conformer based end-to-end streaming ASR system on resource-constrained devices without accuracy degradation. We achieve over 5.26 times faster than realtime (0.19 RTF) speech recognition on smart wearables while minimizing energy consumption and achieving state-of-the-art accuracy. The proposed methods are widely applicable to other transformer-based server-free AI applications. In addition, we provide a complete theory on optimal pre-normalizers that numerically stabilize layer normalization in any Lp-norm using any floating point precision.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WhisperKit: On-device Real-time ASR with Billion-Scale Transformers

    cs.SD 2025-07 conditional novelty 5.0 of 10

    WhisperKit's optimized on-device Whisper Large v3 Turbo streaming system reportedly achieves 0.46 s per-word latency and 2.2% WER, beating cloud baselines in its benchmark.

Pith tools