REVIEW 5 cited by
Moonshine: Speech Recognition for Live Transcription and Voice Commands
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Moonshine: Speech Recognition for Live Transcription and Voice Commands
read the original abstract
This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary Position Embedding (RoPE) instead of traditional absolute position embeddings. The model is trained on speech segments of various lengths, but without using zero-padding, leading to greater efficiency for the encoder during inference time. When benchmarked against OpenAI's Whisper tiny-en, Moonshine Tiny demonstrates a 5x reduction in compute requirements for transcribing a 10-second speech segment while incurring no increase in word error rates across standard evaluation datasets. These results highlight Moonshine's potential for real-time and resource-constrained applications.
Forward citations
Cited by 5 Pith papers
-
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
Asynchronous I/O and Speculative Tool Calling cut latency in tool-calling LLM agents by 1.3-2.2x with only minor accuracy loss on cloud and edge models.
-
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
Speculative Interaction Agents achieve 1.3-2.2x speedups for real-time tool-calling agents via async I/O decoupling and speculative calls, with clock-based training for small edge models.
-
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
FADE adaptively compensates for quantization errors layer-by-layer in ASR models using diagnostic scores from weight geometry and calibration data, yielding lower word error rates at 3- and 4-bit precision.
-
VibeVoice-ASR-BitNet Technical Report
Heterogeneous quantization (INT8 tokenizer + 2-bit ternary LM) makes a 1.5B-parameter LLM-based ASR system run at real-time speed on CPUs with 2.9x compression and modest measured WER increases.
-
BranchShine: Compact Raw-Audio-to-IPA Transcription with a RoPE E-Branchformer Encoder
A 33M-parameter raw-audio CTC model with 19-block RoPE E-Branchformer achieves 9.19% whitespace-insensitive IPA CER on a 16,660-utterance 41-language test set, outperforming a 575M-parameter PhoneticXEUS baseline at 9...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.