Pith. sign in

REVIEW 1 cited by

Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18239 v2 pith:7IXOIMH7 submitted 2024-09-26 cs.SD cs.LGeess.AS

Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables

classification cs.SD cs.LGeess.AS
keywords latencyhearablesenhancementmeanspeechmodelsreal-timesub-millisecond
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FIR filter, enabling sample-by-sample processing to achieve mean algorithmic latency of 0.32 ms to 1.25 ms. With a single microphone, we observe a mean SI-SDRi of 4.1 dB. The approach shows generalization with a DNSMOS increase of 0.2 on unseen audio recordings. We use a lightweight LSTM-based model of 626k parameters to generate FIR taps. Using a real hardware implementation on a low-power DSP, our system can run with 376 MIPS and a mean end-to-end latency of 3.35 ms. In addition, we provide a comparison with existing low-latency spectral masking techniques. We hope this work will enable a better understanding of latency and can be used to improve the comfort and usability of hearables.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Low-latency Assistive Audio Enhancement for Neurodivergent People

    eess.AS 2025-09 conditional novelty 4.0

    Among DSP and ML audio enhancement approaches evaluated on trigger-sound mixtures, Dynamic Range Compression (DRC) attenuates distressing sounds most effectively in both objective metrics and a neurodivergent listening test.