Pith. sign in

REVIEW 1 cited by

Real-Time Target Sound Extraction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.02250 v3 pith:2MRXV47R submitted 2022-11-04 cs.SD cs.LGeess.AS

Real-Time Target Sound Extraction

classification cs.SD cs.LGeess.AS
keywords architecturecausaldecoderdilatedextractionmodelreal-timesound
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present the first neural network model to achieve real-time and streaming target sound extraction. To accomplish this, we propose Waveformer, an encoder-decoder architecture with a stack of dilated causal convolution layers as the encoder, and a transformer decoder layer as the decoder. This hybrid architecture uses dilated causal convolutions for processing large receptive fields in a computationally efficient manner while also leveraging the generalization performance of transformer-based architectures. Our evaluations show as much as 2.2-3.3 dB improvement in SI-SNRi compared to the prior models for this task while having a 1.2-4x smaller model size and a 1.5-2x lower runtime. We provide code, dataset, and audio samples: https://waveformer.cs.washington.edu/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Low-latency Assistive Audio Enhancement for Neurodivergent People

    eess.AS 2025-09 conditional novelty 4.0

    Among DSP and ML audio enhancement approaches evaluated on trigger-sound mixtures, Dynamic Range Compression (DRC) attenuates distressing sounds most effectively in both objective metrics and a neurodivergent listening test.