Pith. sign in

REVIEW 4 cited by

Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.13975 v3 pith:GIAIQMLY submitted 2020-07-28 eess.AS cs.SD

classification eess.AScs.SD
keywords speechnetworksequencestransformerseparationdirectmodelmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent neural network, leading to suboptimal separation performance. In this paper, we propose a dual-path transformer network (DPTNet) for end-to-end speech separation, which introduces direct context-awareness in the modeling for speech sequences. By introduces a improved transformer, elements in speech sequences can interact directly, which enables DPTNet can model for the speech sequences with direct context-awareness. The improved transformer in our approach learns the order information of the speech sequences without positional encodings by incorporating a recurrent neural network into the original transformer. In addition, the structure of dual paths makes our model efficient for extremely long speech sequence modeling. Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (20.6 dB SDR on the public WSj0-2mix data corpus).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  2. Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A dynamic frequency-adaptive knowledge distillation method, using the steepest point in the running maximum of the teacher spectrum as a crossover, improves speech enhancement student models by small PESQ margins over...

  3. DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions

    cs.SD 2025-02 conditional novelty 4.0 of 10

    DCF-Net, a time-frequency target speaker extraction model with a DualStream Fusion Block, reports 21.6 dB SI-SDRi on WSJ0-2Mix and a 0.4% target confusion rate.

  4. EDSep: An Effective Diffusion-Based Method for Speech Source Separation

    eess.AS 2025-01 reject novelty 4.0 of 10

    EDSep, a modified score-matching SDE method with a new denoiser and stochastic sampler, reports higher SI-SDR than DiffSep and mixed results against Conv-TasNet.

Pith tools