Pith. sign in

REVIEW 3 cited by

An efficient encoder-decoder architecture with top-down attention for speech separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15200 v5 pith:2VHVNAZT submitted 2022-09-30 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords attentiontdanettop-downonlysepformerfeaturesinputlayers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks have shown excellent prospects in speech separation tasks. However, obtaining good results while keeping a low model complexity remains challenging in real-world applications. In this paper, we provide a bio-inspired efficient encoder-decoder architecture by mimicking the brain's top-down attention, called TDANet, with decreased model complexity without sacrificing performance. The top-down attention in TDANet is extracted by the global attention (GA) module and the cascaded local attention (LA) layers. The GA module takes multi-scale acoustic features as input to extract global attention signal, which then modulates features of different scales by direct top-down connections. The LA layers use features of adjacent layers as input to extract the local attention signal, which is used to modulate the lateral input in a top-down manner. On three benchmark datasets, TDANet consistently achieved competitive separation performance to previous state-of-the-art (SOTA) methods with higher efficiency. Specifically, TDANet's multiply-accumulate operations (MACs) are only 5\% of Sepformer, one of the previous SOTA models, and CPU inference time is only 10\% of Sepformer. In addition, a large-size version of TDANet obtained SOTA results on three datasets, with MACs still only 10\% of Sepformer and the CPU inference time only 24\% of Sepformer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Independent hard prompt compression frequently retains an answer while deleting the definition or bridge needed to interpret it, and fixed-budget reinsertion of that missing support substantially improves QA accuracy.

  2. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  3. EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis

    eess.IV 2025-06 reject novelty 4.0 of 10

    EAGLE combines Mamba-style state-space blocks with wavelet downsampling and reports 89.76% DSC on private hepatic echinococcosis CT data, but slice-level splitting and missing artifacts undermine the claim.

Pith tools