A lightweight recurrent network trained on 10-second segments can directly process 21 to 121 second mixtures and keep each speaker's utterances in a consistent output stream across silences up to about 40 seconds.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multi-Utterance Speech Separation and Association Trained on Short Segments
A lightweight recurrent network trained on 10-second segments can directly process 21 to 121 second mixtures and keep each speaker's utterances in a consistent output stream across silences up to about 40 seconds.