A ResNet-Conformer network with shared weights and multi-scale attention, trained on synthetic and augmented real audio, improves detection and direction estimation over the DCASE 2024 baseline while distance error stays flat.
STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
A ResNet-Conformer network with shared weights and multi-scale attention, trained on synthetic and augmented real audio, improves detection and direction estimation over the DCASE 2024 baseline while distance error stays flat.