A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy and a new WordNet-based semantic score.
InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The local and global features are both essential for automatic speech recognition (ASR). Many recent methods have verified that simply combining local and global features can further promote ASR performance. However, these methods pay less attention to the interaction of local and global features, and their series architectures are rigid to reflect local and global relationships. To address these issues, this paper proposes InterFormer for interactive local and global features fusion to learn a better representation for ASR. Specifically, we combine the convolution block with the transformer block in a parallel design. Besides, we propose a bidirectional feature interaction module (BFIM) and a selective fusion module (SFM) to implement the interaction and fusion of local and global features, respectively. Extensive experiments on public ASR datasets demonstrate the effectiveness of our proposed InterFormer and its superior performance over the other Transformer and Conformer models.
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy and a new WordNet-based semantic score.