Pith. sign in

REVIEW 2 cited by

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.00466 v1 pith:CDY25KRU submitted 2025-05-31 eess.AS cs.SD

classification eess.AScs.SD
keywords speechtemporalalignmentbrain-assistedextractinformationm3anetmodalities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction

    cs.SD 2025-07 reject novelty 5.0 of 10

    DMF2Mel, a dynamic multiscale fusion network, reports the best mel spectrogram reconstruction scores on SparrKULee, though test-set hyperparameter tuning makes the comparison unreliable.

  2. Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss

    eess.AS 2026-01 conditional novelty 4.0 of 10

    A contrastive Pearson-correlation loss—maximizing attended minus average unattended envelope correlation—improves EEG auditory attention decoding accuracy in most, but not all, tested model/dataset settings.

Pith tools