Pith. sign in

REVIEW 2 cited by

NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.14303 v1 pith:CYF3VXBO submitted 2023-07-26 eess.AS

classification eess.AS
keywords neuroheedspeechsignalsspeakerattendedauditoryextractiononline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities, which the latter can be measured using affordable and non-intrusive electroencephalography (EEG) devices. In this study, we present NeuroHeed, a speaker extraction model that leverages EEG signals to establish a neuronal attractor which is temporally associated with the speech stimulus, facilitating the extraction of the attended speech signal in a cocktail party scenario. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results demonstrate that NeuroHeed effectively extracts brain-attended speech signals, achieving high signal quality, excellent perceptual quality, and intelligibility in a two-speaker scenario.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Audio-Visual Autoregressive Speaker Extraction

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A 0.1M-parameter visual frontend plus an autoregressive acoustic encoder improves online audio-visual speaker extraction by up to 0.9 dB SI-SNRi on simulated LRS3 mixtures, with roughly 0.5 dB attributable to the acou...

  2. ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

    cs.SD 2025-06 conditional novelty 4.0 of 10

    A unified open-source toolkit with pretrained models for speech enhancement, separation, super-resolution, and multimodal target extraction, plus a new face-conditioned extraction model.

Pith tools