REVIEW 3 cited by
ConvConcatNet: a deep convolutional neural network to reconstruct mel spectrogram from the EEG
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To investigate the processing of speech in the brain, simple linear models are commonly used to establish a relationship between brain signals and speech features. However, these linear models are ill-equipped to model a highly dynamic and complex non-linear system like the brain. Although non-linear methods with neural networks have been developed recently, reconstructing unseen stimuli from unseen subjects' EEG is still a highly challenging task. This work presents a novel method, ConvConcatNet, to reconstruct mel-specgrams from EEG, in which the deep convolution neural network and extensive concatenation operation were combined. With our ConvConcatNet model, the Pearson correlation between the reconstructed and the target mel-spectrogram can achieve 0.0420, which was ranked as No.1 in the Task 2 of the Auditory EEG Challenge. The codes and models to implement our work will be available on Github: https://github.com/xuxiran/ConvConcatNet
Forward citations
Cited by 3 Pith papers
-
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
EEG decoding of Mandarin pitch is more accurate for speaker-normalized than raw pitch across multiple speakers, suggesting the brain encodes relative, speaker-independent pitch.
-
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
A multi-task EEG decoder with an auxiliary phoneme predictor decodes listened speech waveforms and phoneme sequences in parallel and reports improvements over prior single-task models.
-
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
SSM2Mel reconstructs continuous speech mel spectrograms from EEG with a modest Pearson correlation gain (0.069 vs 0.063) over prior state of the art.
Discussion (0). Continue with ORCID to comment.