Pith. sign in

REVIEW 3 cited by

ConvConcatNet: a deep convolutional neural network to reconstruct mel spectrogram from the EEG

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04965 v1 pith:TG4SP3VT submitted 2024-01-10 eess.SP cs.LG

classification eess.SPcs.LG
keywords convconcatnetbrainmodelsneuraldeepgithubhighlylinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To investigate the processing of speech in the brain, simple linear models are commonly used to establish a relationship between brain signals and speech features. However, these linear models are ill-equipped to model a highly dynamic and complex non-linear system like the brain. Although non-linear methods with neural networks have been developed recently, reconstructing unseen stimuli from unseen subjects' EEG is still a highly challenging task. This work presents a novel method, ConvConcatNet, to reconstruct mel-specgrams from EEG, in which the deep convolution neural network and extensive concatenation operation were combined. With our ConvConcatNet model, the Pearson correlation between the reconstructed and the target mel-spectrogram can achieve 0.0420, which was ranked as No.1 in the Task 2 of the Auditory EEG Challenge. The codes and models to implement our work will be available on Github: https://github.com/xuxiran/ConvConcatNet

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception

    cs.SD 2025-05 conditional novelty 6.0 of 10

    EEG decoding of Mandarin pitch is more accurate for speaker-normalized than raw pitch across multiple speakers, suggesting the brain encodes relative, speaker-independent pitch.

  2. Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A multi-task EEG decoder with an auxiliary phoneme predictor decodes listened speech waveforms and phoneme sequences in parallel and reports improvements over prior single-task models.

  3. SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

    eess.SP 2025-01 conditional novelty 4.0 of 10

    SSM2Mel reconstructs continuous speech mel spectrograms from EEG with a modest Pearson correlation gain (0.069 vs 0.063) over prior state of the art.

Pith tools