Intermediate wav2vec 2.0 layers provide the best self-supervised speech representations for learning articulatory gestures and imitating speech across speakers.
u rbis, Simon Stone, Patrick H \
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
Intermediate wav2vec 2.0 layers provide the best self-supervised speech representations for learning articulatory gestures and imitating speech across speakers.