Whisper-Tiny used as the audio feature extractor for RAD-NeRF and ER-NeRF talking portraits reduces AFE latency and yields modestly better SyncNet lip-sync scores than DeepSpeech, Wav2Vec 2.0, or HuBERT on three short datasets.
Designing effective training programs for investigative interviewers of children
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
Whisper-Tiny used as the audio feature extractor for RAD-NeRF and ER-NeRF talking portraits reduces AFE latency and yields modestly better SyncNet lip-sync scores than DeepSpeech, Wav2Vec 2.0, or HuBERT on three short datasets.