Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.