CNVSRC 2024 lowers the baseline character error rate for Chinese visual speech recognition from 48.6% to 39.7% (single-speaker) and from 58.4% to 52.2% (multi-speaker), while adding a 200-hour dataset and documenting winning methods.
Only the datasets provided by the organizer (except CN-CVS2-P1) were used, meaning these systems conform to the specifications of the fixed tracks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
CNVSRC 2024 lowers the baseline character error rate for Chinese visual speech recognition from 48.6% to 39.7% (single-speaker) and from 58.4% to 52.2% (multi-speaker), while adding a 200-hour dataset and documenting winning methods.