Pith. sign in

REVIEW 1 cited by

Visual Speech Recognition for Multiple Languages in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.13084 v2 pith:FQAU5IA4 submitted 2022-02-26 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords datadatasetslanguagesmodelmodelsspeechtrainingadvances
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream. Advances in deep learning and the availability of large audio-visual datasets have led to the development of much more accurate and robust VSR models than ever before. However, these advances are usually due to the larger training sets rather than the model design. Here we demonstrate that designing better models is equally as important as using larger training sets. We propose the addition of prediction-based auxiliary tasks to a VSR model, and highlight the importance of hyperparameter optimization and appropriate data augmentations. We show that such a model works for different languages and outperforms all previous methods trained on publicly available datasets by a large margin. It even outperforms models that were trained on non-publicly available datasets containing up to to 21 times more data. We show, furthermore, that using additional training data, even in other languages or with automatically generated transcriptions, results in further improvement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concatenating speech-recognition and active-speaker visual embeddings improves audio-visual speech enhancement in low-SNR multi-speaker settings, and the CPU real-time system is released open source.

Pith tools