Pith. sign in

REVIEW 15 cited by

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.02184 v2 pith:PYGQY5BB submitted 2022-01-05 eess.AS cs.CVcs.SD

classification eess.AScs.CVcs.SD
keywords speechaudio-visualrepresentationhoursav-hubertdatalearninglip-reading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition. On the largest public lip-reading benchmark LRS3 (433 hours), AV-HuBERT achieves 32.5% WER with only 30 hours of labeled data, outperforming the former state-of-the-art approach (33.6%) trained with a thousand times more transcribed video data (31K hours). The lip-reading WER is further reduced to 26.9% when using all 433 hours of labeled data from LRS3 and combined with self-training. Using our audio-visual representation on the same benchmark for audio-only speech recognition leads to a 40% relative WER reduction over the state-of-the-art performance (1.3% vs 2.3%). Our code and models are available at https://github.com/facebookresearch/av_hubert

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

    eess.AS 2026-06 unverdicted novelty 7.0 of 10

    HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only...

  2. Your Multimodal Speech Model Says I Have a Face for Radio

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Multimodal speech models show word error rate differences of up to 4.05 points when the same audio is paired with faces differing in self-declared gender and ethnicity.

  3. Identity-Faithful Audio-Visual Target Speaker Extraction with QIANGDA and VOXBLINK2-AVSE

    eess.AS 2026-08 conditional novelty 6.0 of 10

    A real-scene Mandarin benchmark and a 766-hour curated audio-lip corpus for identity-faithful audio-visual target speaker extraction, with a baseline scoring 0.2261 CER and 82.22% strict identity correctness.

  4. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  5. Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    TAVR-IVD uses multi-object tracking to create per-vehicle tracklets for audio-visual idling detection, claiming higher SNR, temporal stability, explicit spatial alignment, and better domain adaptation than full-frame fusion.

  6. Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

    eess.SP 2026-06 unverdicted novelty 6.0 of 10

    Augmenting diffusion-based visual-conditioned speech enhancement with a contrastive audio-visual loss produces consistent gains in interference suppression and perceptual quality, especially at low SNRs.

  7. Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    Masked multimodal training on sEMG and lipreading reduces word error rate by up to 14 percentage points and improves robustness to modality loss in silent speech synthesis.

  8. Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    IaMSB applies a Schrödinger Bridge in two stages to estimate cross-modal consistency and localize deepfake intervals, reporting 3-10% gains in AP@0.95 especially on single-sided forgeries.

  9. Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    A training-free dual-system framework refines anomaly score ordering on uncertain samples from self-supervised talking head forgery detectors to improve detection performance.

  10. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...

  11. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 conditional novelty 6.0 of 10

    Four-class audio-visual DeepFake detectors misclassify authentic but semantically mismatched audio-video pairs; five-class training plus ImageBind similarity improves models that can learn cross-modal semantics.

  12. HumanOmni-Speaker: Identifying Who said What and When

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    HumanOmni-Speaker introduces a Visual Delta Encoder and VR-SDR benchmark that enable end-to-end speaker diarization and recognition by sampling video at 25 fps and compressing inter-frame motion residuals into 6 token...

  13. Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

    cs.NI 2025-08 unverdicted novelty 6.0 of 10

    DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.

  14. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  15. Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography

    cs.CR 2026-06 unverdicted novelty 3.0 of 10

    Splitting a steganographic payload across audio and video tracks reduces detection rates for both single-mode and multimodal detectors, though multimodal performance gains appear driven mostly by the video stream alone.

Pith tools