REVIEW 15 cited by
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition. On the largest public lip-reading benchmark LRS3 (433 hours), AV-HuBERT achieves 32.5% WER with only 30 hours of labeled data, outperforming the former state-of-the-art approach (33.6%) trained with a thousand times more transcribed video data (31K hours). The lip-reading WER is further reduced to 26.9% when using all 433 hours of labeled data from LRS3 and combined with self-training. Using our audio-visual representation on the same benchmark for audio-only speech recognition leads to a 40% relative WER reduction over the state-of-the-art performance (1.3% vs 2.3%). Our code and models are available at https://github.com/facebookresearch/av_hubert
Forward citations
Cited by 15 Pith papers
-
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis
HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only...
-
Your Multimodal Speech Model Says I Have a Face for Radio
Multimodal speech models show word error rate differences of up to 4.05 points when the same audio is paired with faces differing in self-declared gender and ethnicity.
-
Identity-Faithful Audio-Visual Target Speaker Extraction with QIANGDA and VOXBLINK2-AVSE
A real-scene Mandarin benchmark and a 766-hour curated audio-lip corpus for identity-faithful audio-visual target speaker extraction, with a baseline scoring 0.2261 CER and 82.22% strict identity correctness.
-
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.
-
Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning
TAVR-IVD uses multi-object tracking to create per-vehicle tracklets for audio-visual idling detection, claiming higher SNR, temporal stability, explicit spatial alignment, and better domain adaptation than full-frame fusion.
-
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Augmenting diffusion-based visual-conditioned speech enhancement with a contrastive audio-visual loss produces consistent gains in interference suppression and perceptual quality, especially at low SNRs.
-
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
Masked multimodal training on sEMG and lipreading reduces word error rate by up to 14 percentage points and improves robustness to modality loss in silent speech synthesis.
-
Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization
IaMSB applies a Schrödinger Bridge in two stages to estimate cross-modal consistency and localize deepfake intervals, reporting 3-10% gains in AP@0.95 especially on single-sided forgeries.
-
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
A training-free dual-system framework refines anomaly score ordering on uncertain samples from self-supervised talking head forgery detectors to improve detection performance.
-
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...
-
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
Four-class audio-visual DeepFake detectors misclassify authentic but semantically mismatched audio-video pairs; five-class training plus ImageBind similarity improves models that can learn cross-modal semantics.
-
HumanOmni-Speaker: Identifying Who said What and When
HumanOmni-Speaker introduces a Visual Delta Encoder and VR-SDR benchmark that enable end-to-end speaker diarization and recognition by sampling video at 25 fps and compressing inter-frame motion residuals into 6 token...
-
Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring
DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography
Splitting a steganographic payload across audio and video tracks reduces detection rates for both single-mode and multimodal detectors, though multimodal performance gains appear driven mostly by the video stream alone.
Discussion (0). Sign in to comment.