HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only methods.
Learning audio-visual speech representa- tion by masked multimodal cluster prediction
10 Pith papers cite this work. Polarity classification is still indexing.
years
2026 10representative citing papers
Multimodal speech models show word error rate differences of up to 4.05 points when the same audio is paired with faces differing in self-declared gender and ethnicity.
TAVR-IVD uses multi-object tracking to create per-vehicle tracklets for audio-visual idling detection, claiming higher SNR, temporal stability, explicit spatial alignment, and better domain adaptation than full-frame fusion.
Augmenting diffusion-based visual-conditioned speech enhancement with a contrastive audio-visual loss produces consistent gains in interference suppression and perceptual quality, especially at low SNRs.
Masked multimodal training on sEMG and lipreading reduces word error rate by up to 14 percentage points and improves robustness to modality loss in silent speech synthesis.
IaMSB applies a Schrödinger Bridge in two stages to estimate cross-modal consistency and localize deepfake intervals, reporting 3-10% gains in AP@0.95 especially on single-sided forgeries.
A training-free dual-system framework refines anomaly score ordering on uncertain samples from self-supervised talking head forgery detectors to improve detection performance.
Four-class audio-visual DeepFake detectors misclassify authentic but semantically mismatched audio-video pairs; five-class training plus ImageBind similarity improves models that can learn cross-modal semantics.
HumanOmni-Speaker introduces a Visual Delta Encoder and VR-SDR benchmark that enable end-to-end speaker diarization and recognition by sampling video at 25 fps and compressing inter-frame motion residuals into 6 tokens per frame.
Splitting a steganographic payload across audio and video tracks reduces detection rates for both single-mode and multimodal detectors, though multimodal performance gains appear driven mostly by the video stream alone.
citing papers explorer
-
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis
HoliDubber introduces a patch-based autoregressive diffusion transformer for joint text-guided synthesis of speech and ambient audio in video dubbing, with a new benchmark showing outperformance over prior speech-only methods.
-
Your Multimodal Speech Model Says I Have a Face for Radio
Multimodal speech models show word error rate differences of up to 4.05 points when the same audio is paired with faces differing in self-declared gender and ethnicity.
-
Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning
TAVR-IVD uses multi-object tracking to create per-vehicle tracklets for audio-visual idling detection, claiming higher SNR, temporal stability, explicit spatial alignment, and better domain adaptation than full-frame fusion.
-
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Augmenting diffusion-based visual-conditioned speech enhancement with a contrastive audio-visual loss produces consistent gains in interference suppression and perceptual quality, especially at low SNRs.
-
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
Masked multimodal training on sEMG and lipreading reduces word error rate by up to 14 percentage points and improves robustness to modality loss in silent speech synthesis.
-
Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization
IaMSB applies a Schrödinger Bridge in two stages to estimate cross-modal consistency and localize deepfake intervals, reporting 3-10% gains in AP@0.95 especially on single-sided forgeries.
-
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
A training-free dual-system framework refines anomaly score ordering on uncertain samples from self-supervised talking head forgery detectors to improve detection performance.
-
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
Four-class audio-visual DeepFake detectors misclassify authentic but semantically mismatched audio-video pairs; five-class training plus ImageBind similarity improves models that can learn cross-modal semantics.
-
HumanOmni-Speaker: Identifying Who said What and When
HumanOmni-Speaker introduces a Visual Delta Encoder and VR-SDR benchmark that enable end-to-end speaker diarization and recognition by sampling video at 25 fps and compressing inter-frame motion residuals into 6 tokens per frame.
-
Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography
Splitting a steganographic payload across audio and video tracks reduces detection rates for both single-mode and multimodal detectors, though multimodal performance gains appear driven mostly by the video stream alone.