Pith. sign in

REVIEW 5 cited by

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.10193 v2 pith:UGE2COYG submitted 2024-11-15 cs.CV

classification cs.CV
keywords deepfakedetectiondimodifaudiovisuallocalizationaudio-visualav-deepfake1m
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deepfake technology has rapidly advanced and poses significant threats to information integrity and trust in online multimedia. While significant progress has been made in detecting deepfakes, the simultaneous manipulation of audio and visual modalities, sometimes at small parts or in subtle ways, presents highly challenging detection scenarios. To address these challenges, we present DiMoDif, an audio-visual deepfake detection framework that leverages the inter-modality differences in machine perception of speech, based on the assumption that in real samples -- in contrast to deepfakes -- visual and audio signals coincide in terms of information. DiMoDif leverages features from deep networks that specialize in visual and audio speech recognition to spot frame-level cross-modal incongruities, and in that way to temporally localize the deepfake forgery. To this end, we devise a hierarchical cross-modal fusion network, integrating adaptive temporal alignment modules and a learned discrepancy mapping layer to explicitly model the subtle differences between visual and audio representations. Then, the detection model is optimized through a composite loss function accounting for frame-level detections and fake intervals localization. DiMoDif outperforms the state-of-the-art on the Deepfake Detection task by 30.5 AUC on the highly challenging AV-Deepfake1M, while it performs exceptionally on FakeAVCeleb and LAV-DF. On the Temporal Forgery Localization task, it outperforms the state-of-the-art by 47.88 AP@0.75 on AV-Deepfake1M, and performs on-par on LAV-DF. Code available at https://github.com/mever-team/dimodif.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Skip-scanning Mamba with unified audio-visual sequences reaches 63.4% AP@0.95 on LAV-DF and 63.58% mAP on AV-Deepfake1M by regularizing toward low/mid-frequency forgery cues.

  2. EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

    cs.CV 2026-07 conditional novelty 6.0 of 10

    EVAS localizes sparse multimodal forgeries via multi-stage audio-visual synergy and decoupled boundary-aware refinement, reporting SOTA AP and AR on LAV-DF, AV-Deepfake1M, and TVIL.

  3. Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    UniCaCLF trains temporal instant features to separate real from forged moments relative to each sample's global context, achieving state-of-the-art temporal forgery localization on five public datasets.

  4. DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection

    cs.MM 2025-06 conditional novelty 6.0 of 10

    Proposes new cross-manipulation evaluation protocols for FakeAVCeleb and DeepSpeak v1, shows temporal jittering mitigates a leading-silence shortcut, and introduces the SIMBA baseline.

  5. Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A systematic review of deepfake detection finds a pervasive lack of adversarial robustness evaluation across all modalities and calls for resilient, modality-agnostic detectors.

Pith tools