Pith. sign in

REVIEW 2 cited by

Do You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal Method for Temporal Forgery Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.06228 v2 pith:MBQGOQ7W submitted 2022-04-13 cs.CV

classification cs.CV
keywords deepfakedetectionforgerytemporalaudio-visualcontentdatasetlocalization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to its high societal impact, deepfake detection is getting active attention in the computer vision community. Most deepfake detection methods rely on identity, facial attributes, and adversarial perturbation-based spatio-temporal modifications at the whole video or random locations while keeping the meaning of the content intact. However, a sophisticated deepfake may contain only a small segment of video/audio manipulation, through which the meaning of the content can be, for example, completely inverted from a sentiment perspective. We introduce a content-driven audio-visual deepfake dataset, termed Localized Audio Visual DeepFake (LAV-DF), explicitly designed for the task of learning temporal forgery localization. Specifically, the content-driven audio-visual manipulations are performed strategically to change the sentiment polarity of the whole video. Our baseline method for benchmarking the proposed dataset is a 3DCNN model, termed as Boundary Aware Temporal Forgery Detection (BA-TFD), which is guided via contrastive, boundary matching, and frame classification loss functions. Our extensive quantitative and qualitative analysis demonstrates the proposed method's strong performance for temporal forgery localization and deepfake detection tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio Cross Verification Using Dual Alignment Likelihood Ratio Test

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A dual-alignment likelihood-ratio test verifies short audio queries against trusted reference recordings, reaching near-zero equal error rates for insertion/deletion tampering on the DAPS benchmark.

  2. SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms

    cs.LG 2025-06 reject novelty 4.0 of 10

    A benchmark of 2,126 Instagram videos labeled real or deepfake by uploader disclosure, evaluated with an LLM fact-checking pipeline that reaches 90.4% accuracy but conflates authenticity with factualness.

Pith tools