Pith. sign in

GOTCHA: Real-Time Video Deepfake Detection via Challenge-Response

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

With the rise of AI-enabled Real-Time Deepfakes (RTDFs), the integrity of online video interactions has become a growing concern. RTDFs have now made it feasible to replace an imposter's face with their victim in live video interactions. Such advancement in deepfakes also coaxes detection to rise to the same standard. However, existing deepfake detection techniques are asynchronous and hence ill-suited for RTDFs. To bridge this gap, we propose a challenge-response approach that establishes authenticity in live settings. We focus on talking-head style video interaction and present a taxonomy of challenges that specifically target inherent limitations of RTDF generation pipelines. We evaluate representative examples from the taxonomy by collecting a unique dataset comprising eight challenges, which consistently and visibly degrades the quality of state-of-the-art deepfake generators. These results are corroborated both by humans and a new automated scoring function, leading to 88.6% and 80.1% AUC, respectively. The findings underscore the promising potential of challenge-response systems for explainable and scalable real-time deepfake detection in practical scenarios. We provide access to data and code at \url{https://github.com/mittalgovind/GOTCHA-Deepfakes}.

fields

cs.AI 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

Modeling Human Responses to Multimodal AI Content

cs.AI · 2025-08-14 · unverdicted · novelty 5.0

A 154K-post study reports that humans identify AI content best when text and images are both present and inconsistent, and offers metrics plus an LLM agent for human-aligned responses.

citing papers explorer

Showing 1 of 1 citing paper.

  • Modeling Human Responses to Multimodal AI Content cs.AI · 2025-08-14 · unverdicted · none · ref 57 · internal anchor

    A 154K-post study reports that humans identify AI content best when text and images are both present and inconsistent, and offers metrics plus an LLM agent for human-aligned responses.