Pith. sign in

REVIEW 2 cited by

TruthLens: Visual Grounding for Universal DeepFake Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.15867 v3 pith:MQKX34U5 submitted 2025-03-20 cs.CV cs.AI

TruthLens: Visual Grounding for Universal DeepFake Reasoning

classification cs.CV cs.AI
keywords truthlensgroundingreasoningbeyondbinaryclassificationcontentdeepfake
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Detecting DeepFakes has become a crucial research area as the widespread use of AI image generators enables the effortless creation of face-manipulated and fully synthetic content, while existing methods are often limited to binary classification (real vs. fake) and lack interpretability. To address these challenges, we propose TruthLens, a novel, unified, and highly generalizable framework that goes beyond traditional binary classification, providing detailed, textual reasoning for its predictions. Distinct from conventional methods, TruthLens performs MLLM grounding. TruthLens uses a task-driven representation integration strategy that unites global semantic context from a multimodal large language model (MLLM) with region-specific forensic cues through explicit cross-modal adaptation of a vision-only model. This enables nuanced, region-grounded reasoning for both face-manipulated and fully synthetic content, and supports fine-grained queries such as "Does the eyes/nose/mouth look real or fake?"- capabilities beyond pretrained MLLMs alone. Extensive experiments across diverse datasets demonstrate that TruthLens sets a new benchmark in both forensic interpretability and detection accuracy, generalizing to seen and unseen manipulations alike. By unifying high-level scene understanding with fine-grained region grounding, TruthLens delivers transparent DeepFake forensics, bridging a critical gap in the literature.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

    cs.CV 2026-03 conditional novelty 6.5

    A plan-then-examine MLLM framework with stage-gated part-level forensic injection and part-aware RL rewards outperforms expert and concurrent MLLM deepfake detectors across a hierarchical 5-level generalizability benchmark.

  2. VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

    cs.CV 2026-04 unverdicted novelty 5.0

    VRAG-DFD uses RAG to retrieve forgery knowledge and RL-based training to build critical reasoning in MLLMs, delivering state-of-the-art generalization on deepfake detection tasks.