Pre-trained SigLIP/VideoMAE features with a linear probe or nearest-neighbor distance separate real videos from text-to-video model outputs, reaching about 90% average F1 on the new VID-AID benchmark, but much lower on the newest closed models.
V-jepa: Latent video prediction for visual represen- tation learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
Pre-trained SigLIP/VideoMAE features with a linear probe or nearest-neighbor distance separate real videos from text-to-video model outputs, reaching about 90% average F1 on the new VID-AID benchmark, but much lower on the newest closed models.