Pith. sign in

hub Canonical reference

Is space-time attention all you need for video understanding?

Canonical reference. 83% of citing Pith papers cite this work as background.

14 Pith papers citing it
1,359 external citations · external index
Background 83% of classified citations

hub tools

citation-role summary

background 5 baseline 1

citation-polarity summary

representative citing papers

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection

cs.CV · 2025-12-11 · conditional · novelty 8.0

RobustSora benchmark demonstrates that current AI video detectors rely heavily on visible watermarks, with average accuracy drops of 6.6 percentage points when watermarks are erased and increased false alarms when watermarks are spoofed onto real videos.

Video Diffusion Models

cs.CV · 2022-04-07 · unverdicted · novelty 7.0

A diffusion model for video generation extends image architectures with joint image-video training and improved conditional sampling, delivering first large-scale text-to-video results and state-of-the-art performance on video prediction and unconditional generation benchmarks.

citing papers explorer

Showing 14 of 14 citing papers.