An SNN-based detector combining multi-channel pseudo-event residuals with frozen semantic features reaches 93.14% mean accuracy on unseen generators under the Pika-trained GenVideo protocol.
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3roles
background 2representative citing papers
UNICA unifies motion planning, rigging, physical simulation, and rendering into a single skeleton-free neural framework that produces next-frame 3D avatar geometry from action inputs and renders it with Gaussian splatting.
VideoPhy benchmark shows state-of-the-art text-to-video models follow physical commonsense and text prompts in only 39.6% of cases for the best model.
citing papers explorer
-
Detecting AI-Generated Videos with Spiking Neural Networks
An SNN-based detector combining multi-channel pseudo-event residuals with frozen semantic features reaches 93.14% mean accuracy on unseen generators under the Pika-trained GenVideo protocol.
-
UNICA: A Unified Neural Framework for Controllable 3D Avatars
UNICA unifies motion planning, rigging, physical simulation, and rendering into a single skeleton-free neural framework that produces next-frame 3D avatar geometry from action inputs and renders it with Gaussian splatting.
-
VideoPhy: Evaluating Physical Commonsense for Video Generation
VideoPhy benchmark shows state-of-the-art text-to-video models follow physical commonsense and text prompts in only 39.6% of cases for the best model.