REVIEW 16 cited by
Adversarial Video Generation on Complex Datasets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative models of natural images have progressed towards high fidelity samples by the strong leveraging of scale. We attempt to carry this success to the field of video modeling by showing that large Generative Adversarial Networks trained on the complex Kinetics-600 dataset are able to produce video samples of substantially higher complexity and fidelity than previous work. Our proposed model, Dual Video Discriminator GAN (DVD-GAN), scales to longer and higher resolution videos by leveraging a computationally efficient decomposition of its discriminator. We evaluate on the related tasks of video synthesis and video prediction, and achieve new state-of-the-art Fr\'echet Inception Distance for prediction for Kinetics-600, as well as state-of-the-art Inception Score for synthesis on the UCF-101 dataset, alongside establishing a strong baseline for synthesis on Kinetics-600.
Forward citations
Cited by 16 Pith papers
-
FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers
A DiT-based portrait animation model transfers implicit facial expressions to one or more characters using a masked cross-attention mechanism, supported by a new multi-face dataset and benchmark.
-
Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation
TANGO avoids "terminal points" in autoregressive video generation by optimizing a LoRA adapter at test time until the model's one-step-ahead predicted noise looks like isotropic Gaussian noise.
-
ELT: Elastic Looped Transformers for Visual Generation
Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.
-
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
TARDIS is a transformer world model trained on STRIDE, a graph-structured street-view dataset, with claimed abilities in controllable image generation, georeferencing, self-driving actions, and temporal simulation.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
A training-free method that lifts 2D object trajectories into camera poses with depth and uses a frozen camera-control video model to achieve more accurate and 3D-aware object motion, including rotation.
-
Extended Field of View Analysis for VideoGAN-based Trajectory Generation
A video GAN trained on semantic top-down traffic videos generates 15–25 m field-of-view scenes whose speed, acceleration, spacing, and time-to-collision statistics resemble real Waymo data, with inference below 20 ms.
-
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.
-
Seeing Voices: Generating A-Roll Video from Audio with Mirage
Mirage generates photorealistic A-roll videos of people speaking directly from audio, using only joint self-attention over audio, text, and video tokens.
-
LIA-X: Interpretable Latent Portrait Animator
Adding an L1 sparsity penalty to the motion dictionary of the LIA portrait animator produces disentangled, human-interpretable motion vectors that support controllable image and video editing and scale to roughly one ...
-
Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.
-
DTSGAN: Learning Dynamic Textures via Spatiotemporal Generative Adversarial Network
DTSGAN adapts SinGAN-style multi-scale generation to video with 3D convolutions and a sliding-window data update, claiming improved dynamic texture synthesis and diversity.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
-
Efficient Continuous Video Flow Model for Video Prediction
The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.
-
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
CVP trains a network to reverse a continuous interpolation between past and future frames, reporting competitive FVD scores and 25-step sampling on KTH, BAIR, Human3.6M, and UCF101.
-
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
NBP generates video rows in parallel with block-wise attention, achieving 11x faster inference and better FVD than next-token autoregressive models.
Discussion (0). Continue with ORCID to comment.