REVIEW 31 cited by
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically experiment on simple or small datasets, where the generalization ability is quite limited. In this work, we propose GODIVA, an open-domain text-to-video pretrained model that can generate videos from text in an auto-regressive manner using a three-dimensional sparse attention mechanism. We pretrain our model on Howto100M, a large-scale text-video dataset that contains more than 136 million text-video pairs. Experiments show that GODIVA not only can be fine-tuned on downstream video generation tasks, but also has a good zero-shot capability on unseen texts. We also propose a new metric called Relative Matching (RM) to automatically evaluate the video generation quality. Several challenges are listed and discussed as future work.
Forward citations
Cited by 31 Pith papers
-
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
RDPO builds preference pairs by reverse-sampling real video latents with a pre-trained generator, then fine-tunes with Flow-DPO, improving physics consistency metrics on two video models.
-
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
Causal Forcing uses an autoregressive teacher for ODE initialization in diffusion distillation to close the causal attention gap and deliver better real-time video generation than Self Forcing.
-
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
BadVideo shows that text-to-video models can be backdoored by poisoning fine-tuning data with temporally distributed malicious content that evades frame-based moderation.
-
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.
-
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
A video editor trained on randomly distorted objects, then steered by adaptive noise at inference, is claimed to surpass dedicated and unified editors across eight tasks with far less training.
-
Neural Scene Designer: Self-Styled Semantic Image Manipulation
NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.
-
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
MiraMo turns a static image into a video using a linear-attention transformer that learns inter-frame motion residuals, with DCT-based noise refinement and a user-controllable dynamics knob.
-
Don't Forget your Inverse DDIM for Image Editing
SAGE edits real images with text by using self-attention maps from inverse DDIM as guidance, which preserves unedited regions without per-image optimization.
-
FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis
FreePCA improves training-free long video generation by projecting global and local attention features into PCA space, selecting consistent appearance components from global and motion components from local, then prog...
-
MV-Crafter: An Intelligent System for Music-guided Video Generation
MV-Crafter generates beat-synchronized music videos from music and a text theme by combining LLM-based scripting, diffusion video generation, and a dynamic beat-matching warping algorithm.
-
OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
OmniPhysGS lets each Gaussian in a 3D scene select from 12 expert material models, supervised by a text-to-video diffusion model, to generate dynamics for multiple materials.
-
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Controlled scaling shows a 2.3B self-attention U-ViT matches or slightly outperforms SDXL U-Net and larger cross-attention DiT variants, while long captions and larger datasets improve text-image alignment.
-
BrushEdit: All-In-One Image Inpainting and Editing
BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.
-
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
TIV-Diffusion adds object-centric slot alignment to a diffusion-based image-to-video generator and reports improved alignment and temporal-consistency metrics on MNIST, CATER, and Bridge datasets.
-
VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
VSD2M, a 2.09 million sample bilingual sticker dataset with animated GIFs, plus a Spatial Temporal Interaction layer, improves animated sticker generation over standard video diffusion baselines.
-
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
An iterative design-generate-redesign pipeline with four specialized LLM agents and self-routing correction improves compositional text-to-video generation on T2V-CompBench, with the largest gains in object numeracy.
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.
-
PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control
PainterNet is a diffusion-model plugin that uses local prompts, attention supervision, and diverse masks to improve text-consistent image inpainting.
-
I2VControl: Disentangled and Unified Video Motion Synthesis Control
I2VControl unifies camera, drag, and brush controls into a single point-trajectory-based adapter for image-to-video diffusion models, enabling conflict-free combined motion control.
-
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
ABot-World-0 claims real-time, long-horizon interactive world rollout on a single desktop GPU using raw keyboard actions, distillation, and low-bit inference, but the results cannot yet be independently checked becaus...
-
Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model
AIGVEval combines BLIP, 3D Swin Transformer, and SlowFast features with a LoRA-tuned LLM to predict AI-generated video quality, hitting second place on the NTIRE 2025 Track 2 leaderboard.
-
Semantics-Guided Generative Image Compression
Decoder-side ClipSeg segmentation and content-adaptive diffusion steps improve MISC-based ultra-low-bitrate generative image compression in perceptual quality and speed.
-
VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
VGDFR speeds up video diffusion generation by adaptively merging low-motion frames in latent space and adjusting positional embeddings, achieving up to 3x faster inference with modest quality changes.
-
RepVideo: Rethinking Cross-Layer Representation for Video Generation
A feature cache and gating mechanism that aggregates neighboring layer outputs in a video diffusion transformer improves temporal coherence and spatial accuracy.
-
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.
-
Stable Score Distillation
SSD is a diffusion score-distillation loss for text-guided 2D and 3D editing that combines a CFG cross-prompt term, a null-text cross-trajectory regularizer, and a prompt-enhancement term to stabilize edits.
-
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
MTADiffusion improves text-guided object inpainting by training on a new 5M-image mask-text dataset with edge prediction and style-consistency losses.
-
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
PAROAttention permutes tokens along frame, height, and width axes to make visual attention block-wise, enabling sparse and INT8/INT4 quantized attention with near-baseline generation quality.
-
A Survey of Interactive Generative Video
A survey that divides interactive generative video research into five modules: generation, control, memory, dynamics, and intelligence.
-
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
HuViDPO claims the first DPO-based alignment for text-to-video generation, but its loss reduces to the known DPO-SDXL objective and the evaluation is not reproducible.
-
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).
Discussion (0). Continue with ORCID to comment.