Pith. sign in

REVIEW 31 cited by

GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.14806 v1 pith:OU26JZVX submitted 2021-04-30 cs.CV

classification cs.CV
keywords godivavideosgeneratinggenerationmodelopen-domainproposetext
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically experiment on simple or small datasets, where the generalization ability is quite limited. In this work, we propose GODIVA, an open-domain text-to-video pretrained model that can generate videos from text in an auto-regressive manner using a three-dimensional sparse attention mechanism. We pretrain our model on Howto100M, a large-scale text-video dataset that contains more than 136 million text-video pairs. Experiments show that GODIVA not only can be fine-tuned on downstream video generation tasks, but also has a good zero-shot capability on unseen texts. We also propose a new metric called Relative Matching (RM) to automatically evaluate the video generation quality. Several challenges are listed and discussed as future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

    cs.CV 2025-06 conditional novelty 8.0 of 10

    RDPO builds preference pairs by reverse-sampling real video latents with a pre-trained generator, then fine-tunes with Flow-DPO, improving physics consistency metrics on two video models.

  2. Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    Causal Forcing uses an autoregressive teacher for ODE initialization in diffusion distillation to close the causal attention gap and deliver better real-time video generation than Self Forcing.

  3. BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation

    cs.CV 2025-04 conditional novelty 7.0 of 10

    BadVideo shows that text-to-video models can be backdoored by poisoning fine-tuning data with temporally distributed malicious content that evades frame-based moderation.

  4. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.

  5. O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A video editor trained on randomly distorted objects, then steered by adaptive noise at inference, is claimed to surpass dedicated and unified editors across eight tasks with far less training.

  6. Neural Scene Designer: Self-Styled Semantic Image Manipulation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.

  7. Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers

    cs.CV 2025-08 conditional novelty 6.0 of 10

    MiraMo turns a static image into a video using a linear-attention transformer that learns inter-frame motion residuals, with DCT-based noise refinement and a user-controllable dynamics knob.

  8. Don't Forget your Inverse DDIM for Image Editing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SAGE edits real images with text by using self-attention maps from inverse DDIM as guidance, which preserves unedited regions without per-image optimization.

  9. FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FreePCA improves training-free long video generation by projecting global and local attention features into PCA space, selecting consistent appearance components from global and motion components from local, then prog...

  10. MV-Crafter: An Intelligent System for Music-guided Video Generation

    cs.HC 2025-04 conditional novelty 6.0 of 10

    MV-Crafter generates beat-synchronized music videos from music and a text theme by combining LLM-based scripting, diffusion video generation, and a dynamic beat-matching warping algorithm.

  11. OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    OmniPhysGS lets each Gaussian in a 3D scene select from 12 expert material models, supervised by a text-to-video diffusion model, to generate dynamics for multiple materials.

  12. Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Controlled scaling shows a 2.3B self-attention U-ViT matches or slightly outperforms SDXL U-Net and larger cross-attention DiT variants, while long captions and larger datasets improve text-image alignment.

  13. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  14. TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TIV-Diffusion adds object-centric slot alignment to a diffusion-based image-to-video generator and reports improved alignment and temporal-consistency metrics on MNIST, CATER, and Bridge datasets.

  15. VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation

    cs.HC 2024-12 conditional novelty 6.0 of 10

    VSD2M, a 2.09 million sample bilingual sticker dataset with animated GIFs, plus a Spatial Temporal Interaction layer, improves animated sticker generation over standard video diffusion baselines.

  16. GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An iterative design-generate-redesign pipeline with four specialized LLM agents and self-routing correction improves compositional text-to-video generation on T2V-CompBench, with the largest gains in object numeracy.

  17. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  18. PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PainterNet is a diffusion-model plugin that uses local prompts, attention supervision, and diverse masks to improve text-consistent image inpainting.

  19. I2VControl: Disentangled and Unified Video Motion Synthesis Control

    cs.CV 2024-11 conditional novelty 6.0 of 10

    I2VControl unifies camera, drag, and brush controls into a single point-trajectory-based adapter for image-to-video diffusion models, enabling conflict-free combined motion control.

  20. ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ABot-World-0 claims real-time, long-horizon interactive world rollout on a single desktop GPU using raw keyboard actions, distillation, and low-bit inference, but the results cannot yet be independently checked becaus...

  21. Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    AIGVEval combines BLIP, 3D Swin Transformer, and SlowFast features with a LoRA-tuned LLM to predict AI-generated video quality, hitting second place on the NTIRE 2025 Track 2 leaderboard.

  22. Semantics-Guided Generative Image Compression

    eess.IV 2025-05 conditional novelty 5.0 of 10

    Decoder-side ClipSeg segmentation and content-adaptive diffusion steps improve MISC-based ultra-low-bitrate generative image compression in perceptual quality and speed.

  23. VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate

    cs.CV 2025-04 conditional novelty 5.0 of 10

    VGDFR speeds up video diffusion generation by adaptively merging low-motion frames in latent space and adjusting positional embeddings, achieving up to 3x faster inference with modest quality changes.

  24. RepVideo: Rethinking Cross-Layer Representation for Video Generation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A feature cache and gating mechanism that aggregates neighboring layer outputs in a video diffusion transformer improves temporal coherence and spatial accuracy.

  25. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

  26. Stable Score Distillation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    SSD is a diffusion score-distillation loss for text-guided 2D and 3D editing that combines a CFG cross-prompt term, a null-text cross-trajectory regularizer, and a prompt-enhancement term to stabilize edits.

  27. MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

    cs.CV 2025-06 conditional novelty 4.0 of 10

    MTADiffusion improves text-guided object inpainting by training on a new 5M-image mask-text dataset with edge prediction and style-consistency losses.

  28. PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    PAROAttention permutes tokens along frame, height, and width axes to make visual attention block-wise, enabling sparse and INT8/INT4 quantized attention with near-baseline generation quality.

  29. A Survey of Interactive Generative Video

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A survey that divides interactive generative video research into five modules: generation, control, memory, dynamics, and intelligence.

  30. HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment

    cs.CV 2025-02 reject novelty 4.0 of 10

    HuViDPO claims the first DPO-based alignment for text-to-video generation, but its loss reduces to the known DPO-SDXL objective and the evaluation is not reproducible.

  31. Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).

Pith tools