Pith. sign in

REVIEW 16 cited by

VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09047 v1 pith:34XB6TJB submitted 2024-01-17 cs.CV

classification cs.CV
keywords high-qualitymodelsvideomodulesvideoslow-qualityspatialtemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with minimal noise, excellent details, and high aesthetic scores. However, these models rely on large-scale, well-filtered, high-quality videos that are not accessible to the community. Many existing research works, which train models using the low-quality WebVid-10M dataset, struggle to generate high-quality videos because the models are optimized to fit WebVid-10M. In this work, we explore the training scheme of video models extended from Stable Diffusion and investigate the feasibility of leveraging low-quality videos and synthesized high-quality images to obtain a high-quality video model. We first analyze the connection between the spatial and temporal modules of video models and the distribution shift to low-quality videos. We observe that full training of all modules results in a stronger coupling between spatial and temporal modules than only training temporal modules. Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model. Evaluations are conducted to demonstrate the superiority of the proposed method, particularly in picture quality, motion, and concept composition.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.

  2. Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures

    cs.AR 2026-05 unverdicted novelty 7.0 of 10

    First systematic column-level sparsity profiling across seven diffusion workloads reveals element-level sparsity overstates hardware savings by up to 78 points and identifies a three-way taxonomy of concentration vs. ...

  3. MoRight: Motion Control Done Right

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    MoRight disentangles object and camera motion via canonical-view specification and temporal cross-view attention, while decomposing motion into active user-driven and passive consequence components to learn and apply ...

  4. GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

    cs.CV 2026-02 unverdicted novelty 7.0 of 10

    GT-SVJ turns video generative models into self-supervised reward judges via EBM reformulation and contrastive training on controlled synthetic degradations, claiming SOTA on GenAI-Bench and MonteBench with 30K annotations.

  5. Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

    cs.CV 2024-11 unverdicted novelty 7.0 of 10

    VideoRepair detects text-video misalignments via MLLM-generated questions and performs localized, region-preserving refinement to improve alignment in existing T2V diffusion models.

  6. VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An MLLM extracts transferable physical cues from a reference video and conditions a pretrained I2V generator so new scenes follow that physics without exhaustive prompts.

  7. Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks

    cs.HC 2026-05 conditional novelty 6.0 of 10

    Fully aligned instructional videos for physical tasks yield 11.1% better completion quality and 15.5% faster times, with four decomposable visual attributes whose isolated misalignments degrade performance without use...

  8. TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    TS-Attn dynamically separates and rearranges attention in existing text-to-video models to improve temporal consistency and prompt adherence for videos with multiple sequential actions.

  9. DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    DVAR turns video authenticity detection into an iterative debate between a generative hypothesis agent and a natural mechanism agent, resolved via minimum description length and a knowledge base for better generalizat...

  10. Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    A training-free framework uses physics-violating counterfactual prompts and Synchronized Decoupled Guidance to suppress implausible motions in diffusion-based video generation while preserving photorealism.

  11. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  12. Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

    cs.CV 2025-07 unverdicted novelty 6.0 of 10

    Geometry Forcing aligns video diffusion representations with geometric foundation model features via angular cosine and scale regression objectives to improve 3D consistency in generated videos.

  13. VideoPhy: Evaluating Physical Commonsense for Video Generation

    cs.CV 2024-06 conditional novelty 6.0 of 10

    VideoPhy benchmark shows state-of-the-art text-to-video models follow physical commonsense and text prompts in only 39.6% of cases for the best model.

  14. Retrieval-Driven Training-Free AI-Generated Video Attribution

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free retrieval pipeline using adaptive color transforms, multi-scale quantized residuals, and temporal aggregation attributes AI-generated videos to one of eight generators with 84.6% Rank-1 and 78.3% mAP o...

  15. RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Reversing bit-plane weights ('bit-reversed image') plus a gradient-selected 32×32 patch lets a small ResNet detect AI-generated images with state-of-the-art accuracy on many benchmarks.

  16. 3D Reconstruction Techniques in the Manufacturing Domain: Applications, Research Opportunities and Use Cases

    cs.CV 2026-04 unverdicted novelty 2.0 of 10

    A survey of 106 papers finds quality inspection dominates 3D reconstruction use in manufacturing at 40 percent of applications, with a shift toward hybrid sensor systems and a noted gap in unified frameworks.

Pith tools