Pith. sign in

REVIEW 15 cited by

Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15818 v3 pith:7GSVLARA submitted 2023-09-27 cs.CV

classification cs.CV
keywords vdmsshow-1modelvideogenerationlatent-basedlow-resolutionpixel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs). However, previous methods either rely solely on pixel-based VDMs, which come with high computational costs, or on latent-based VDMs, which often struggle with precise text-video alignment. In this paper, we are the first to propose a hybrid model, dubbed as Show-1, which marries pixel-based and latent-based VDMs for text-to-video generation. Our model first uses pixel-based VDMs to produce a low-resolution video of strong text-video correlation. After that, we propose a novel expert translation method that employs the latent-based VDMs to further upsample the low-resolution video to high resolution, which can also remove potential artifacts and corruptions from low-resolution videos. Compared to latent VDMs, Show-1 can produce high-quality videos of precise text-video alignment; Compared to pixel VDMs, Show-1 is much more efficient (GPU memory usage during inference is 15G vs 72G). Furthermore, our Show-1 model can be readily adapted for motion customization and video stylization applications through simple temporal attention layer finetuning. Our model achieves state-of-the-art performance on standard video generation benchmarks. Our code and model weights are publicly available at https://github.com/showlab/Show-1.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReNeg: Learning Negative Embedding with Reward Guidance

    cs.CV 2024-12 conditional novelty 7.0 of 10

    ReNeg optimizes a negative text embedding with reward feedback and classifier-free guidance in the training loop, improving image-video generation quality over null-text and handcrafted negative prompts.

  2. Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Adding point-tracking supervision to video diffusion features reduces appearance drift in generated videos while preserving generation quality.

  3. Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model

    cs.GR 2025-07 conditional novelty 6.0 of 10

    Elevate3D refines low-quality 3D models by alternating high-frequency-guided texture redrawing with monocular-normal-driven geometry correction.

  4. NeoBabel: A Multilingual Open Tower for Visual Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.

  5. Enhance-A-Video: Better Generated Video for Free

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.

  6. Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    HALO aligns text-to-video diffusion models by jointly optimizing patch-level and video-level DPO rewards, with modest and partially self-referential benchmark gains.

  7. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  8. TransPixeler: Advancing Text-to-Video Generation with Transparency

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A LoRA-based adaptation of DiT video generators that jointly outputs aligned RGB and alpha channels via extra tokens, shared positions, and attention masking.

  9. JOG3R: Towards 3D-Consistent Video Generators

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Jointly training a video diffusion model with a 3D point map reconstruction head improves the 3D consistency of generated videos and yields usable camera pose estimates on static scenes.

  10. Grid: Omni Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.

  11. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.

  12. GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An iterative design-generate-redesign pipeline with four specialized LLM agents and self-routing correction improves compositional text-to-video generation on T2V-CompBench, with the largest gains in object numeracy.

  13. AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.

  14. ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Continual pre-training of a text-to-video model with block duplication plus an LLM-conditioned cross-attention improves benchmark scores, but aggregate gains hide several per-dimension regressions.

  15. Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment

    cs.CV 2025-01 conditional novelty 4.0 of 10

    MSA-VQA combines CLIP-based prompt checking with cross-attention over frames to predict human quality scores for AI-generated videos, reporting state-of-the-art numbers on the T2VQA-DB benchmark.

Pith tools