REVIEW 15 cited by
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs). However, previous methods either rely solely on pixel-based VDMs, which come with high computational costs, or on latent-based VDMs, which often struggle with precise text-video alignment. In this paper, we are the first to propose a hybrid model, dubbed as Show-1, which marries pixel-based and latent-based VDMs for text-to-video generation. Our model first uses pixel-based VDMs to produce a low-resolution video of strong text-video correlation. After that, we propose a novel expert translation method that employs the latent-based VDMs to further upsample the low-resolution video to high resolution, which can also remove potential artifacts and corruptions from low-resolution videos. Compared to latent VDMs, Show-1 can produce high-quality videos of precise text-video alignment; Compared to pixel VDMs, Show-1 is much more efficient (GPU memory usage during inference is 15G vs 72G). Furthermore, our Show-1 model can be readily adapted for motion customization and video stylization applications through simple temporal attention layer finetuning. Our model achieves state-of-the-art performance on standard video generation benchmarks. Our code and model weights are publicly available at https://github.com/showlab/Show-1.
Forward citations
Cited by 15 Pith papers
-
ReNeg: Learning Negative Embedding with Reward Guidance
ReNeg optimizes a negative text embedding with reward feedback and classifier-free guidance in the training loop, improving image-video generation quality over null-text and handcrafted negative prompts.
-
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
Adding point-tracking supervision to video diffusion features reduces appearance drift in generated videos while preserving generation quality.
-
Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
Elevate3D refines low-quality 3D models by alternating high-frequency-guided texture redrawing with monocular-normal-driven geometry correction.
-
NeoBabel: A Multilingual Open Tower for Visual Generation
A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.
-
Enhance-A-Video: Better Generated Video for Free
Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.
-
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models
HALO aligns text-to-video diffusion models by jointly optimizing patch-level and video-level DPO rewards, with modest and partially self-referential benchmark gains.
-
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.
-
TransPixeler: Advancing Text-to-Video Generation with Transparency
A LoRA-based adaptation of DiT video generators that jointly outputs aligned RGB and alpha channels via extra tokens, shared positions, and attention masking.
-
JOG3R: Towards 3D-Consistent Video Generators
Jointly training a video diffusion model with a 3D point map reconstruction head improves the 3D consistency of generated videos and yields usable camera pose estimates on static scenes.
-
Grid: Omni Visual Generation
GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.
-
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.
-
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
An iterative design-generate-redesign pipeline with four specialized LLM agents and self-routing correction improves compositional text-to-video generation on T2V-CompBench, with the largest gains in object numeracy.
-
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.
-
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
Continual pre-training of a text-to-video model with block duplication plus an LLM-conditioned cross-attention improves benchmark scores, but aggregate gains hide several per-dimension regressions.
-
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
MSA-VQA combines CLIP-based prompt checking with cross-attention over frames to predict human quality scores for AI-generated videos, reporting state-of-the-art numbers on the T2VQA-DB benchmark.
Discussion (0). Continue with ORCID to comment.