REVIEW 2 major objections 2 minor 323 cited by
Make-A-Video: Text-to-Video Generation without Text-Video Data
T0 review · 2 major / 2 minor · reviewed 2026-05-11 · grok-4.3
Pith's one-line read A method turns text into videos by extending image generators with motion learned separately from unlabeled footage.
desk verdict Make-A-Video shows a workable split between image appearance and video motion to skip paired text-video data, but the SOTA claim sits on asserted results rather than displayed evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Spatial-temporal decomposition of U-Net and attention tensors together with a multi-stage pipeline of video decoder, interpolation, and super-resolution models.
What would settle it
A side-by-side evaluation on the same text prompts where Make-A-Video outputs show more flickering, unnatural object trajectories, or lower text-video alignment scores than models trained directly on paired text-video data.
Extended reading notes
Core claim
Make-A-Video decomposes the temporal U-Net and attention tensors into separate spatial and temporal approximations and then runs a spatial-temporal pipeline that includes a video decoder, an interpolation model, and two super-resolution models. The system re-uses a pre-trained text-to-image model for visual content and text alignment while adding motion learned from unsupervised video. The outcome is state-of-the-art text-to-video output in resolution, frame rate, text faithfulness, and overall quality, achieved without any paired text-video training data.
Load-bearing premise
Motion patterns taken from unlabeled video can be added to a text-to-image model through these modules without creating visible motion artifacts or weakening how well the output matches the original text prompt.
Editorial extensions
If this is right
- Text-to-video training becomes faster because visual and language representations are reused rather than learned from scratch.
- Paired text-video datasets are no longer required to reach competitive performance.
- The generated videos carry over the aesthetic variety and fantastical content already present in current text-to-image systems.
- High-resolution and high-frame-rate results are produced by chaining the dedicated interpolation and super-resolution stages.
Reading between the lines
- The same separation of appearance learning from motion learning could be tried on other data-scarce generation tasks such as 3D or audio synthesis.
- Modular pipelines like this one may reduce the total compute needed when extending image models to new domains.
- The approach opens a route to video editing or animation tools that start from a single text prompt and then refine motion independently.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Make-A-Video, a text-to-video generation method that transfers progress from text-to-image (T2I) models by learning appearance and text alignment from paired text-image data while acquiring motion dynamics from unsupervised video footage. It introduces a spatial-temporal decomposition of the U-Net and attention tensors, combined with a multi-stage pipeline (video decoder, temporal interpolation, and super-resolution models) to produce high-resolution, high-frame-rate videos without requiring paired text-video data. The central claim is that this yields state-of-the-art results in spatial/temporal resolution, text faithfulness, and perceptual quality, as measured by both qualitative examples and quantitative metrics.
Significance. If the quantitative claims hold, the work is significant because it demonstrates a practical route to high-quality T2V generation that sidesteps the scarcity of paired text-video data, accelerates training by reusing T2I representations, and inherits the diversity of modern image generators. The decomposition approach and modular pipeline are reusable for other video synthesis tasks and could reduce compute barriers in the field.
major comments (2)
- [§4] §4 (Experiments): The SOTA claim is central but rests on quantitative comparisons whose details (specific metrics such as FVD, CLIP similarity, or human preference scores, exact baselines, and effect sizes) are not summarized in the abstract and must be verified against prior T2V methods; without these numbers and ablations on the spatial-temporal modules, the superiority cannot be assessed.
- [§3.2] §3.2 (Spatial-Temporal Decomposition): The approximation of full temporal U-Net and attention tensors in space and time is described at a high level; the paper must supply the precise tensor factorization or insertion points (e.g., which layers receive the temporal attention) to confirm that motion transfer occurs without degrading text conditioning or introducing systematic artifacts.
minor comments (2)
- [Abstract] The abstract and introduction could more explicitly list the quantitative metrics and baselines used to support the SOTA statement.
- [Figures] Figure captions for qualitative results should include the exact text prompts and frame counts to aid reproducibility.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address each major comment below with clarifications from the paper and propose targeted revisions to strengthen the presentation of our results and technical details.
read point-by-point responses
-
Referee: [§4] §4 (Experiments): The SOTA claim is central but rests on quantitative comparisons whose details (specific metrics such as FVD, CLIP similarity, or human preference scores, exact baselines, and effect sizes) are not summarized in the abstract and must be verified against prior T2V methods; without these numbers and ablations on the spatial-temporal modules, the superiority cannot be assessed.
Authors: We agree that a concise summary of the key quantitative results would improve accessibility. Section 4 reports FVD, CLIP similarity, and human preference scores against baselines including CogVideo and other recent T2V methods, with effect sizes and ablations on the spatial-temporal modules detailed in Tables 1-3 and Section 4.3 (plus appendix). The abstract states the SOTA outcome but does not list the numbers. We will revise the abstract to include a brief summary of the primary metrics and baselines while retaining the existing detailed comparisons in the experiments section. revision: partial
-
Referee: [§3.2] §3.2 (Spatial-Temporal Decomposition): The approximation of full temporal U-Net and attention tensors in space and time is described at a high level; the paper must supply the precise tensor factorization or insertion points (e.g., which layers receive the temporal attention) to confirm that motion transfer occurs without degrading text conditioning or introducing systematic artifacts.
Authors: We appreciate this request for greater precision. Section 3.2 describes the decomposition of the U-Net and attention tensors into separate spatial and temporal factors, with temporal attention inserted after spatial attention in the decoder blocks to enable motion modeling while preserving the pretrained text-image conditioning pathway. To address the comment directly, we will add a detailed diagram and explicit layer specifications (including tensor shapes and insertion points) in the revised Section 3.2. revision: yes
Circularity Check
No significant circularity in derivation chain
full rationale
The paper presents Make-A-Video as a pipeline that inherits appearance from external pretrained T2I models and motion from separate unsupervised video data. It describes a spatial-temporal decomposition of U-Net/attention tensors plus a multi-stage generation pipeline (video decoder, interpolation, super-resolution). No load-bearing step reduces by construction to a self-fit, self-definition, or self-citation chain; the central claim is a concrete engineering combination of independent pretrained components rather than a tautological prediction. The SOTA assertion rests on external qualitative/quantitative evaluation, not internal re-derivation of inputs.
Assumptions & free parameters
assumptions (1)
- domain assumption Decomposing full temporal U-Net and attention tensors into separate spatial and temporal approximations preserves sufficient modeling capacity for coherent video generation.
Cite this review
Pith. "Pith review of Make-A-Video: Text-to-Video Generation without Text-Video Data." pith.science (2026). https://pith.science/paper/RMT4HSRN
@misc{pith2026220914792,
author = {Pith},
title = {Pith review of: Make-A-Video: Text-to-Video Generation without Text-Video Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/RMT4HSRN}},
note = {Machine review of arXiv:2209.14792}
}
read the original abstract
We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in aesthetic, fantastical depictions, etc.) of today's image generation models. We design a simple yet effective way to build on T2I models with novel and effective spatial-temporal modules. First, we decompose the full temporal U-Net and attention tensors and approximate them in space and time. Second, we design a spatial temporal pipeline to generate high resolution and frame rate videos with a video decoder, interpolation model and two super resolution models that can enable various applications besides T2V. In all aspects, spatial and temporal resolution, faithfulness to text, and quality, Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures.
Forward citations
Showing 60 of 323 Pith papers that cite this
-
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
PhyAVBench supplies the first benchmark and contrastive metric that measures whether text-to-audio-video models respect real-world audio physics across controlled prompt pairs.
-
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
RDPO builds preference pairs by reverse-sampling real video latents with a pre-trained generator, then fine-tunes with Flow-DPO, improving physics consistency metrics on two video models.
-
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
H2R-Bench is a benchmark that evaluates whether video generation models can transfer human manipulation demonstrations into robot videos with correct embodiment, contact, and task completion, and it finds current mode...
-
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.
-
OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
OrbitQuant is a data-agnostic PTQ technique for DiTs that uses RPBH rotation in a normalized basis to enable a single codebook across all inputs, achieving SOTA low-bit performance on FLUX.1, CogVideoX and similar models.
-
HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention
HCMS partitions multi-head attention into chunks and pipelines them across dual CUDA streams to overlap communication and computation, delivering 10-17.5% speedup over Ulysses for 31K-56K token sequences.
-
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Goku supplies a 2M-scale dataset, synthesis pipeline, decoupled dual-branch model, and 1000-case benchmark for multi-task instruction-based video editing, reporting up to 8% gains in instruction following.
-
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
None of ten tested video-generation models reliably remembers objects after occlusion in dynamic scenes; static-camera videos inflate consistency scores.
-
Net-Ev$^2$: A Generative Simulator for Network Event Evolution
Net-Ev² proposes a two-stage generative simulator with structure-guided masked pre-training and topology-aware diffusion using graph U-Net down/upsampling to model network event evolution from text inputs, plus a new ...
-
FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
FadeMem introduces distance-aware KV memory consolidation for autoregressive video diffusion that builds a temporal hierarchy with power-law merging to preserve short-term dynamics and long-range coherence under fixed...
-
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
Dream.exe evaluates 8 video generation models on 101 manipulation tasks by converting generated videos into executable robot trajectories in a simulator, finding measurable success rates that visual metrics do not predict.
-
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
SafeGen-Bench is a benchmark with 10 malicious categories that evaluates conditional T2V models on paired start frames and text prompts, finding unsafety scores up to 44.5 and 80% guardrail failure rate.
-
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
MBench is a new benchmark that quantifies long-term memory in video world models via three hierarchical consistency dimensions evaluated on curated real videos.
-
YoCausal: How Far is Video Generation from World Model? A Causality Perspective
YoCausal benchmark shows video diffusion models detect the arrow of time but lack genuine causal understanding relative to humans.
-
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
CRONOS benchmark shows recent open-source video generators fail to preserve physical consistency under controlled changes to viewpoint, scene, object category, and appearance.
-
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
MSAVBench is the first comprehensive benchmark for multi-shot audio-video generation, spanning video, audio, shot, and reference dimensions with an adaptive evaluation framework that reaches 91.5% Spearman correlation...
-
Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls
Aero-World adapts a pretrained latent diffusion transformer for action-conditioned aerial video generation by injecting inertial action tokens and using a frozen latent-space Physics Probe for inertial consistency sup...
-
Functionalization via Structure Completion and Motion Rectification
Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture wi...
-
StreamingEffect: Real-Time Human-Centric Video Effect Generation
StreamingEffect enables real-time 720p human-centric video effect generation on one GPU via teacher-student distillation, keyframe control, and a new 130K video dataset.
-
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
TeDiO regularizes temporal diagonals in diffusion transformer attention maps to produce smoother video motion while keeping per-frame quality intact.
-
R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow
R-DMesh uses a VAE with a learned rectification jump offset and Triflow Attention inside a rectified-flow diffusion transformer to produce video-aligned 4D meshes despite initial pose misalignment.
-
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
GTA generates 3D worlds from single images via a two-stage video diffusion process that prioritizes geometry before appearance to improve structural consistency.
-
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
V2V-Zero adapts frozen VLMs for visual conditioning via hidden states from specification pages, scoring 0.85 on GenEval and 32.7 on a new seven-task benchmark while revealing capability hierarchies in attribute bindin...
-
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
Recursive generative retraining with pluralistic preferences converges to a stable diverse distribution that satisfies a weighted Nash bargaining solution.
-
Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges
Structured diffusion bridges with alignment constraints achieve near fully-paired quality in modality translation while working effectively in unpaired and semi-paired regimes.
-
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
Encoding cameras as pixel-aligned raxels lets one video diffusion model jointly denoise video and trajectories, supporting pose estimation, controlled generation, and joint synthesis.
-
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding
Omni-NegCLIP improves CLIP's negation understanding by up to 52.65% on presence-based and 12.50% on absence-based tasks through front-layer fine-tuning with specialized contrastive losses.
-
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
ChopGrad truncates backpropagation to local frame windows in video diffusion models, reducing memory from linear in frame count to constant while enabling pixel-wise loss fine-tuning.
-
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
FrameDiT proposes Matrix Attention for DiTs to achieve SOTA video generation with improved temporal coherence and efficiency comparable to local factorized attention.
-
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
Causal Forcing uses an autoregressive teacher for ODE initialization in diffusion distillation to close the causal attention gap and deliver better real-time video generation than Self Forcing.
-
One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
One-to-All Animation enables alignment-free character animation and image pose transfer via self-supervised outpainting reformulation, reference extraction, hybrid fusion attention, identity-robust pose control, and t...
-
ASTRA: Let Arbitrary Subjects Transform in Video Editing
ASTRA is a plug-and-play training-free method for precise multi-subject video editing that uses prompt-guided multimodal alignment and prior-based mask retargeting to avoid attention dilution and boundary issues.
-
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.
-
GUAVA: Generalizable Upper Body 3D Gaussian Avatar
From a single image, GUAVA builds an animatable upper-body 3D Gaussian avatar in one forward pass, using a new hybrid SMPLX/FLAME template and inverse texture mapping, then renders it in real time.
-
STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
STP4D directly denoises 4D Gaussian splatting attributes with DDIM conditioned on time-varying text embeddings, producing high-fidelity 4D assets in about 4.6 seconds.
-
History-Guided Video Diffusion
DFoT enables flexible history conditioning in video diffusion, with history guidance methods that boost temporal consistency and support long rollouts.
-
Fast Video Generation with Sliding Tile Attention
Sliding tile attention (STA) replaces full 3D attention in video diffusion transformers with dense tile-local windows, achieving 1.89x training-free and up to 3.53x fine-tuned end-to-end speedups on HunyuanVideo with ...
-
Every Image Listens, Every Image Dances: Music-Driven Image Animation
MuseDance animates a reference image into a music-synchronized dance video conditioned only on the audio track and a text description, and contributes a new 2,904-video dataset.
-
Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images
Stem uses a conditional diffusion model to infer spatially resolved gene expression from H&E histology images, outperforming regression baselines on several datasets.
-
Can Generative Video Models Help Pose Estimation?
Generating intermediate frames with a video model and feeding them to DUSt3R improves pairwise pose estimation for low-overlap images, when a medoid-based self-consistency score selects the best generated video.
-
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.
-
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
Adding point-tracking supervision to video diffusion features reduces appearance drift in generated videos while preserving generation quality.
-
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
VideoRepair detects text-video misalignments via MLLM-generated questions and performs localized, region-preserving refinement to improve alignment in existing T2V diffusion models.
-
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
Embedding Sanitizer removes inappropriate concepts directly from text prompt embeddings with per-token scores, claiming state-of-the-art robustness against adversarial prompts.
-
RoboDreamer: Learning Compositional World Models for Robot Imagination
RoboDreamer factorizes video generation using language primitives to achieve compositional generalization in robot world models, outperforming monolithic baselines on unseen goals in RT-X.
-
Learning Interactive Real-World Simulators
UniSim learns a universal real-world simulator from orchestrated diverse datasets, enabling zero-shot deployment of policies trained purely in simulation.
-
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
A new shared video-image tokenizer enables large language models to surpass diffusion models on standard visual generation benchmarks.
-
Generative Semantic Communication: Diffusion Models Beyond Bit Recovery
A generative semantic communication system that sends compressed semantic information and uses diffusion models with spatially-adaptive normalizations to reconstruct high-quality, semantically consistent images even u...
-
Imagen Video: High Definition Video Generation with Diffusion Models
Imagen Video generates high-definition text-conditional videos via a cascade of base and super-resolution diffusion models, achieving high fidelity and controllability.
-
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
A hybrid NoPE visual diffusion backbone with HeteroP scaling yields ~7× pretraining compute efficiency versus matched full attention and near-stable zero-shot 5s→30s video extrapolation.
-
TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
TPD restores suppressed late-segment events in text-to-video diffusion by projecting classifier-free guidance onto a frame- and timestep-selective lower bound along a temporal-counterfactual direction.
-
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
LiveEdit distills a bidirectional video foundation model into a unidirectional streaming editor via three-stage training plus mask caching to reach 12.66 FPS with stable edits.
-
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
Ultra-low-bitrate image decoding is cast as one-step next-frame prediction from a compact anchor using adapted video diffusion priors, yielding large perceptual bitrate savings versus DiffC.
-
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
SafeCA reduces text-to-video jailbreak success by roughly 20% relative to T2VShield by masking anomalous cross-attention activations using clean-prompt statistics.
-
Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking
Uni4R couples optimal-transport motion priors with ODE integration to predict kinematically coherent point trajectories and reconstructions at arbitrary video timestamps, achieving reported SOTA on four tracking and t...
-
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
A frozen video forgery detector can detect AI-generated videos using only 211 selected anchor neurons and a linear classifier.
-
CoT-Edit: Let CoT Guide Instruction Video Editing
CoT-Edit achieves state-of-the-art instruction-based video editing by generating bounding boxes and enriched instructions with a CoT-enhanced multimodal planner, which guide mask-based diffusion editing.
-
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
TARS splits videos into clip pairs for self-supervised camera learning, adds text-driven viewpoint labels, and restricts scarce paired training to high-noise timesteps.
Reference graph
Works this paper leans on
-
[2]
Language Models are Few-Shot Learners
URL https://arxiv.org/abs/2005.14165. Franc ¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1251–1258,
work page Pith review arXiv 2005
-
[3]
arXiv preprint arXiv:2204.14217 , eprint =
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. arXiv preprint arXiv:2204.14217,
-
[4]
Make-a-scene: Scene-based text-to-image generation with human priors.ArXiv, abs/2203.13131, 2022
URLhttps://arxiv. org/abs/2203.13131. Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. ECCV,
-
[5]
doi: 10.1109/CVPRW50498.2020.00193. Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. NIPS,
-
[6]
Denoising Diffusion Probabilistic Models
URL https://arxiv.org/abs/2006.11239. Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models,
work page Pith review arXiv 2006
-
[7]
URL https://arxiv.org/abs/2204.03458. Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee. Inferring semantic layout for hierarchical text-to-image synthesis. In CVPR, pp. 7986–7994,
-
[8]
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
URL https://arxiv.org/ abs/2205.15868. Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin. Video generation from text. In AAAI, volume 32,
-
[9]
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019a. URL http://arxiv.org/abs/1907.11692. Yue Liu, Xin Wang, Yitian Yuan, and Wenwu Zhu. Cross-modal dual learning for sentence-to- video ge...
work page Pith review arXiv 1907
Show all 15 references
-
[10]
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741,
-
[11]
Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless
URL https://arxiv.org/abs/ 2204.06125. Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless. Film: Frame interpolation for large motion. arXiv preprint arXiv:2202.04901,
-
[12]
Masaki Saito, Shunta Saito, Masanori Koyama, and Sosuke Kobayashi
URL https://arxiv.org/abs/ 2205.11487. Masaki Saito, Shunta Saito, Masanori Koyama, and Sosuke Kobayashi. Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan. International Journal of Computer Vision, 128(10):2586–2606,
-
[13]
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402,
-
[14]
org/abs/1706.03762
URL https://arxiv. org/abs/1706.03762. Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021a. Chenfei Wu, Jian Liang, Lei Ji, Fa...
-
[15]
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627,
-
[16]
Scaling autoregressive models for content-rich text-to-image generation, 2022a
12 Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-...
Reviewed May 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.