REVIEW 3 major objections 2 minor 1 cited by
HumanGenesis proposes a closed-loop, four-agent pipeline that turns monocular video of a person into a 3D-consistent representation, enabling photorealistic, intention-driven human motion synthesis for text prompts, reenactment, and novel p
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in this review.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection The submission pairs a HumanGenesis abstract with the full text of a different paper (OneVAE), so none of the claimed state-of-the-art results are backed by methods, equations, or experiments. the 3 major comments →
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
HumanGenesis claims that the two persistent failures of synthetic human dynamics—geometric inconsistency and coarse reconstruction, plus poor motion generalization and scene inharmonization—can be solved by treating reconstruction and generation as one closed loop. The framework's contribution is a four-agent architecture: a Reconstructor that turns monocular video into a 3D-consistent human-scene representation using 3D Gaussian Splatting with deformation decomposition; a Critique Agent that refines poor regions over multiple rounds of multimodal-LLM reflection; a Pose Guider that generates expressive pose sequences with time-aware parametric encoders; and a Video Harmonizer that renders th
What carries the argument
The load-bearing mechanism is the closed-loop collaboration of four agents, anchored by the Reconstructor's 3D Gaussian Splatting representation with deformation decomposition, and closed by the Video Harmonizer's Back-to-4D feedback into that representation. The Critique Agent supplies the loop's error signal: multi-round reflection by a multimodal large language model identifies and refines poor regions, while the Pose Guider provides motion generalization through time-aware parametric encoders. The central object carrying the argument is the 3D Gaussian Splatting human-scene representation, because both critique and feedback operate on it; the generated video is only as consistent as that
Load-bearing premise
The whole framework depends on being able to recover a genuinely 3D-consistent human-scene representation from a single monocular video; if that reconstruction is coarse or wrong, the critique, pose, and video agents all inherit the error and the claimed geometric fidelity collapses.
What would settle it
Record a monocular video of a person against a calibrated 3D scene with known ground truth, run HumanGenesis, and measure the Reconstructor's geometric error against the scan; if the error is large, or if disabling the Critique Agent produces no measurable drop in reconstruction or video quality, the four-agent claim is unsupported.
If this is right
- Text-guided and reenactment video of people could become geometrically consistent, not just photorealistic in 2D.
- A single monocular clip could drive a 3D-consistent representation that supports novel poses, enabling controllable character animation from casual video.
- Generative video becomes self-correcting: each synthesis round can improve the underlying reconstruction through the Back-to-4D loop.
- Multi-round multimodal-LLM critique offers a general way to use reasoning models to polish 3D reconstructions, not just final images.
Where Pith is reading between the lines
- Editorial inference: if the pipeline delivers on its claim, the boundary between reconstruction and generation blurs—generative video becomes a tool for 3D scene understanding, and 3D reconstruction becomes a generator's internal representation.
- Editorial inference: the Critique Agent suggests a general recipe: use a vision-language model as a feedback loop for geometric refinement in other ill-posed 3D tasks, such as single-image or single-video scene reconstruction.
- Editorial inference: the focus on intention-driven motion implies the framework could eventually take high-level behavioral descriptions and produce geometrically consistent video, a step toward language-driven character animation.
- Editorial inference: a direct testable extension is to feed the same monocular video with deliberately incorrect critique prompts; if output quality drops measurably, the Critique Agent is load-bearing rather than decorative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript, as submitted, consists of an abstract for a system called HumanGenesis, described as a four-agent framework (Reconstructor, Critique Agent, Pose Guider, Video Harmonizer) for photorealistic synthetic human dynamics from monocular video, and a full text that is titled 'OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better' by Yupeng Zhou et al., with header arXiv:2508.09857v1. The full text contains no mention of HumanGenesis, none of the four agents, and no methods, equations, experiments, or comparisons for the claimed tasks (text-guided synthesis, video reenactment, novel-pose generalization). The abstract's final sentence claims state-of-the-art performance on these tasks, but the submitted body provides no supporting evidence whatsoever for that claim. The paper must be assessed on this mismatch: the central claims of HumanGenesis are entirely unsupported by the manuscript's actual content.
Significance. If the HumanGenesis claims were substantiated by a corresponding methods and experiments section, the proposed integration of 3D Gaussian Splatting-based reconstruction with MLLM-based critique and hybrid diffusion rendering could be of interest to the human-centric video generation community. However, as submitted, the manuscript provides no verifiable technical content for HumanGenesis. There are no architectural details, no equations, no training procedures, no benchmark results, and no ablations. The only experimental content in the full text belongs to an unrelated video-VAE paper (OneVAE) and cannot be used as evidence for HumanGenesis. Thus the significance of the submitted manuscript is currently nil: it is impossible to evaluate correctness, novelty, or usefulness of the claimed framework.
major comments (3)
- [Abstract, final sentence] The central claim—'HumanGenesis achieves state-of-the-art performance on tasks including text-guided synthesis, video reenactment, and novel-pose generalization'—is unsupported by the submitted full text. The full text is titled 'OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better' with arXiv:2508.09857v1, and it contains no occurrence of 'HumanGenesis', none of the four agents (Reconstructor, Critique Agent, Pose Guider, Video Harmonizer), and no experiments on the claimed tasks. This is an absence-of-evidence problem at the load-bearing level: the manuscript's only quantitative results (Table 1, Panda70M reconstruction metrics) pertain to OneVAE and are irrelevant to HumanGenesis's text-guided synthesis, video reenactment, or novel-pose generalization claims.
- [Full text (entirety)] The full text provides no methods for HumanGenesis. In particular, the Reconstructor's 3D Gaussian Splatting with deformation decomposition, the Critique Agent's multi-round MLLM reflection, the Pose Guider's time-aware parametric encoders, and the Video Harmonizer's Back-to-4D feedback loop are described only in the abstract. No equations, architecture diagrams, or algorithm pseudocode are given for any of these components. Consequently, the reader cannot audit the proposed pipeline, reproduce it, or assess the stated geometric fidelity and scene-integration improvements. The manuscript is not self-contained even at the level of a technical abstract; it lacks the essential methodological content expected of a paper making such performance claims.
- [Full text, §4 (Quantitative Comparison)] Even if one treated the full text as part of the submission, the experiments in §4 evaluate video autoencoder reconstruction quality (PSNR, SSIM, LPIPS, FVD) against tokenization baselines. None of these experiments involve text-guided synthesis, video reenactment, or novel-pose generalization. Therefore the full text cannot substantiate the abstract's state-of-the-art assertion. The mismatch is not a minor presentation issue but a fundamental disconnect between the claimed contribution and the submitted manuscript content.
minor comments (2)
- [Header, full text] The full text header indicates arXiv:2508.09857v1, which does not match the arXiv identifier of the submitted manuscript (2508.09858). This discrepancy should be checked with the authors/submission system, as it likely indicates a file-upload or packaging error.
- [Throughout] The full text contains inconsistent spacing in the model name (e.g., 'OneV AE' in the title versus 'OneVAE' elsewhere). This is a typographical issue in the submitted body, though it is secondary to the major content mismatch.
Circularity Check
No circular derivation identifiable: the HumanGenesis abstract contains no equations, fits, or self-citations, and the supplied full text is a different paper (OneVAE), so there is no derivation chain to audit.
full rationale
The claimed subject is HumanGenesis (arXiv:2508.09858), whose abstract describes a four-agent pipeline and states state-of-the-art performance, but provides no equations, ablations, or benchmark numbers. The supplied full text is titled 'OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better' (arXiv:2508.09857v1) and concerns video tokenizers; it never mentions HumanGenesis, the four agents, or the claimed tasks. Therefore there is no concrete derivation chain in which an output is equal to an input by construction, no fitted parameter renamed as a prediction, and no load-bearing self-citation that can be quoted. The absence of supporting methods and experiments makes the abstract's SOTA claim unverifiable, but unverifiable is not circular. Under the hard rule that circularity must be evidenced by a specific reduction or self-citation chain, no circular step can be identified. Score 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Monocular video contains sufficient information to build a 3D-consistent human-scene representation via 3D Gaussian Splatting with deformation decomposition
- domain assumption Multi-round MLLM-based reflection improves reconstruction fidelity
- domain assumption Diffusion-based hybrid rendering can produce photorealistic video consistent with the reconstructed geometry
Cite this review
Pith. "Pith review of HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics." pith.science (2026). https://pith.science/paper/GTLMDQNF
@misc{pith2026250809858,
author = {Pith},
title = {Pith review of: HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics},
year = {2026},
howpublished = {\url{https://pith.science/paper/GTLMDQNF}},
note = {Machine review of arXiv:2508.09858}
}
read the original abstract
\textbf{Synthetic human dynamics} aims to generate photorealistic videos of human subjects performing expressive, intention-driven motions. However, current approaches face two core challenges: (1) \emph{geometric inconsistency} and \emph{coarse reconstruction}, due to limited 3D modeling and detail preservation; and (2) \emph{motion generalization limitations} and \emph{scene inharmonization}, stemming from weak generative capabilities. To address these, we present \textbf{HumanGenesis}, a framework that integrates geometric and generative modeling through four collaborative agents: (1) \textbf{Reconstructor} builds 3D-consistent human-scene representations from monocular video using 3D Gaussian Splatting and deformation decomposition. (2) \textbf{Critique Agent} enhances reconstruction fidelity by identifying and refining poor regions via multi-round MLLM-based reflection. (3) \textbf{Pose Guider} enables motion generalization by generating expressive pose sequences using time-aware parametric encoders. (4) \textbf{Video Harmonizer} synthesizes photorealistic, coherent video via a hybrid rendering pipeline with diffusion, refining the Reconstructor through a Back-to-4D feedback loop. HumanGenesis achieves state-of-the-art performance on tasks including text-guided synthesis, video reenactment, and novel-pose generalization, significantly improving expressiveness, geometric fidelity, and scene integration.
Forward citations
Cited by 1 Pith paper
-
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.
Reference graph
Works this paper leans on
-
[1]
Lumiere: A space-time diffusion model for video generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, et al. Lumiere: A space-time diffusion model for video generation. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 4
work page 2024
-
[2]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023. 3
work page 2023
-
[3]
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018. 3
Pith/arXiv arXiv 2018
-
[4]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. 4
work page 2024
-
[5]
Free-form video inpainting with 3d gated convolution and temporal patchgan
Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, and Winston Hsu. Free-form video inpainting with 3d gated convolution and temporal patchgan. In Proceedings of the IEEE/CVF international conference on computer vision, 2019. 7
work page 2019
-
[6]
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 2, 3
Pith/arXiv arXiv 2023
-
[7]
Od-vae: An omni-dimensional video compressor for improving latent video diffusion model
Liuhan Chen, Zongjian Li, Bin Lin, Bin Zhu, Qian Wang, Shenghai Yuan, Xing Zhou, Xinhua Cheng, and Li Yuan. Od-vae: An omni-dimensional video compressor for improving latent video diffusion model. arXiv preprint arXiv:2409.01199, 2024. 3
Pith/arXiv arXiv 2024
-
[8]
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13320–13331, 2024. 7
work page 2024
-
[9]
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811, 2025. 2 10
Pith/arXiv arXiv 2025
-
[10]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024. 2, 3
work page 2024
-
[11]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 2, 3, 4, 5
work page 2021
-
[12]
Long video generation with time-agnostic vqgan and time-sensitive transformer
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. InEuropean Conference on Computer Vision, pages 102–118. Springer, 2022. 3
work page 2022
-
[13]
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 3
2014
-
[14]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
work page 2020
-
[15]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. Advances in Neural Information Processing Systems, 35:8633–8646, 2022. 4
work page 2022
-
[16]
Dive: Dit-based video generation with enhanced control
Junpeng Jiang, Gangyi Hong, Lijun Zhou, Enhui Ma, Hengtong Hu, Xia Zhou, Jie Xiang, Fan Liu, Kaicheng Yu, Haiyang Sun, et al. Dive: Dit-based video generation with enhanced control. arXiv preprint arXiv:2409.01595, 2024. 4
Pith/arXiv arXiv 2024
-
[17]
Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Yang Jin, Zhicheng Sun, Kun Xu, Liwei Chen, Hao Jiang, Quzhe Huang, Chengru Song, Yuliang Liu, Di Zhang, Yang Song, et al. Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization. arXiv preprint arXiv:2402.03161, 2024. 3
arXiv 2024
-
[18]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410, 2019. 3
work page 2019
-
[19]
Auto-encoding variational bayes, 2013
Diederik P Kingma, Max Welling, et al. Auto-encoding variational bayes, 2013. 4
work page 2013
-
[20]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603,
-
[21]
PKU-Yuan Lab and Tuzhan AI etc. Open-sora-plan, April 2024. 4
work page 2024
-
[22]
Efficient spatially sparse inference for conditional gans and diffusion models
Muyang Li, Ji Lin, Chenlin Meng, Stefano Ermon, Song Han, and Jun-Yan Zhu. Efficient spatially sparse inference for conditional gans and diffusion models. Advances in neural information processing systems, 35:28858–28873, 2022. 3
work page 2022
-
[23]
Autoregressive image generation without vector quantization
Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. Autoregressive image generation without vector quantization. Advances in Neural Information Processing Systems, 37:56424–56445, 2024. 2
work page 2024
-
[24]
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model. arXiv preprint arXiv:2411.17459, 2024. 4
Pith/arXiv arXiv 2024
-
[25]
Open-sora plan: Open-source large video generation model
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131,
-
[26]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 7
Pith/arXiv arXiv 2017
-
[27]
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, et al. Step-video-t2v technical report: The practice, challenges, and future of video foundation model. arXiv preprint arXiv:2502.10248, 2025. 2
Pith/arXiv arXiv 2025
-
[28]
Finite scalar quantization: Vq-vae made simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505, 2023. 2, 3, 4, 7
Pith/arXiv arXiv 2023
-
[29]
Cosmos tokenizer: A suite of image and video neural tokenizers, November 2024
NVIDIA. Cosmos tokenizer: A suite of image and video neural tokenizers, November 2024. 1, 2, 3, 4, 6, 8
work page 2024
-
[30]
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rom- bach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952,
-
[31]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 3
work page 2018
-
[32]
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 3
Pith/arXiv arXiv 2022
-
[33]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821–8831. Pmlr, 2021. 3 11
work page 2021
-
[34]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 3, 4
work page 2022
-
[35]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35:36479–36494, 2022. 3
2022
-
[36]
Taming scalable visual tokenizer for autoregressive image generation
Fengyuan Shi, Zhuoyan Luo, Yixiao Ge, Yujiu Yang, Ying Shan, and Limin Wang. Taming scalable visual tokenizer for autoregressive image generation. arXiv preprint arXiv:2412.02692, 2024. 3, 4, 7
Pith/arXiv arXiv 2024
-
[37]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792,
-
[38]
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[39]
Vidtok: A versatile and open-source video tokenizer
Anni Tang, Tianyu He, Junliang Guo, Xinle Cheng, Li Song, and Jiang Bian. Vidtok: A versatile and open-source video tokenizer. arXiv preprint arXiv:2412.13061, 2024. 3, 6
Pith/arXiv arXiv 2024
-
[40]
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024. 2
Pith/arXiv arXiv 2024
-
[41]
Fvd: A new metric for video generation
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. 2019. 7
work page 2019
-
[42]
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. 3
work page 2017
-
[43]
Wan: Open and advanced large-scale video generative models
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 2, 5
Pith/arXiv arXiv 2025
-
[44]
Modelscope text-to-video technical report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 3
Pith/arXiv arXiv 2023
-
[45]
Omnitokenizer: A joint image- video tokenizer for visual generation
Junke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Zuxuan Wu, and Yu-Gang Jiang. Omnitokenizer: A joint image- video tokenizer for visual generation. Advances in Neural Information Processing Systems, 37:28281–28295, 2024. 2, 3, 8, 9
work page 2024
-
[46]
Emu3: Next-token prediction is all you need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024. 1, 2, 8, 9
Pith/arXiv arXiv 2024
-
[47]
Loong: Generating minute-level long videos with autoregressive language models
Yuqing Wang, Tianwei Xiong, Daquan Zhou, Zhijie Lin, Yang Zhao, Bingyi Kang, Jiashi Feng, and Xihui Liu. Loong: Generating minute-level long videos with autoregressive language models. arXiv preprint arXiv:2410.02757, 2024. 6
Pith/arXiv arXiv 2024
-
[48]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 2004. 7
work page 2004
-
[49]
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7623–7633, 2023. 4
work page 2023
-
[50]
Elastictok: Adaptive tokenization for image and video
Wilson Yan, Matei Zaharia, V olodymyr Mnih, Pieter Abbeel, Aleksandra Faust, and Hao Liu. Elastictok: Adaptive tokenization for image and video. arXiv preprint arXiv:2410.08368, 2024. 3
Pith/arXiv arXiv 2024
-
[51]
Videogpt: Video generation using vq-vae and transformers
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021. 2, 3
Pith/arXiv arXiv 2021
-
[52]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2, 3, 4
Pith/arXiv arXiv 2024
-
[53]
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 3
Pith/arXiv arXiv 2022
-
[54]
Magvit: Masked generative video transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, et al. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10459–10469, 2023. 3
work page 2023
-
[55]
Language model beats diffusion–tokenizer is key to visual generation
Lijun Yu, José Lezama, Nitesh B Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023. 1, 3 12
Pith/arXiv arXiv 2023
-
[56]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 7
work page 2018
-
[57]
Cv-vae: A compatible video vae for latent generative video models
Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv-vae: A compatible video vae for latent generative video models. arXiv preprint arXiv:2405.20279, 2024. 1, 3, 4, 7, 8, 9
Pith/arXiv arXiv 2024
-
[58]
Chuanxia Zheng and Andrea Vedaldi. Online clustered codebook. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22798–22807, 2023. 3
work page 2023
-
[59]
Open-sora: Democratizing efficient video production for all, March 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, March 2024. 4, 5, 8, 9
work page 2024
-
[60]
Addressing representation collapse in vector quantized models with one linear layer
Yongxin Zhu, Bocheng Li, Yifei Xin, and Linli Xu. Addressing representation collapse in vector quantized models with one linear layer. arXiv preprint arXiv:2411.02038, 2024. 7 13
arXiv 2024
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.