REVIEW 3 major objections 6 minor 2 cited by
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Sliding iterative denoising, alternating across viewpoints and time, produces 4D-consistent novel-view human videos from sparse-view inputs.
desk verdict Sliding iterative denoising is a genuinely useful idea, but the missing timestep schedule for overlapping windows makes the core algorithm unreproducible as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The sliding iterative denoising process is the central mechanism: a context window of length $W$ slides over the $(N+M) \times T$ latent grid with stride $S$, alternately along the spatial (multi-view) and temporal (video) dimensions, performing a few denoising steps per window position and reversing direction to enable bidirectional context aggregation. This is what extends the diffusion model's receptive field without exceeding GPU memory. The second mechanism is skeleton-Plücker mixed conditioning, where 3D skeletons projected into each view are encoded into the latent space and concatenated with Plücker coordinates, giving pixel-aligned human pose and camera control that resolve front-back ambiguity and improve pose accuracy.
What would settle it
Regenerate a test sequence with per-sample noise levels tracked exactly (each latent denoised according to its true step count) and compare the output consistency against the paper's shared-timestep sliding schedule; if the sliding variant does not produce fewer inter-window inconsistencies than independent-window denoising with the same total compute, the claimed large-receptive-field benefit is not supported.
Extended reading notes
Core claim
The central claim is that consistency in long-sequence multi-view video generation can be achieved by treating all target views and frames as a single 4D latent grid and denoising it with a sliding window that iterates over the spatial and temporal dimensions. Each sample in the grid is a latent encoding the image, camera pose, and human pose for a given viewpoint and timestamp. The window slides counter-clockwise then clockwise in the spatial dimension and similarly in the temporal dimension, and each sample receives a total of $D = 2 \times P \times W / S$ denoising steps, equal to the diffusion inference steps. The authors argue this gives the model a large receptive field, lets nearby samples share more joint denoising steps (matching the correlation structure of 4D data), and keeps GPU memory bounded by the window size. They further condition generation on projected 3D skeletons plus Plücker coordinates to constrain pose, and feed the resulting dense-view videos into a 4D Gaussian Splatting reconstruction pipeline. Experiments show consistent improvements over the compared baselines at both 4-view and 8-view settings.
Load-bearing premise
The load-bearing premise is that during sliding, latents that have undergone different numbers of denoising passes can be fed together with a single shared timestep, and the resulting schedule still behaves like a well-defined diffusion sampling process.
Editorial extensions
If this is right
- The method enables 1024p, 4D-consistent novel-view human video synthesis from as few as 4 input views, with visual quality reported to be comparable to dense 48-view reconstruction.
- Alternating spatial and temporal denoising with sliding windows introduces a smoothness inductive bias that removes the group-boundary jumps seen in multi-group and median-filtering denoising strategies.
- Skeleton conditioning combined with Plücker coordinates reduces camera misalignment and front-back ambiguity, improving the accuracy of generated human poses.
- The generated dense-view videos can be fed into an off-the-shelf 4D Gaussian Splatting pipeline, yielding a real-time renderable free-viewpoint video representation.
- The processed DNA-Rendering dataset with recalibrated cameras, color correction, foreground masks, and estimated skeletons is planned for release, which could support future sparse-view human reconstruction research.
Reading between the lines
- The sliding schedule implicitly assumes that all latents in a window can be denoised with a single shared timestep even though samples have undergone different numbers of denoising passes across sliding iterations; a reader could test whether this assumption holds and whether it affects generation quality.
- The sliding iterative denoising mechanism is not human-specific and could transfer to other multi-view video diffusion tasks, such as general dynamic scenes, where long-sequence consistency is also a bottleneck.
- The paper sets window length and stride to fixed values ($W=3$, $S=1$) without ablating them; the trade-off between consistency and inference time across these hyperparameters is a testable extension.
- Releasing the processed DNA-Rendering dataset could standardize sparse-view human evaluation, since the reproduced baseline CAT4D was retrained on the same processed data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Diffuman4D proposes a two-stage system for free-viewpoint video synthesis of humans from sparse-view video. In the first stage, a latent diffusion model, initialized from Stable Diffusion 2.1 and fine-tuned on a processed DNA-Rendering subset, generates dense multi-view videos. Input views, skeleton renderings, Plücker coordinates, and a conditional mask are concatenated in latent space. The central novelty is a sliding iterative denoising schedule (Sec. 3.2) in which a window of length W slides with stride S over a latent grid, applying P denoising steps per window, first along the spatial dimension and then along the temporal dimension; the paper claims each latent receives D = 2PW/S denoising steps, yielding a large receptive field at bounded GPU cost. In the second stage, the synthesized videos are used to reconstruct a 4D Gaussian Splatting representation with LongVolcap. Experiments on DNA-Rendering and ActorsHQ compare against LongVolcap, GauHuman, GPS-Gaussian, and a reproduced CAT4D, with ablations over the denoising strategy and conditioning scheme.
Significance. The contribution is potentially significant: if the sliding schedule is well-defined, it offers a practical way to extend multi-view video diffusion to long sequences while keeping memory affordable, and the paper's integration with a 4DGS reconstruction backend is sensible. The paper has clear strengths: evaluation on two public datasets, ablations for both main components, a reproduced CAT4D baseline trained under the same settings, and a plan to release a processed DNA-Rendering dataset. The reported numbers, if reproducible, are substantially better than the baselines. However, the central sampling schedule is under-specified in a way that affects the validity of the method as a diffusion process, and the empirical support would be stronger with error bars and a direct consistency metric.
major comments (3)
- [Sec. 3.2] The sliding iterative denoising schedule is not a well-defined diffusion sampling procedure as written. With overlapping windows (W > S), a new window contains latents that have undergone different numbers of prior denoising updates (e.g., with W=3, S=1, samples 2 and 3 have one more update than sample 4 when the window moves), yet the paper never specifies the diffusion timestep or noise level assigned to each latent in each network call. The formula D = 2PW/S counts network applications, but it cannot be mapped to a standard DDIM/DDPM or DPM-Solver++ schedule unless per-sample timestep bookkeeping is provided; the supplementary's statement that DPM-Solver++ with 24 steps is used does not resolve this because W, S, and P are not reported. Please provide exact pseudocode for the schedule, including per-window/per-sample timestep assignment, boundary handling at sequence ends, and the integer constraints on D/2 and W/S.
- [Sec. 4.4 / Table 2] The quantitative evidence for the central claim is incomplete: the denoising-strategy ablation reports only PSNR, SSIM, and LPIPS, which are per-frame image-similarity metrics and do not directly measure spatio-temporal consistency. Since the paper's core claim is that sliding iterative denoising enhances 4D consistency, the evaluation should include a direct consistency metric, such as temporal flicker, cross-view re-projection error, or a warping-based metric, rather than relying on per-frame averages and qualitative figures.
- [Tables 1–3] No error bars or significance tests are reported for any of the quantitative comparisons, although the method is stochastic (diffusion sampling with classifier-free guidance). The claim that Diffuman4D 'significantly outperforms' the baselines would be more defensible with means and variances over multiple independent sampling runs and a statement of the number of trials; this applies especially to the small-scale ablations in Tables 2 and 3.
minor comments (6)
- [Table 1] The table title contains a typo: 'DNA-Rednering' should be 'DNA-Rendering'.
- [Fig. 6 caption] The caption says that 'GPS-Gaussian uses 8 input views while all other methods use 4 input views,' which is inconsistent with Table 1 where GPS-Gaussian is evaluated in both 4-view and 8-view settings; please clarify the input-view counts for each method in each panel.
- [Supplementary Sec. A] There are formatting errors such as 'with24 sampling steps' (missing space) and '4D reconstrcution' (misspelled); similar spacing/encoding issues affect 'V AE' in Sec. 3.1 and 'Pl ¨ucker' throughout.
- [Sec. 3.2 and Fig. 3b] The notation is inconsistent: the text uses N for target views and T for frames, while Fig. 3b says M=2, N=5 for an 'M-view, N-frame video'; please unify the notation for the number of views and frames.
- [Sec. 3.2] The formula for total denoising steps is typeset differently as 'D = 2×P ×W/S' in the text and 'D = 2P W/S' in Fig. 3; please use one consistent form and state the required divisibility conditions (e.g., W/S and D/2 integer).
- [Sec. 4.2] Please clarify whether the reproduced CAT4D† uses its original sliding-window strategy or an adapted version of the proposed alternating spatial/temporal procedure, since the current wording says it uses 'the same sampling sequences and conditional-view selection strategy described in Sec. 3.2' but not the same denoising schedule.
Circularity Check
No significant circularity: the derivation chain is self-contained and evaluated against external baselines.
full rationale
I found no step where an output is defined in terms of an input, no fitted parameter later reported as a prediction, and no load-bearing appeal to a self-citation. The sliding iterative denoising (Sec. 3.2) defines D = 2*P*W/S as a counting identity for how many P-step windows cover each sample; that is a definition used to set the number of sampling steps, not a derivation of the result from the result. The ablation in Table 2 compares sliding iterative against multi-group and median filtering on the same data, and all evaluations use external baselines (CAT4D, GauHuman, GPS-Gaussian, LongVolcap) on public datasets. LongVolcap is an author-provided reconstruction tool used as a shared pipeline stage, not a parameter fitted to the target; using it does not make the novel-view comparison circular. The unspecified per-sample timestep bookkeeping flagged in Sec. 3.2 is a potential correctness/well-definedness issue in the sampling procedure, but it is not an instance of a claim reducing to its own inputs. No circularity score above 0 is warranted.
Assumptions & free parameters
free parameters (4)
- Window length W and stride S
- Denoising steps per window P
- Classifier-free guidance scale =
3.0
- DPM-Solver++ sampling steps =
24
assumptions (4)
- domain assumption Pretrained Stable Diffusion 2.1 and its autoencoder provide a suitable generative prior for multi-view human video.
- domain assumption 3D skeletons estimated by Sapiens and triangulation from sparse views are accurate enough to condition generation.
- ad hoc to paper All latents in a sliding window share a single timestep, or the model can handle mixed noise levels within a batch.
- domain assumption The DNA-Rendering test sequences used for evaluation are representative and not identity-overlapping with the 1,000 training sequences.
Cite this review
Pith. "Pith review of Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models." pith.science (2026). https://pith.science/paper/POGBTCKR
@misc{pith2026250713344,
author = {Pith},
title = {Pith review of: Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/POGBTCKR}},
note = {Machine review of arXiv:2507.13344}
}
read the original abstract
This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel viewpoints. However, the generated videos from these models often lack spatio-temporal consistency, thus degrading view synthesis quality. In this paper, we propose a novel sliding iterative denoising process to enhance the spatio-temporal consistency of the 4D diffusion model. Specifically, we define a latent grid in which each latent encodes the image, camera pose, and human pose for a certain viewpoint and timestamp, then alternately denoising the latent grid along spatial and temporal dimensions with a sliding window, and finally decode the videos at target viewpoints from the corresponding denoised latents. Through the iterative sliding, information flows sufficiently across the latent grid, allowing the diffusion model to obtain a large receptive field and thus enhance the 4D consistency of the output, while making the GPU memory consumption affordable. The experiments on the DNA-Rendering and ActorsHQ datasets demonstrate that our method is able to synthesize high-quality and consistent novel-view videos and significantly outperforms the existing approaches. See our project page for interactive demos and video results: https://diffuman4d.github.io/ .
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras
SparseCam4D achieves spatio-temporally consistent high-fidelity 4D reconstruction from sparse cameras via a Spatio-Temporal Distortion Field that corrects inconsistencies in generative observations.
-
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
Splatography improves dynamic 3D reconstruction from sparse multi-view videos by splitting foreground and background Gaussian representations and applying tailored deformation learning for each.
Reference graph
Works this paper leans on
-
[1]
Easymocap - make human motion capture easier. Github,
-
[2]
Creation of 3d human avatar using kinect
Kairat Aitpayev and Jaafar Gaber. Creation of 3d human avatar using kinect. Asian Transactions on Fundamentals of Electronics, Communication & Multimedia , 1(5):12–24,
-
[3]
Tc4d: Trajectory-conditioned text-to-4d generation
Sherwin Bahmani, Xian Liu, Wang Yifan, Ivan Sko- rokhodov, Victor Rong, Ziwei Liu, Xihui Liu, Jeong Joon Park, Sergey Tulyakov, Gordon Wetzstein, et al. Tc4d: Trajectory-conditioned text-to-4d generation. In European Conference on Computer Vision , pages 53–72. Springer,
-
[4]
4d-fy: Text-to-4d generation using hybrid score dis- tillation sampling
Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gordon Wetzstein, Leonidas Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, and David B Lin- dell. 4d-fy: Text-to-4d generation using hybrid score dis- tillation sampling. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 7996–8006, 2024. 3
2024
-
[5]
Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences
Federica Bogo, Michael J Black, Matthew Loper, and Javier Romero. Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences. In Proceedings of the IEEE international conference on computer vision , pages 2300–2308, 2015. 2
2015
-
[6]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, et al. Video generation models as world simulators. 2024. URL https://openai. com/research/video- generation-models-as-world-simulators, 3:1, 2024. 3
2024
-
[7]
Hexplane: A fast representa- tion for dynamic scenes
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 130–141, 2023. 2
2023
-
[8]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In CVPR, 2024. 3
work page 2024
Show all 86 references
-
[9]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024. 3
2024 arXiv
-
[10]
Dna- rendering: A diverse neural actor repository for high-fidelity human-centric rendering
Wei Cheng, Ruixiang Chen, Wanqi Yin, Siming Fan, Keyu Chen, Honglin He, Huiwen Luo, Zhongang Cai, Jingbo Wang, Yang Gao, Zhengming Yu, Zhengyu Lin, Daxuan Ren, Lei Yang, Ziwei Liu, Chen Change Loy, Chen Qian, Wayne Wu, Dahua Lin, Bo Dai, and Kwan-Yee Lin. Dna- rendering: A div...
-
[11]
High-quality streamable free-viewpoint video
Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (ToG) , 34(4):1–13,
-
[12]
Obja- verse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Obja- verse: A universe of annotated 3d objects. arXiv preprint arXiv:2212.08051, 2022. 2
2022 arXiv
-
[13]
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Chris- tian Laforte, Vikram V oleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl V ondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Obja...
2023 arXiv
-
[14]
4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes
Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wen- zheng Chen, and Baoquan Chen. 4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes. In ACM SIGGRAPH 2024 Conference Papers , pages 1–11,
2024
-
[15]
Fast dynamic radiance fields with time-aware neural vox- els
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xi- aopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. Fast dynamic radiance fields with time-aware neural vox- els. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9,
2022
-
[16]
K-planes: Explicit radiance fields in space, time, and appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 12479–12488, 2023. 2
2023
-
[17]
Massively parallel multiview stereopsis by surface normal diffusion
Silvano Galliani, Katrin Lasinger, and Konrad Schindler. Massively parallel multiview stereopsis by surface normal diffusion. 2015. 2
2015
-
[18]
Srinivasan, Jonathan T
Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srinivasan, Jonathan T. Barron, and Ben Poole. Cat3d: Create anything in 3d with multi-view diffusion models. In NeurIPS 2024,
2024
-
[19]
Studio production system for dynamic 3d con- tent
Oliver Grau. Studio production system for dynamic 3d con- tent. In Visual Communications and Image Processing 2003, pages 80–89. SPIE, 2003. 2
2003
-
[20]
Viewdiff: 3d-consistent image generation with text-to-image models
Lukas H ¨ollein, Aljaˇz Boˇziˇc, Norman M¨uller, David Novotny, Hung-Yu Tseng, Christian Richardt, Michael Zollh ¨ofer, and Matthias Nießner. Viewdiff: 3d-consistent image generation with text-to-image models. In Proceedings of the IEEE/CVF conference on computer vision and pa...
2024
-
[21]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 3
2022 arXiv
-
[22]
Gauhuman: Articulated gaus- sian splatting from monocular human videos
Shoukang Hu and Ziwei Liu. Gauhuman: Articulated gaus- sian splatting from monocular human videos. arXiv preprint arXiv:, 2023. 3, 6, 7
2023
-
[23]
Humanrf: High-fidelity neural radiance fields for humans in motion
Mustafa Is ¸ık, Martin R ¨unz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes Agapito, and Matthias Nießner. Humanrf: High-fidelity neural radiance fields for humans in motion. ACM Transactions on Graphics (TOG), 42(4):1–12, 2023. 6, 7, 8, 2
2023
-
[24]
Pl ¨ucker coordinates for lines in the space.Prob- lem Solver Techniques for Applied Computer Science, Com- S-477/577 Course Handout, 3, 2020
Yan-Bin Jia. Pl ¨ucker coordinates for lines in the space.Prob- lem Solver Techniques for Applied Computer Science, Com- S-477/577 Course Handout, 3, 2020. 3
2020
-
[25]
Virtualized reality: Constructing virtual worlds from real scenes
Takeo Kanade, Peter Rander, and PJ Narayanan. Virtualized reality: Constructing virtual worlds from real scenes. IEEE multimedia, 4(1):34–47, 1997. 2
1997
-
[26]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 2, 1
2023
-
[27]
Sapiens: Foundation for human vision mod- els
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vision mod- els. arXiv preprint arXiv:2408.12569, 2024. 5, 1
2024 arXiv
-
[28]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ,
-
[29]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 3
2024 arXiv
-
[30]
A theory of shape by space carving
Kiriakos N Kutulakos and Steven M Seitz. A theory of shape by space carving. International journal of computer vision , 38:199–218, 2000. 1
2000
-
[31]
Vivid-zoo: Multi-view video generation with diffusion model.Advances in Neural Information Processing Systems, 37:62189–62222,
Bing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai, Biao Zhang, Peter Wonka, and Bernard Ghanem. Vivid-zoo: Multi-view video generation with diffusion model.Advances in Neural Information Processing Systems, 37:62189–62222,
-
[32]
Neural 3d video synthesis from multi-view video
Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. In Proceedings of the IEEE/CVF conference on computer vi- si...
2022
-
[33]
Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling
Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3
2024
-
[34]
Efficient neural radiance fields for interactive free-viewpoint video
Haotong Lin, Sida Peng, Zhen Xu, Yunzhi Yan, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Efficient neural radiance fields for interactive free-viewpoint video. In SIGGRAPH Asia Conference Proceedings, 2022. 3
2022
-
[35]
Real-time high-resolution background matting
Shanchuan Lin, Andrey Ryabtsev, Soumyadip Sengupta, Brian L Curless, Steven M Seitz, and Ira Kemelmacher- Shlizerman. Real-time high-resolution background matting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8762–8771, 2021. 1
2021
-
[36]
Raft-stereo: Multilevel recurrent field transforms for stereo matching
Lahav Lipson, Zachary Teed, and Jia Deng. Raft-stereo: Multilevel recurrent field transforms for stereo matching. In 2021 International Conference on 3D Vision (3DV) , pages 218–227. IEEE, 2021. 3
2021
-
[37]
Zero-1-to-3: Zero-shot one image to 3d object, 2023
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023. 3
2023
-
[38]
Mvsgaussian: Fast generalizable gaussian splatting recon- struction from multi-view stereo
Tianqi Liu, Guangcong Wang, Shoukang Hu, Liao Shen, Xinyi Ye, Yuhang Zang, Zhiguo Cao, Wei Li, and Ziwei Liu. Mvsgaussian: Fast generalizable gaussian splatting recon- struction from multi-view stereo. In European Conference on Computer Vision, pages 37–53. Springer, 2024. 3
2024
-
[39]
Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age. arXiv preprint arXiv:2309.03453, 2023. 3
2023 arXiv
-
[40]
Smpl: A skinned multi- person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi- person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023. 2
2023
-
[41]
Dpm-solver: A fast ode solver for diffu- sion probabilistic model sampling in around 10 steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffu- sion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022. 1
2022 arXiv
-
[42]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 2
2021
-
[43]
Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time
Richard A Newcombe, Dieter Fox, and Steven M Seitz. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 343–352,
-
[44]
Effi- cient4d: Fast dynamic 3d object generation from a single- view video
Zijie Pan, Zeyu Yang, Xiatian Zhu, and Li Zhang. Effi- cient4d: Fast dynamic 3d object generation from a single- view video. arXiv preprint arXiv:2401.08742, 2024. 3
2024
-
[45]
Barron, Sofien Bouaziz, Dan B Goldman, Steven M
Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. ICCV, 2021. 3
2021
-
[46]
Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields. ACM Trans. Graph., 40(6), 2021. 3
2021
-
[47]
Ani- matable neural radiance fields for modeling dynamic human bodies
Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Ani- matable neural radiance fields for modeling dynamic human bodies. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 14314–14323, 2021. 3
2021
-
[48]
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. InProceedings of the IEEE/CVF conference on computer vision and ...
2021
-
[49]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 3
2022 arXiv
-
[50]
D-NeRF: Neural Radiance Fields for Dynamic Scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural Radiance Fields for Dynamic Scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2020. 3
2020
-
[51]
Dreamgaussian4d: Genera- tive 4d gaussian splatting
Jiawei Ren, Liang Pan, Jiaxiang Tang, Chi Zhang, Ang Cao, Gang Zeng, and Ziwei Liu. Dreamgaussian4d: Genera- tive 4d gaussian splatting. arXiv preprint arXiv:2312.17142,
-
[52]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022. 1
2022
-
[53]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 2, 1
2016
-
[54]
Pixelwise view selection for un- structured multi-view stereo
Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for un- structured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016. 2, 1
2016
-
[55]
Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering
Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 166...
2023
-
[56]
Rapid avatar capture and simulation using commodity depth sensors
Ari Shapiro, Andrew Feng, Ruizhe Wang, Hao Li, Mark Bo- las, Gerard Medioni, and Evan Suma. Rapid avatar capture and simulation using commodity depth sensors. Computer Animation and Virtual Worlds, 25(3-4):201–211, 2014. 2
2014
-
[57]
Mvdream: Multi-view diffusion for 3d gen- eration
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d gen- eration. arXiv preprint arXiv:2308.16512, 2023. 3
2023 arXiv
-
[58]
Text-to-4d dy- namic scene generation
Uriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual, Iurii Makarov, Filippos Kokkinos, Naman Goyal, Andrea Vedaldi, Devi Parikh, Justin Johnson, et al. Text-to-4d dy- namic scene generation. arXiv preprint arXiv:2301.11280 ,
-
[59]
Virtual view synthesis of people from multiple view video sequences
Jonathan Starck and Adrian Hilton. Virtual view synthesis of people from multiple view video sequences. Graphical Models, 67(6):600–620, 2005. 2
2005
-
[60]
Surface capture for performance-based animation
Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE computer graphics and applications, 27(3):21–31, 2007. 2
2007
-
[61]
Scanning 3d full human bodies using kinects
Jing Tong, Jin Zhou, Ligang Liu, Zhigeng Pan, and Hao Yan. Scanning 3d full human bodies using kinects. IEEE trans- actions on visualization and computer graphics, 18(4):643– 650, 2012. 2
2012
-
[62]
Diffusers: State-of-the-art diffu- sion models
Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. Diffusers: State-of-the-art diffu- sion models. https://github.com/huggingface/ diffusers...
2022
-
[63]
Fourier plenoctrees for dynamic radiance field ren- dering in real-time
Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yan- shun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. Fourier plenoctrees for dynamic radiance field ren- dering in real-time. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognitio...
2022
-
[64]
Imagedream: Image-prompt multi-view diffusion for 3d generation
Peng Wang and Yichun Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation. arXiv preprint arXiv:2312.02201, 2023. 3
2023 arXiv
-
[65]
Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasac- chi, and David J. Fleet. Controlling space and time with dif- fusion models. arXiv preprint arXiv:2407.07860, 2024. 2, 4
2024 arXiv
-
[66]
Hu- mannerf: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern Recognition , pages 162...
2022
-
[67]
4d gaussian splatting for real-time dynamic scene render- ing
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene render- ing. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 2...
2024
-
[68]
Barron, and Aleksander Holynski
Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, and Aleksander Holynski. Cat4d: Create anything in 4d with multi-view video diffu- sion models. arXiv:2411.18613, 2024. 2, 3, 4, 6, 7
2024 arXiv
-
[69]
Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency
Yiming Xie, Chun-Han Yao, Vikram V oleti, Huaizu Jiang, and Varun Jampani. Sv4d: Dynamic 3d content generation with multi-frame and multi-view consistency. arXiv preprint arXiv:2407.17470, 2024. 2, 3
2024 arXiv
-
[70]
Easyvolcap: Accelerating neural volumetric video research
Zhen Xu, Tao Xie, Sida Peng, Haotong Lin, Qing Shuai, Zhiyuan Yu, Guangzhao He, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. Easyvolcap: Accelerating neural volumetric video research. 2023. 5
2023
-
[71]
Relightable and animatable neural avatar from sparse-view video
Zhen Xu, Sida Peng, Chen Geng, Linzhan Mou, Zihan Yan, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. Relightable and animatable neural avatar from sparse-view video. In CVPR, 2024. 3
2024
-
[72]
4k4d: Real-time 4d view synthesis at 4k resolution
Zhen Xu, Sida Peng, Haotong Lin, Guangzhao He, Jiaming Sun, Yujun Shen, Hujun Bao, and Xiaowei Zhou. 4k4d: Real-time 4d view synthesis at 4k resolution. InCVPR, 2024. 2, 1
2024
-
[73]
Representing long volumet- ric video with temporal gaussian hierarchy
Zhen Xu, Yinghao Xu, Zhiyuan Yu, Sida Peng, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. Representing long volumet- ric video with temporal gaussian hierarchy. ACM Transac- tions on Graphics, 43(6), 2024. 2, 4, 5, 6, 7, 1, 3
2024
-
[74]
Diffusion2: Dynamic 3d content generation via score composition of orthogonal diffusion models
Zeyu Yang, Zijie Pan, Chun Gu, and Li Zhang. Diffusion2: Dynamic 3d content generation via score composition of orthogonal diffusion models. arXiv e-prints, pages arXiv– 2404, 2024. 3
2024
-
[75]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 3
2024 arXiv
-
[76]
Real- time photorealistic dynamic scene representation and render- ing with 4d gaussian splatting
Zeyu Yang, Hongye Yang, Zijie Pan, and Li Zhang. Real- time photorealistic dynamic scene representation and render- ing with 4d gaussian splatting. In International Conference on Learning Representations (ICLR), 2024. 2, 1
2024
-
[77]
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vi- sion (ECCV), pages 767–783, 2018. 3
2018
-
[78]
Recurrent mvsnet for high-resolution multi-view stereo depth inference
Yao Yao, Zixin Luo, Shiwei Li, Tianwei Shen, Tian Fang, and Long Quan. Recurrent mvsnet for high-resolution multi-view stereo depth inference. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5525–5534, 2019. 3
2019
-
[79]
4dgen: Grounded 4d content gen- eration with spatial-temporal consistency
Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao, and Yunchao Wei. 4dgen: Grounded 4d content gen- eration with spatial-temporal consistency. arXiv preprint arXiv:2312.17225, 2023. 3
2023 arXiv
-
[80]
Stag4d: Spatial-temporal anchored generative 4d gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu, Yuanxun Lu, Youtian Lin, Hao Zhu, Weiming Hu, Xun Cao, and Yao Yao. Stag4d: Spatial-temporal anchored generative 4d gaussians. In Eu- ropean Conference on Computer Vision , pages 163–179. Springer, 2024. 3
2024
-
[81]
4diffusion: Multi-view video dif- fusion model for 4d generation
Haiyu Zhang, Xinyuan Chen, Yaohui Wang, Xihui Liu, Yun- hong Wang, and Yu Qiao. 4diffusion: Multi-view video dif- fusion model for 4d generation. Advances in Neural Infor- mation Processing Systems, 37:15272–15295, 2025. 3
2025
-
[82]
Cameras as rays: Pose estimation via ray diffusion
Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. arXiv preprint arXiv:2402.14817, 2024. 4
2024 arXiv
-
[83]
Animate124: Animating one im- age to 4d dynamic scene
Yuyang Zhao, Zhiwen Yan, Enze Xie, Lanqing Hong, Zhen- guo Li, and Gim Hee Lee. Animate124: Animating one im- age to 4d dynamic scene. arXiv preprint arXiv:2311.14603,
-
[84]
Bilateral refer- ence for high-resolution dichotomous image segmentation
Peng Zheng, Dehong Gao, Deng-Ping Fan, Li Liu, Jorma Laaksonen, Wanli Ouyang, and Nicu Sebe. Bilateral refer- ence for high-resolution dichotomous image segmentation. CAAI Artificial Intelligence Research, 3:9150038, 2024. 1
2024
-
[85]
Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2024
-
[86]
A unified approach for text- and image-guided 4d scene generation
Yufeng Zheng, Xueting Li, Koki Nagano, Sifei Liu, Otmar Hilliges, and Shalini De Mello. A unified approach for text- and image-guided 4d scene generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7300–7309, 2024. 3 Diffuman4D:...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.