REVIEW 4 major objections 5 minor 94 references
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A double texture unprojection splits geometry from appearance, enabling real-time photorealistic 4K free-view rendering from sparse RGB video.
desk verdict Well-engineered method with a genuinely new double-unprojection idea, but the strongest claim of consistent, significant outperformance is contradicted by the paper's own Table 1. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the double texture unprojection itself: a function that warps pixels from camera space into the UV texel space of a human body template, computes a per-texel visibility mask from normals, depth, and segmentation, and fuses multi-view colors. The first unprojection lands on an LBS-posed template that is not yet deformed; GeoNet, a UNet, reads that first map plus a non-root normal map and outputs per-vertex deformations in canonical space, trained with Chamfer distance to NeuS2 point clouds plus smoothness regularizers. The deformed geometry is reposed and used for a second unprojection, and GauNet, a second UNet, reads the cleaner second map and predicts Gaussian parameters, namely displacements, spherical harmonics, scales, rotations, and opacities, per texel; a scale-refinement step multiplies predicted scales by the maximum edge-stretching ratios of the LBS deformation to avoid artifacts. Everything runs in 2D texture space so that both networks are lightweight and the whole pipeline stays real-time.
What would settle it
Construct a test where the skeletal pose is deliberately corrupted by a known amount, or the subject wears loose non-rigid clothing, before the first unprojection, and measure whether Chamfer distance to ground-truth point clouds and final PSNR degrade sharply once the pose error exceeds a threshold; if the degradation tracks the distortion of the first texture map, that tolerance is the limit of the double-unprojection design.
Extended reading notes
Core claim
The central discovery is that unprojecting the sparse RGB views twice, first onto a linear-blend-skinned template and then onto that same template after an estimated deformation, decouples coarse geometric deformation estimation from appearance synthesis, and this decoupling is what makes high-quality real-time rendering possible. The first map is heavily distorted but still encodes enough about surface deformation for the geometry network to predict vertex offsets; the second map, produced using the corrected geometry, has fewer ghosting artifacts and better alignment, so the Gaussian network only has to learn small residual displacements. The result is a pipeline built entirely from 2D CNNs in texture space that outputs photorealistic 4K novel views from one to four cameras and runs at up to 42 FPS on an RTX 3090, outperforming ENeRF, DVA, HoloChar, and GHG on standard metrics.
Load-bearing premise
Stage one assumes that the first texture map, obtained from a merely skin-posed template that is not yet deformed, stays aligned enough with the true body surface for the geometry network to extract useful deformation information; the paper calls this map heavily distorted but does not measure how much distortion is tolerable.
Editorial extensions
If this is right
- Real-time telepresence and free-viewpoint replay become feasible with as few as four RGB cameras, since the whole pipeline runs at 42 FPS on a single RTX 3090.
- Because geometry is conditioned on images rather than only on motion, the method generalizes to out-of-distribution poses where motion-only avatars fail.
- Separating geometry recovery from appearance synthesis means each stage is a simpler regression problem; improving the deformed geometry directly improves the second unprojection and the final rendering.
- The Gaussian scale refinement removes pose-dependent stretching artifacts, so the method stays stable under strong LBS-induced scale changes.
- Higher-resolution texture maps, such as 512x512, further improve fidelity at some speed cost, and the method scales with GPU power.
Reading between the lines
- If the first unprojection's tolerance for misalignment is the binding constraint, then adding a coarse pose-refinement step before GeoNet, or supervising GeoNet to be robust to synthetic distortions, should extend the pipeline to looser clothing and stronger pose errors without changing the two-stage design.
- The same double-unprojection principle may transfer to other template-based actors, such as hands, animals, or garments, wherever a parametric template and sparse views are available and the first map is distorted but informative.
- Since the method already runs feed-forward, combining it with an online motion estimator could remove the need for an external motion-capture stage, making the entire capture-to-render system sparse and real-time.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Double Unprojected Textures (DUT), a two-stage method for sparse-view real-time human rendering. In the first stage, an LBS-posed body template is used to unproject sparse-view images into a texture map, from which a GeoNet predicts vertex displacements. The deformed template is then used for a second, more accurate texture unprojection, and a GauNet predicts 3D Gaussian parameters in texture space, followed by a Gaussian scale refinement that compensates for LBS-induced scale changes. The method is evaluated on DynaCap S3, ASH S22, THUman4.0 S2618, and a newly collected out-of-distribution subject, with PSNR/SSIM/LPIPS comparisons against ENeRF, DVA, HoloChar, and GHG, and with runtime measurements on a single RTX3090.
Significance. The proposed two-stage architecture is well motivated and the paper contains several useful components: a concrete double-unprojection scheme, an image-conditioned deformation network in texel space, and a Gaussian scale refinement with a dedicated ablation (Table 3). The runtime measurements, including the detailed per-module decomposition in the supplementary, are a strength, as is the attempt to evaluate on out-of-distribution motions. If the empirical claims were properly qualified, this would be a solid contribution to real-time sparse-view human rendering. However, the central advertised result—consistent and significant outperformance over all baselines—is not established by the reported experiments, as detailed in the major comments.
major comments (4)
- [Abstract; Sec. 4.1; Table 1] The abstract, Sec. 4.1, and the caption of Table 1 repeatedly state that DUT 'consistently outperforms' all baselines for both variants. This is contradicted by the paper's own Table 1 on S22 at 4K: DVA has PSNR 31.2019, while Ours has 30.6427 and Ours-Large has 30.8126. Both DUT variants therefore trail DVA on PSNR, and the unqualified 'consistently' claim is false. The authors should either relax the claim to the cells where DUT leads or provide a corrected comparison that explains this exception.
- [Sec. 4.1; Table 1] Table 1 and Sec. 4.1 report point estimates without any error bars, confidence intervals, or significance tests. The abstract's 'significantly surpasses' and Sec. 4.1's 'clear improvement' are not supported by the statistics shown, and even the cells where DUT leads cannot be interpreted as significant without variance information. Please report per-cell variance or repeated-run standard deviations and either perform a significance test or remove the word 'significantly'.
- [Evaluation Protocol (Sec. 4); supplementary Sec. E] The evaluation protocol for S2618 is not a clean generalization test. The main text states, 'For S2618... we include condition views for supervision,' and the supplementary states, 'The training views, condition views, and evaluation views do not overlap, except for S2618.' Since the S2618 rows of Table 1 are part of the 'consistently outperform' claim, this train/test overlap must be disclosed in the main text or the S2618 rows should be removed or explicitly treated as a within-training-distribution evaluation.
- [Sec. 3.2; supplementary Sec. I] Sec. 3.2 introduces the first texture unprojection as 'heavily distorted' but 'sufficient to learn coarse deformations,' yet no experiment quantifies the tolerance to misalignment. Because the second unprojection (Sec. 3.3) inherits geometry errors from the first stage, this assumption is load-bearing for the robustness claim. The supplementary motion-sensitivity analysis perturbs joint angles but does not directly vary the initial template-to-image misalignment or texture distortion. Please add an experiment that degrades the initial alignment (e.g., pose noise, shape perturbation, or a coarser template) and reports Chamfer distance and rendering metrics, so the operating range of the method is explicit.
minor comments (5)
- [Sec. 4.1] The phrase 'consistently better visually results' should be 'consistently better visual results.'
- [Throughout; Table 1] The method name DVA appears with an extra space as 'DV A' in several places, including Table 1 and Fig. 5; please fix the typesetting.
- [Eq. (13)] The selection of the maximum edge scaling ratio and the clamp to 1.0 are heuristics; a sentence explaining why the maximum rather than the mean is appropriate would improve clarity.
- [Supplementary Table 4] Supplementary Table 4 is labelled 'Quantitative Ablation' but contains runtime measurements; renaming it 'Runtime Ablation' would avoid confusion with the quality ablations in Tables 2 and 3.
- [Sec. 4 Evaluation Protocol] The number of training frames and testing frames per subject is not stated; adding these numbers would improve reproducibility.
Circularity Check
No significant circularity: DUT's method and evaluation rest on external benchmarks and ablations, not on self-citation or fitted predictions.
full rationale
The paper's derivation chain is not circular. The proposed Double Unprojected Textures (DUT) is an architectural contribution: a first texture unprojection on an LBS-posed template feeds a geometry network (GeoNet), whose predicted deformation is used for a second unprojection that conditions a Gaussian appearance network (GauNet). Each stage has its own supervision: GeoNet is trained with Chamfer distance to ground-truth point clouds reconstructed by NeuS2, and GauNet is trained with image reconstruction losses. These are external supervisory signals, not quantities that the paper claims to predict. The ablations in Tables 2 and 3 test each component (image conditioning, normal conditioning, number of views, double unprojection, Gaussian scale refinement) against independently constructed variants, and the main comparison in Table 1 evaluates against external baselines (ENeRF, DVA, HoloChar, GHG) on public datasets. No equation or design choice reduces to its own input: the first unprojected texture is used to estimate geometry, and the second unprojected texture is a re-projection using that geometry; this is a feed-forward cascade, not a fixed-point or self-referential construction. The use of prior work by the same authors (NeuS2 for ground-truth point clouds, ASH as an implementation reference, HoloChar as a baseline) is methodological and does not carry the argument: even if those citations were absent, the central claim of outperforming prior methods would stand or fall on the reported comparisons and ablations. Possible concerns about the evaluation protocol (e.g., S2618 condition views used for supervision) or about the 'consistently outperforms' wording relative to Table 1 are correctness or fairness issues, not circularity. Per the review rules, non-circular self-citation does not raise the circularity score.
Assumptions & free parameters
free parameters (9)
- lambda_Lap =
1.0
- lambda_Iso =
0.1, 0.5 for hands
- lambda_Nc =
0.001
- lambda_SSIM =
0.1
- lambda_IDMRF =
0.01
- lambda_Reg =
0.005, 0.01 for large map
- visibility threshold delta =
0.17
- visibility threshold epsilon =
0.02
- Gaussian scale refinement clamp =
max(..., 1.0)
assumptions (6)
- domain assumption A posed human template obtained from motion parameters via LBS lies close enough to the true body surface that the first unprojection preserves usable image evidence for deformation estimation.
- domain assumption Texel-space 2D CNNs can regress geometry and Gaussian parameters from unprojected texture and normal maps.
- domain assumption NeuS2 point clouds are an adequate ground-truth geometry for Chamfer supervision of the deformation network.
- domain assumption The visibility map thresholds delta and epsilon correctly separate occluded from visible texels.
- domain assumption Per-frame body motion M is available and sufficiently accurate at inference.
- standard math Linear blend skinning and 3D Gaussian splatting are accepted as correct deformation and rendering models.
Cite this review
Pith. "Pith review of Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures." pith.science (2026). https://pith.science/paper/FC5TMZ6O
@misc{pith2026241213183,
author = {Pith},
title = {Pith review of: Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures},
year = {2026},
howpublished = {\url{https://pith.science/paper/FC5TMZ6O}},
note = {Machine review of arXiv:2412.13183}
}
read the original abstract
Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and appearance, or completely ignore sparse image information for geometry estimation, significantly harming visual quality and robustness to unseen body poses. To address these issues, we present Double Unprojected Textures, which at the core disentangles coarse geometric deformation estimation from appearance synthesis, enabling robust and photorealistic 4K rendering in real-time. Specifically, we first introduce a novel image-conditioned template deformation network, which estimates the coarse deformation of the human template from a first unprojected texture. This updated geometry is then used to apply a second and more accurate texture unprojection. The resulting texture map has fewer artifacts and better alignment with input views, which benefits our learning of finer-level geometry and appearance represented by Gaussian splats. We validate the effectiveness and efficiency of the proposed method in quantitative and qualitative experiments, which significantly surpasses other state-of-the-art methods. Project page: https://vcai.mpi-inf.mpg.de/projects/DUT/
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Scape: shape completion and animation of people
Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Se- bastian Thrun, Jim Rodgers, and James Davis. Scape: shape completion and animation of people. In ACM SIGGRAPH 2005 Papers, pages 408–416. 2005. 2
2005
-
[2]
Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences
Federica Bogo, Michael J Black, Matthew Loper, and Javier Romero. Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences. In Proceedings of the IEEE international conference on computer vision, pages 2300–2308, 2015. 2
2015
-
[3]
Distance transformations in digital im- ages
Gunilla Borgefors. Distance transformations in digital im- ages. Computer vision, graphics, and image processing, 34 (3):344–371, 1986. 15
1986
-
[4]
Egoavatar: Egocentric view-driven and photorealistic full- body avatars
Jianchun Chen, Jian Wang, Yinda Zhang, Rohit Pandey, Thabo Beeler, Marc Habermann, and Christian Theobalt. Egoavatar: Egocentric view-driven and photorealistic full- body avatars. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 14
2024
-
[5]
High-quality streamable free-viewpoint video
Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (ToG), 34(4):1–13,
-
[6]
A volumetric method for building complex models from range images
Brian Curless and Marc Levoy. A volumetric method for building complex models from range images. InProceedings of the 23rd annual conference on Computer graphics and interactive techniques, pages 303–312, 1996. 3
1996
-
[7]
Performance capture from sparse multi-view video
Edilson De Aguiar, Carsten Stoll, Christian Theobalt, Naveed Ahmed, Hans-Peter Seidel, and Sebastian Thrun. Performance capture from sparse multi-view video. In ACM SIGGRAPH 2008 papers, pages 1–10. 2008. 2
2008
-
[8]
Fusion4d: Real-time performance capture of challeng- ing scenes
Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Taylor, et al. Fusion4d: Real-time performance capture of challeng- ing scenes. ACM Transactions on Graphics (ToG), 35(4): 1–13, 2016. 2
2016
Show all 94 references
-
[9]
Motion2fusion: Real-time volumetric performance capture
Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, and Shahram Izadi. Motion2fusion: Real-time volumetric performance capture. ACM Transactions on Graphics (ToG), 36(6):1–16, 2017. 3
2017
-
[10]
Fof: Learning fourier occupancy field for monocular real- time human reconstruction
Qiao Feng, Yebin Liu, Yu-Kun Lai, Jingyu Yang, and Kun Li. Fof: Learning fourier occupancy field for monocular real- time human reconstruction. Advances in Neural Information Processing Systems, 35:7397–7409, 2022. 3
2022
-
[11]
Capturing and animation of body and clothing from monocular video
Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J Black, and Timo Bolkart. Capturing and animation of body and clothing from monocular video. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022. 14
2022
-
[12]
Motion capture using joint skeleton tracking and surface estimation
Juergen Gall, Carsten Stoll, Edilson De Aguiar, Christian Theobalt, Bodo Rosenhahn, and Hans-Peter Seidel. Motion capture using joint skeleton tracking and surface estimation. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 1746–1753. Ieee, 2009. 2
2009
-
[13]
A latent implicit 3d shape model for multiple levels of detail
Benoit Guillard, Marc Habermann, Christian Theobalt, and Pascal Fua. A latent implicit 3d shape model for multiple levels of detail. arXiv preprint arXiv:2409.06231, 2024. 6
2024 arXiv
-
[14]
Human performance capture from monocular video in the wild
Chen Guo, Xu Chen, Jie Song, and Otmar Hilliges. Human performance capture from monocular video in the wild. In 2021 International Conference on 3D Vision (3DV), pages 889–898. IEEE, 2021. 2
2021
-
[15]
Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition
Chen Guo, Tianjian Jiang, Xu Chen, Jie Song, and Ot- mar Hilliges. Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12858–12868, 2023. 2
2023
-
[16]
The re- lightables: V olumetric performance capture of humans with realistic relighting
Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts- Escolano, Rohit Pandey, Jason Dourgarian, et al. The re- lightables: V olumetric performance capture of humans with realistic relighting. ACM Transactions on Graphics (To...
2019
-
[17]
Livecap: Real-time human performance capture from monocular video
Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger- ard Pons-Moll, and Christian Theobalt. Livecap: Real-time human performance capture from monocular video. ACM Transactions On Graphics (TOG), 38(2):1–17, 2019. 2, 3
2019
-
[18]
Deepcap: Monoc- ular human performance capture using weak supervision
Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger- ard Pons-Moll, and Christian Theobalt. Deepcap: Monoc- ular human performance capture using weak supervision. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020. 2, 3, 4
2020
-
[19]
Real-time deep dynamic characters
Marc Habermann, Lingjie Liu, Weipeng Xu, Michael Zoll- hoefer, Gerard Pons-Moll, and Christian Theobalt. Real-time deep dynamic characters. ACM Transactions on Graphics (ToG), 40(4):1–16, 2021. 2, 3, 4, 5, 7, 8, 15, 17
2021
-
[20]
Hd- humans: A hybrid approach for high-fidelity digital hu- mans
Marc Habermann, Lingjie Liu, Weipeng Xu, Gerard Pons- Moll, Michael Zollhoefer, and Christian Theobalt. Hd- humans: A hybrid approach for high-fidelity digital hu- mans. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 6(3):1–23, 2023. 2
2023
-
[21]
Survey of texture mapping
Paul S Heckbert. Survey of texture mapping. IEEE computer graphics and applications, 6(11):56–67, 1986. 4
1986
-
[22]
Real- time 3d motion capture
Thanarat Horprasert, Ismail Haritaoglu, Christopher Wren, David Harwood, Larry Davis, and Alex Pentland. Real- time 3d motion capture. In Second workshop on perceptual interfaces. Citeseer, 1998. 3
1998
-
[23]
Tech: Text- guided reconstruction of lifelike clothed humans
Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Jiaxiang Tang, Deng Cai, and Justus Thies. Tech: Text- guided reconstruction of lifelike clothed humans. In 2024 International Conference on 3D Vision (3DV), pages 1531–
2024
-
[24]
Deep volumetric video from very sparse multi-view perfor- mance capture
Zeng Huang, Tianye Li, Weikai Chen, Yajie Zhao, Jun Xing, Chloe LeGendre, Linjie Luo, Chongyang Ma, and Hao Li. Deep volumetric video from very sparse multi-view perfor- mance capture. In Proceedings of the European Conference on Computer Vision (ECCV), pages 336–354, 2018. 2
2018
-
[25]
Neuralhofu- sion: Neural volumetric rendering under human-object in- teractions
Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kai- wen Guo, Minye Wu, Jingyi Yu, and Lan Xu. Neuralhofu- sion: Neural volumetric rendering under human-object in- teractions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155– 616...
2022
-
[26]
Self-intersection removal in triangular mesh offsetting
Wonhyung Jung, Hayong Shin, and Byoung K Choi. Self-intersection removal in triangular mesh offsetting. Computer-Aided Design and Applications, 1(1-4):477–484,
-
[27]
Real- time animation of realistic virtual humans
Prem Kalra, Nadia Magnenat-Thalmann, Laurent Moccozet, Gael Sannier, Amaury Aubel, and Daniel Thalmann. Real- time animation of realistic virtual humans. IEEE Computer Graphics and Applications, 18(5):42–56, 1998. 3
1998
-
[28]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[29]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. 3
2023 arXiv
-
[30]
Neural human performer: Learning generalizable ra- diance fields for human performance rendering
Youngjoong Kwon, Dahun Kim, Duygu Ceylan, and Henry Fuchs. Neural human performer: Learning generalizable ra- diance fields for human performance rendering. Advances in Neural Information Processing Systems, 34:24741–24752,
-
[31]
Deliffas: Deformable light fields for fast avatar synthesis
Youngjoong Kwon, Lingjie Liu, Henry Fuchs, Marc Haber- mann, and Christian Theobalt. Deliffas: Deformable light fields for fast avatar synthesis. Advances in Neural Information Processing Systems, 2023. 2
2023
-
[32]
Gener- alizable human gaussians for sparse view synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang, Francisco Vicente Carrasco, Albert Mosella- Montoro, Jianjin Xu, Shingo Takagi, Daeil Kim, et al. Gener- alizable human gaussians for sparse view synthesis. ECCV,
-
[33]
Modular primi- tives for high-performance differentiable rendering
Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primi- tives for high-performance differentiable rendering. ACM Transactions on Graphics, 39(6), 2020. 15
2020
-
[34]
Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation
JP Lewis, Matt Cordner, and Nickson Fong. Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 165–172, 2000. 2, 3, 7
2000
-
[35]
Monocular real-time volumetric per- formance capture
Ruilong Li, Yuliang Xiu, Shunsuke Saito, Zeng Huang, Kyle Olszewski, and Hao Li. Monocular real-time volumetric per- formance capture. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pages 49–67. Springer, 2020. 3
2020
-
[36]
A survey of convolutional neural networks: analy- sis, applications, and prospects
Zewen Li, Fan Liu, Wenjie Yang, Shouheng Peng, and Jun Zhou. A survey of convolutional neural networks: analy- sis, applications, and prospects. IEEE transactions on neural networks and learning systems, 33(12):6999–7019, 2021. 4
2021
-
[37]
Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling
Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19711–19722, 2024. 5
2024
-
[38]
Efficient neural radiance fields for interactive free-viewpoint video
Haotong Lin, Sida Peng, Zhen Xu, Yunzhi Yan, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Efficient neural radiance fields for interactive free-viewpoint video. In SIGGRAPH Asia Conference Proceedings, 2022. 1, 6, 7, 15
2022
-
[39]
Real-time high-resolution background matting
Shanchuan Lin, Andrey Ryabtsev, Soumyadip Sengupta, Brian L Curless, Steven M Seitz, and Ira Kemelmacher- Shlizerman. Real-time high-resolution background matting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8762–8771, 2021. 3
2021
-
[40]
Neural actor: Neural free-view synthesis of human actors with pose con- trol
Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose con- trol. ACM Trans. Graph.(ACM SIGGRAPH Asia), 2021. 4, 17
2021
-
[41]
Markerless motion capture of inter- acting characters using multi-view image segmentation
Yebin Liu, Carsten Stoll, Juergen Gall, Hans-Peter Seidel, and Christian Theobalt. Markerless motion capture of inter- acting characters using multi-view image segmentation. In CVPR 2011, pages 1249–1256. Ieee, 2011. 2
2011
-
[42]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 6
2015
-
[43]
Decoupled weight decay regularization
I Loshchilov. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5
2017 arXiv
-
[44]
Image-based visual hulls
Wojciech Matusik, Chris Buehler, Ramesh Raskar, Steven J Gortler, and Leonard McMillan. Image-based visual hulls. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 369–374, 2000. 2
2000
-
[45]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 1
2020
-
[46]
A real time anatomical converter for human motion capture
Tom Molet, Ronan Boulic, and Daniel Thalmann. A real time anatomical converter for human motion capture. In Computer Animation and Simulation’96: Proceedings of the Eurographics Workshop in Poitiers, France, August 31–September 1, 1996, pages 79–94. Springer, 1996. 3
1996
-
[47]
Holoportation: Virtual 3d teleportation in real-time
Sergio Orts-Escolano, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip L Davidson, Sameh Khamis, Mingsong Dou, et al. Holoportation: Virtual 3d teleportation in real-time. In Proceedings of the 29th annual symposium on user interfa...
2016
-
[48]
Ash: Animatable gaus- sian splats for efficient and photoreal human rendering
Haokai Pang, Heming Zhu, Adam Kortylewski, Christian Theobalt, and Marc Habermann. Ash: Animatable gaus- sian splats for efficient and photoreal human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1165–1175,
-
[49]
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 15
2017
-
[50]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...
2019
-
[51]
Shell maps
Serban D Porumbescu, Brian Budge, Louis Feng, and Ken- neth I Joy. Shell maps. ACM Transactions on Graphics (TOG), 24(3):626–633, 2005. 3
2005
-
[52]
Drivable volumetric avatars using texel-aligned features
Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito, Chenglei Wu, Tomas Simon, Shih-En Wei, Kaiwen Guo, Zhe Cao, Fabian Prada, Jason Saragih, et al. Drivable volumetric avatars using texel-aligned features. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–9, 2022. 1, 2, 3, ...
2022
-
[53]
Model-based outdoor perfor- mance capture
Nadia Robertini, Dan Casas, Helge Rhodin, Hans-Peter Sei- del, and Christian Theobalt. Model-based outdoor perfor- mance capture. In Proceedings of the 2016 International Conference on 3D Vision (3DV 2016), 2016. 2
2016
-
[54]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...
2015
-
[55]
Pifu: Pixel-aligned implicit function for high-resolution clothed human digi- tization
Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Mor- ishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digi- tization. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2304–2314, 2019. 2, 3
2019
-
[56]
Diffhu- man: Probabilistic photorealistic 3d reconstruction of hu- mans
Akash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona, Andrei Zanfir, and Cristian Sminchisescu. Diffhu- man: Probabilistic photorealistic 3d reconstruction of hu- mans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page...
2024
-
[57]
Floren: Real-time high-quality human performance rendering via appearance flow using sparse rgb cameras
Ruizhi Shao, Liliang Chen, Zerong Zheng, Hongwen Zhang, Yuxiang Zhang, Han Huang, Yandong Guo, and Yebin Liu. Floren: Real-time high-quality human performance rendering via appearance flow using sparse rgb cameras. In SIGGRAPH Asia 2022 Conference Papers, pages 1–10,
2022
-
[58]
Diffustereo: High quality human recon- struction via diffusion-based stereo using sparse cameras
Ruizhi Shao, Zerong Zheng, Hongwen Zhang, Jingxiang Sun, and Yebin Liu. Diffustereo: High quality human recon- struction via diffusion-based stereo using sparse cameras. In European Conference on Computer Vision, pages 702–720. Springer, 2022. 2
2022
-
[59]
Holo- ported characters: Real-time free-viewpoint rendering of humans from sparse rgb cameras
Ashwath Shetty, Marc Habermann, Guoxing Sun, Diogo Lu- vizon, Vladislav Golyanik, and Christian Theobalt. Holo- ported characters: Real-time free-viewpoint rendering of humans from sparse rgb cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2024
-
[60]
Surface capture for performance-based animation
Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE computer graphics and applications, 27(3):21–31, 2007. 2
2007
-
[61]
Fast articulated motion tracking us- ing a sums of gaussians body model
Carsten Stoll, Nils Hasler, Juergen Gall, Hans-Peter Seidel, and Christian Theobalt. Fast articulated motion tracking us- ing a sums of gaussians body model. In 2011 International Conference on Computer Vision, pages 951–958. IEEE,
2011
-
[62]
Embedded deformation for shape manipulation
Robert W Sumner, Johannes Schmid, and Mark Pauly. Embedded deformation for shape manipulation. In ACM siggraph 2007 papers, pages 80–es. 2007. 4
2007
-
[63]
Neural free-viewpoint performance rendering under complex human-object interactions
Guoxing Sun, Xin Chen, Yizhang Chen, Anqi Pang, Pei Lin, Yuheng Jiang, Lan Xu, Jingyi Yu, and Jingya Wang. Neural free-viewpoint performance rendering under complex human-object interactions. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4651–4660,
-
[64]
Metacap: Meta-learning priors from multi-view imagery for sparse-view human per- formance capture and rendering
Guoxing Sun, Rishabh Dabral, Pascal Fua, Christian Theobalt, and Marc Habermann. Metacap: Meta-learning priors from multi-view imagery for sparse-view human per- formance capture and rendering. In ECCV, 2024. 1, 2
2024
-
[65]
3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos
Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing. 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...
2024
-
[66]
The Captury
TheCaptury. The Captury. http://www.thecaptury. com/, 2023. 3, 14, 16
2023
-
[67]
A parallel framework for silhouette-based human motion capture
Christian Theobalt, Joel Carranza, Marcus A Magnor, and Hans-Peter Seidel. A parallel framework for silhouette-based human motion capture. In VMV, pages 207–214. Citeseer,
-
[68]
Combining 3d flow fields with silhouette- based human motion capture for immersive video.Graphical Models, 66(6):333–351, 2004
Christian Theobalt, Joel Carranza, Marcus A Magnor, and Hans-Peter Seidel. Combining 3d flow fields with silhouette- based human motion capture for immersive video.Graphical Models, 66(6):333–351, 2004. 2
2004
-
[69]
Neural-gif: Neural generalized implicit func- tions for animating people in clothing
Garvita Tiwari, Nikolaos Sarafianos, Tony Tung, and Ger- ard Pons-Moll. Neural-gif: Neural generalized implicit func- tions for animating people in clothing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11708–11718, 2021. 2, 4
2021
-
[70]
Articulated mesh animation from multi-view sil- houettes
Daniel Vlasic, Ilya Baran, Wojciech Matusik, and Jovan Popovi´c. Articulated mesh animation from multi-view sil- houettes. In Acm Siggraph 2008 papers, pages 1–9. 2008. 2
2008
-
[71]
Im- proved laplacian smoothing of noisy surface meshes
J ¨org V ollmer, Robert Mencl, and Heinrich Mueller. Im- proved laplacian smoothing of noisy surface meshes. In Computer graphics forum, pages 131–138. Wiley Online Li- brary, 1999. 4
1999
-
[72]
Metaavatar: Learning animatable clothed human models from few depth images
Shaofei Wang, Marko Mihajlovic, Qianli Ma, Andreas Geiger, and Siyu Tang. Metaavatar: Learning animatable clothed human models from few depth images. Advances in Neural Information Processing Systems, 34:2810–2822,
-
[73]
Arah: Animatable volume rendering of articulated human sdfs
Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu Tang. Arah: Animatable volume rendering of articulated human sdfs. In European Conference on Computer Vision,
-
[74]
Image inpainting via generative multi-column convolu- tional neural networks
Yi Wang, Xin Tao, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Image inpainting via generative multi-column convolu- tional neural networks. In Advances in Neural Information Processing Systems, pages 331–340, 2018. 5, 14
2018
-
[75]
Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction
Yiming Wang, Qin Han, Marc Habermann, Kostas Dani- ilidis, Christian Theobalt, and Lingjie Liu. Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 1, 4, 14
2023
-
[76]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 5, 6
2004
-
[77]
Hu- mannerf: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF conference on computer vision and pattern Recognition, pages 16210–...
2022
-
[78]
Multi-view neural human rendering
Minye Wu, Yuehao Wang, Qiang Hu, and Jingyi Yu. Multi-view neural human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1682–1691, 2020. 1
2020
-
[79]
Driv- able avatar clothing: Faithful full-body telepresence with dy- namic clothing driven by sparse rgb-d input
Donglai Xiang, Fabian Prada, Zhe Cao, Kaiwen Guo, Chen- glei Wu, Jessica Hodgins, and Timur Bagautdinov. Driv- able avatar clothing: Faithful full-body telepresence with dy- namic clothing driven by sparse rgb-d input. In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023. 14, 17
2023
-
[80]
Unstructuredfusion: realtime 4d geometry and texture recon- struction using commercial rgbd cameras
Lan Xu, Zhuo Su, Lei Han, Tao Yu, Yebin Liu, and Lu Fang. Unstructuredfusion: realtime 4d geometry and texture recon- struction using commercial rgbd cameras. IEEE transactions on pattern analysis and machine intelligence, 42(10):2508– 2522, 2019. 2
2019
-
[81]
Point- nerf: Point-based neural radiance fields
Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point- nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5438–5448, 2022. 3
2022
-
[82]
Monoperfcap: Human performance capture from monocular video
Weipeng Xu, Avishek Chatterjee, Michael Zollh ¨ofer, Helge Rhodin, Dushyant Mehta, Hans-Peter Seidel, and Christian Theobalt. Monoperfcap: Human performance capture from monocular video. ACM Transactions on Graphics (ToG), 37 (2):1–15, 2018. 2
2018
-
[83]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024. 8
2024
-
[84]
Generalizable neural voxels for fast human radiance fields
Taoran Yi, Jiemin Fang, Xinggang Wang, and Wenyu Liu. Generalizable neural voxels for fast human radiance fields. arxiv:2303.15387, 2023. 1, 2
2023 arXiv
-
[85]
Body- fusion: Real-time capture of human motion and surface ge- ometry using a single depth camera
Tao Yu, Kaiwen Guo, Feng Xu, Yuan Dong, Zhaoqi Su, Jian- hui Zhao, Jianguo Li, Qionghai Dai, and Yebin Liu. Body- fusion: Real-time capture of human motion and surface ge- ometry using a single depth camera. In Proceedings of the IEEE International Conference on Computer Visio...
2017
-
[86]
Double- fusion: Real-time capture of human performances with in- ner body shapes from a single depth sensor
Tao Yu, Zerong Zheng, Kaiwen Guo, Jianhui Zhao, Qiong- hai Dai, Hao Li, Gerard Pons-Moll, and Yebin Liu. Double- fusion: Real-time capture of human performances with in- ner body shapes from a single depth sensor. In Proceedings of the IEEE conference on computer vision and pa...
2018
-
[87]
Function4d: Real-time human volumetric capture from very sparse consumer rgbd sen- sors
Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qiong- hai Dai, and Yebin Liu. Function4d: Real-time human volumetric capture from very sparse consumer rgbd sen- sors. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR2021), 2021. 3
2021
-
[88]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 6
2018
-
[89]
Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2024
-
[90]
Deepmulticap: Perfor- mance capture of multiple characters using sparse multiview cameras
Yang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu, Zerong Zheng, Qionghai Dai, and Yebin Liu. Deepmulticap: Perfor- mance capture of multiple characters using sparse multiview cameras. In IEEE Conference on Computer Vision (ICCV 2021), 2021. 2
2021
-
[91]
Structured local radiance fields for human avatar modeling
Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yan- dong Guo, and Yebin Liu. Structured local radiance fields for human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 5, 6, 7
2022
-
[92]
Trihuman: a real-time and controllable tri- plane representation for detailed human geometry and ap- pearance synthesis
Heming Zhu, Fangneng Zhan, Christian Theobalt, and Marc Habermann. Trihuman: a real-time and controllable tri- plane representation for detailed human geometry and ap- pearance synthesis. ACM Transactions on Graphics, 44(1): 1–17, 2024. 2
2024
-
[93]
Surface splatting
Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Surface splatting. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 371–378, 2001. 3 Figure 9. Given four-view image streams and body motions from the dis...
2001
-
[2024]
1, 2, 3, 4, 5, 6, 7, 13, 14, 15
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.