Pith. sign in

REVIEW 4 major objections 5 minor 94 references

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A double texture unprojection splits geometry from appearance, enabling real-time photorealistic 4K free-view rendering from sparse RGB video.

desk verdict Well-engineered method with a genuinely new double-unprojection idea, but the strongest claim of consistent, significant outperformance is contradicted by the paper's own Table 1. read the letter →

arxiv 2412.13183 v3 pith:FC5TMZ6O submitted 2024-12-17 cs.CV

classification cs.CV
keywords free-viewrenderingsparse-viewRGBvideotextureunprojectionhumanperformancecapture3DGaussiansplattingtemplatedeformationreal-timenovelviewsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that free-viewpoint rendering of a moving person from only a handful of RGB cameras can be made both photorealistic and real-time if the problem is split into two texture-unprojection steps instead of one. The first unprojection maps the sparse images onto a roughly posed body template; a lightweight network reads that distorted map to estimate coarse surface deformations. The deformed template then receives a second unprojection that is far better aligned, and a second network predicts fine-scale 3D Gaussian splats from it. In experiments on three benchmarks, the method outperforms state-of-the-art sparse-view renderers in PSNR, SSIM, and LPIPS while running at tens of frames per second on a single GPU, and it keeps working on out-of-distribution motions such as a standing long jump.

What carries the argument

The load-bearing object is the double texture unprojection itself: a function that warps pixels from camera space into the UV texel space of a human body template, computes a per-texel visibility mask from normals, depth, and segmentation, and fuses multi-view colors. The first unprojection lands on an LBS-posed template that is not yet deformed; GeoNet, a UNet, reads that first map plus a non-root normal map and outputs per-vertex deformations in canonical space, trained with Chamfer distance to NeuS2 point clouds plus smoothness regularizers. The deformed geometry is reposed and used for a second unprojection, and GauNet, a second UNet, reads the cleaner second map and predicts Gaussian parameters, namely displacements, spherical harmonics, scales, rotations, and opacities, per texel; a scale-refinement step multiplies predicted scales by the maximum edge-stretching ratios of the LBS deformation to avoid artifacts. Everything runs in 2D texture space so that both networks are lightweight and the whole pipeline stays real-time.

What would settle it

Construct a test where the skeletal pose is deliberately corrupted by a known amount, or the subject wears loose non-rigid clothing, before the first unprojection, and measure whether Chamfer distance to ground-truth point clouds and final PSNR degrade sharply once the pose error exceeds a threshold; if the degradation tracks the distortion of the first texture map, that tolerance is the limit of the double-unprojection design.

Watch

Extended reading notes

Core claim

The central discovery is that unprojecting the sparse RGB views twice, first onto a linear-blend-skinned template and then onto that same template after an estimated deformation, decouples coarse geometric deformation estimation from appearance synthesis, and this decoupling is what makes high-quality real-time rendering possible. The first map is heavily distorted but still encodes enough about surface deformation for the geometry network to predict vertex offsets; the second map, produced using the corrected geometry, has fewer ghosting artifacts and better alignment, so the Gaussian network only has to learn small residual displacements. The result is a pipeline built entirely from 2D CNNs in texture space that outputs photorealistic 4K novel views from one to four cameras and runs at up to 42 FPS on an RTX 3090, outperforming ENeRF, DVA, HoloChar, and GHG on standard metrics.

Load-bearing premise

Stage one assumes that the first texture map, obtained from a merely skin-posed template that is not yet deformed, stays aligned enough with the true body surface for the geometry network to extract useful deformation information; the paper calls this map heavily distorted but does not measure how much distortion is tolerable.

Editorial extensions

If this is right

  • Real-time telepresence and free-viewpoint replay become feasible with as few as four RGB cameras, since the whole pipeline runs at 42 FPS on a single RTX 3090.
  • Because geometry is conditioned on images rather than only on motion, the method generalizes to out-of-distribution poses where motion-only avatars fail.
  • Separating geometry recovery from appearance synthesis means each stage is a simpler regression problem; improving the deformed geometry directly improves the second unprojection and the final rendering.
  • The Gaussian scale refinement removes pose-dependent stretching artifacts, so the method stays stable under strong LBS-induced scale changes.
  • Higher-resolution texture maps, such as 512x512, further improve fidelity at some speed cost, and the method scales with GPU power.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the first unprojection's tolerance for misalignment is the binding constraint, then adding a coarse pose-refinement step before GeoNet, or supervising GeoNet to be robust to synthetic distortions, should extend the pipeline to looser clothing and stronger pose errors without changing the two-stage design.
  • The same double-unprojection principle may transfer to other template-based actors, such as hands, animals, or garments, wherever a parametric template and sparse views are available and the first map is distorted but informative.
  • Since the method already runs feed-forward, combining it with an online motion estimator could remove the need for an external motion-capture stage, making the entire capture-to-render system sparse and real-time.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Double Unprojected Textures (DUT), a two-stage method for sparse-view real-time human rendering. In the first stage, an LBS-posed body template is used to unproject sparse-view images into a texture map, from which a GeoNet predicts vertex displacements. The deformed template is then used for a second, more accurate texture unprojection, and a GauNet predicts 3D Gaussian parameters in texture space, followed by a Gaussian scale refinement that compensates for LBS-induced scale changes. The method is evaluated on DynaCap S3, ASH S22, THUman4.0 S2618, and a newly collected out-of-distribution subject, with PSNR/SSIM/LPIPS comparisons against ENeRF, DVA, HoloChar, and GHG, and with runtime measurements on a single RTX3090.

Significance. The proposed two-stage architecture is well motivated and the paper contains several useful components: a concrete double-unprojection scheme, an image-conditioned deformation network in texel space, and a Gaussian scale refinement with a dedicated ablation (Table 3). The runtime measurements, including the detailed per-module decomposition in the supplementary, are a strength, as is the attempt to evaluate on out-of-distribution motions. If the empirical claims were properly qualified, this would be a solid contribution to real-time sparse-view human rendering. However, the central advertised result—consistent and significant outperformance over all baselines—is not established by the reported experiments, as detailed in the major comments.

major comments (4)
  1. [Abstract; Sec. 4.1; Table 1] The abstract, Sec. 4.1, and the caption of Table 1 repeatedly state that DUT 'consistently outperforms' all baselines for both variants. This is contradicted by the paper's own Table 1 on S22 at 4K: DVA has PSNR 31.2019, while Ours has 30.6427 and Ours-Large has 30.8126. Both DUT variants therefore trail DVA on PSNR, and the unqualified 'consistently' claim is false. The authors should either relax the claim to the cells where DUT leads or provide a corrected comparison that explains this exception.
  2. [Sec. 4.1; Table 1] Table 1 and Sec. 4.1 report point estimates without any error bars, confidence intervals, or significance tests. The abstract's 'significantly surpasses' and Sec. 4.1's 'clear improvement' are not supported by the statistics shown, and even the cells where DUT leads cannot be interpreted as significant without variance information. Please report per-cell variance or repeated-run standard deviations and either perform a significance test or remove the word 'significantly'.
  3. [Evaluation Protocol (Sec. 4); supplementary Sec. E] The evaluation protocol for S2618 is not a clean generalization test. The main text states, 'For S2618... we include condition views for supervision,' and the supplementary states, 'The training views, condition views, and evaluation views do not overlap, except for S2618.' Since the S2618 rows of Table 1 are part of the 'consistently outperform' claim, this train/test overlap must be disclosed in the main text or the S2618 rows should be removed or explicitly treated as a within-training-distribution evaluation.
  4. [Sec. 3.2; supplementary Sec. I] Sec. 3.2 introduces the first texture unprojection as 'heavily distorted' but 'sufficient to learn coarse deformations,' yet no experiment quantifies the tolerance to misalignment. Because the second unprojection (Sec. 3.3) inherits geometry errors from the first stage, this assumption is load-bearing for the robustness claim. The supplementary motion-sensitivity analysis perturbs joint angles but does not directly vary the initial template-to-image misalignment or texture distortion. Please add an experiment that degrades the initial alignment (e.g., pose noise, shape perturbation, or a coarser template) and reports Chamfer distance and rendering metrics, so the operating range of the method is explicit.
minor comments (5)
  1. [Sec. 4.1] The phrase 'consistently better visually results' should be 'consistently better visual results.'
  2. [Throughout; Table 1] The method name DVA appears with an extra space as 'DV A' in several places, including Table 1 and Fig. 5; please fix the typesetting.
  3. [Eq. (13)] The selection of the maximum edge scaling ratio and the clamp to 1.0 are heuristics; a sentence explaining why the maximum rather than the mean is appropriate would improve clarity.
  4. [Supplementary Table 4] Supplementary Table 4 is labelled 'Quantitative Ablation' but contains runtime measurements; renaming it 'Runtime Ablation' would avoid confusion with the quality ablations in Tables 2 and 3.
  5. [Sec. 4 Evaluation Protocol] The number of training frames and testing frames per subject is not stated; adding these numbers would improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: DUT's method and evaluation rest on external benchmarks and ablations, not on self-citation or fitted predictions.

full rationale

The paper's derivation chain is not circular. The proposed Double Unprojected Textures (DUT) is an architectural contribution: a first texture unprojection on an LBS-posed template feeds a geometry network (GeoNet), whose predicted deformation is used for a second unprojection that conditions a Gaussian appearance network (GauNet). Each stage has its own supervision: GeoNet is trained with Chamfer distance to ground-truth point clouds reconstructed by NeuS2, and GauNet is trained with image reconstruction losses. These are external supervisory signals, not quantities that the paper claims to predict. The ablations in Tables 2 and 3 test each component (image conditioning, normal conditioning, number of views, double unprojection, Gaussian scale refinement) against independently constructed variants, and the main comparison in Table 1 evaluates against external baselines (ENeRF, DVA, HoloChar, GHG) on public datasets. No equation or design choice reduces to its own input: the first unprojected texture is used to estimate geometry, and the second unprojected texture is a re-projection using that geometry; this is a feed-forward cascade, not a fixed-point or self-referential construction. The use of prior work by the same authors (NeuS2 for ground-truth point clouds, ASH as an implementation reference, HoloChar as a baseline) is methodological and does not carry the argument: even if those citations were absent, the central claim of outperforming prior methods would stand or fall on the reported comparisons and ablations. Possible concerns about the evaluation protocol (e.g., S2618 condition views used for supervision) or about the 'consistently outperforms' wording relative to Table 1 are correctness or fairness issues, not circularity. Per the review rules, non-circular self-citation does not raise the circularity score.

Assumptions & free parameters 9 free parameters · 6 assumptions · 0 invented entities

The contributions rest on a chain of empirical modeling assumptions rather than a closed-form derivation. The free parameters are manual loss weights, visibility thresholds, and a scale clamp; none follows from first principles. The axioms concern template alignment, texel-space sufficiency, NeuS2 ground truth, and motion reliability. No new physical entities are introduced: the double-unprojected texture and refining scale map are computational artifacts whose validity is only measured through downstream rendering metrics.

free parameters (9)
  • lambda_Lap = 1.0
    Weight on the Laplacian loss in Eq. 7, chosen by hand to balance Chamfer distance against surface smoothness.
  • lambda_Iso = 0.1, 0.5 for hands
    Weight on the isometry loss in Eq. 7; the hand setting is chosen manually to preserve finger structure.
  • lambda_Nc = 0.001
    Weight on the normal consistency loss in Eq. 7, chosen by hand.
  • lambda_SSIM = 0.1
    SSIM loss weight in Eq. 16, chosen by hand.
  • lambda_IDMRF = 0.01
    IDMRF perceptual loss weight in Eq. 16, chosen by hand.
  • lambda_Reg = 0.005, 0.01 for large map
    Gaussian displacement regularization weight in Eq. 16, chosen by hand.
  • visibility threshold delta = 0.17
    Normal-difference visibility threshold in Eq. 18, set once for all experiments with no sensitivity analysis.
  • visibility threshold epsilon = 0.02
    Depth-difference visibility threshold in Eq. 19, set once for all experiments with no sensitivity analysis.
  • Gaussian scale refinement clamp = max(..., 1.0)
    Eq. 13 forces refining scales to be at least 1.0 to avoid vanishing gradients; this choice biases the appearance network and is not derived from data.
assumptions (6)
  • domain assumption A posed human template obtained from motion parameters via LBS lies close enough to the true body surface that the first unprojection preserves usable image evidence for deformation estimation.
    Invoked in Sec. 3.2 and Fig. 3; the paper calls the first map 'heavily distorted' and asserts it is sufficient to learn coarse deformations without quantifying the tolerable distortion.
  • domain assumption Texel-space 2D CNNs can regress geometry and Gaussian parameters from unprojected texture and normal maps.
    Core architectural premise throughout Sec. 3.2 and 3.3, motivated by prior works [32, 52, 59] but not proven here.
  • domain assumption NeuS2 point clouds are an adequate ground-truth geometry for Chamfer supervision of the deformation network.
    Used in the geometry loss LGeo (Eq. 7); the limitations section admits this supervision requires time-consuming per-frame reconstruction.
  • domain assumption The visibility map thresholds delta and epsilon correctly separate occluded from visible texels.
    Appears in Eqs. 3, 18, and 19; thresholds are set once and no sensitivity analysis is provided.
  • domain assumption Per-frame body motion M is available and sufficiently accurate at inference.
    Sec. 3 treats skeletal motion as an input; the supplementary tests corrupted elbow motion up to 60 degrees but the pipeline itself does not estimate pose.
  • standard math Linear blend skinning and 3D Gaussian splatting are accepted as correct deformation and rendering models.
    The method builds on LBS [34] and 3DGS [28] without proving or testing their error behavior, so DUT inherits their artifacts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures." pith.science (2026). https://pith.science/paper/FC5TMZ6O

@misc{pith2026241213183,
  author       = {Pith},
  title        = {Pith review of: Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FC5TMZ6O}},
  note         = {Machine review of arXiv:2412.13183}
}
read the original abstract

Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and appearance, or completely ignore sparse image information for geometry estimation, significantly harming visual quality and robustness to unseen body poses. To address these issues, we present Double Unprojected Textures, which at the core disentangles coarse geometric deformation estimation from appearance synthesis, enabling robust and photorealistic 4K rendering in real-time. Specifically, we first introduce a novel image-conditioned template deformation network, which estimates the coarse deformation of the human template from a first unprojected texture. This updated geometry is then used to apply a second and more accurate texture unprojection. The resulting texture map has fewer artifacts and better alignment with input views, which benefits our learning of finer-level geometry and appearance represented by Gaussian splats. We validate the effectiveness and efficiency of the proposed method in quantitative and qualitative experiments, which significantly surpasses other state-of-the-art methods. Project page: https://vcai.mpi-inf.mpg.de/projects/DUT/

Figures

Figures reproduced from arXiv: 2412.13183 by the authors.

Figure 1
Figure 1. We propose Double Unprojected Textures (DUT), a new method to synthesize photoreal 4K novel-view renderings in real￾time. Our method consistently beats baseline approaches [32, 38, 52, 59] in terms of rendering quality and inference speed. Moreover, it generalizes to, both, in-distribution (IND) motions, i.e. dancing, and out-of-distribution (OOD) motions, i.e. standing long jump. Abstract Real-time free-view human … view at source ↗
Figure 2
Figure 2. Overview of DUT. Given sparse-view images and respective motion, DUT predicts coarse template geometry and fine-grained 3D Gaussians. We first unproject images onto the posed template to obtain a texture map, which is fed into GeoNet to estimate deformations of the template in canonical pose. We then unproject images again onto the posed and deformed template to obtain a less-distorted texture map, which serves as i… view at source ↗
Figure 3
Figure 3. Unprojected Texture Maps. Performing a second tex￾ture unprojection using the deformed template leads to less ghost￾ing artifacts and better geometric alignment. previous stage significantly improves the rendering quality of the learned Gaussians as we outline in the following. Inspired by previous methods [32, 37, 48] that use nor￾mal maps, position maps, or partial texture maps to es￾timate 3D Gaussians, we propos… view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Qualitative Results. Here, we demonstrate novel-view rendering results on novel poses at 4K or full resolutions. Note that DUT faithfully captures cloth wrinkles, body details on hands and faces. Besides, it is robust to diverse poses and self occlusions. S3 Input Imag…
Figure 5
Figure 5. Figure 5: Qualitative Comparison. Given sparse-view images with novel poses, our method captures sharper and more faithful appearance details including facial expressions, hand gestures, and cloth wrinkles at 4K or full resolution, compared to prior works, i.e. GHG [32], HoloCha…
Figure 6
Figure 6. Figure 6: Qualitative Comparison. Compared to HoloChar [59], our method successfully generalizes to out-of-distribution (OOD) motions, e.g. long standing jump. fer distance (CD), surface regularity (SR) [13], and self￾intersection (SI) [26] to evaluate the accuracy and smooth￾ne…
Figure 7
Figure 7. Figure 7: Qualitative Ablation. Here, we compare the perfor￾mance of our design choices and of baseline methods, DDC [19] and LBS [34], in terms of template deformations. Our full model, using, both, RGB and normal features, achieves better accuracy. Methods Backbone Tex Res CD↓…
Figure 8
Figure 8. Figure 8: Qualitative Ablation. We evaluate the design choices of texel Gaussian prediction module. Compared to baseline variants, our full method generates better visual results. Ours-Large achieves even better performance, at the cost of slower inference speed. is also quantit…
Figure 9
Figure 9. Figure 9: Given four-view image streams and body motions from [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Compared to the visibility map computed by GHG [ [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Quantitative Comparison. We quantitatively compare the rendering results of our method and HoloChar [59] on out-of￾distribution human motions. Our method produces consistently better rendering results. Methods Motion Tex Res PSNR ↑ SSIM ↑ LPIPS ↓ Ours-Large Sparse 512…
Figure 12
Figure 12. Figure 12: Motion Sensitivity Analysis to Increasing Errors. Our method still outputs reasonable results to different level of motion capture errors at inference [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Qualitative results. We perform our method on a loose and long-hair subject, it manages to capture the coarse deformations of hair and dress and produces faithful rendering results. out-of-distribution motions. Our method produces consis￾tently better results, and als…
Figure 14
Figure 14. Figure 14: Qualitative results. Results with finetuned model under novel lighting. After finetuning, DUT still runs in feed-forward manner. Deformed Canonical Template Multi-view Images Posed Template (Undeformed) Undeformed Texture Map Difference Map [PITH_FULL_IMAGE:figures/f…
Figure 15
Figure 15. Figure 15: Illustration of Undeformed Texture Map. The distortions of undeformed (first) texture maps are directly related with deformations on the canonical template [PITH_FULL_IMAGE:figures/full_fig_p018_15.png]
Figure 16
Figure 16. Figure 16: Topology Change. Results of taking off cloth. overfit the uneven lightings of studio and shadows on the body, which are challenging for such simple representation. Integrating ray tracing or physically based rendering may reduce the color fluctuation. Topology Change.…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 69 canonical work pages

  1. [1]

    Scape: shape completion and animation of people

    Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Se- bastian Thrun, Jim Rodgers, and James Davis. Scape: shape completion and animation of people. In ACM SIGGRAPH 2005 Papers, pages 408–416. 2005. 2

  2. [2]

    Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences

    Federica Bogo, Michael J Black, Matthew Loper, and Javier Romero. Detailed full-body reconstructions of moving peo- ple from monocular rgb-d sequences. In Proceedings of the IEEE international conference on computer vision, pages 2300–2308, 2015. 2

  3. [3]

    Distance transformations in digital im- ages

    Gunilla Borgefors. Distance transformations in digital im- ages. Computer vision, graphics, and image processing, 34 (3):344–371, 1986. 15

  4. [4]

    Egoavatar: Egocentric view-driven and photorealistic full- body avatars

    Jianchun Chen, Jian Wang, Yinda Zhang, Rohit Pandey, Thabo Beeler, Marc Habermann, and Christian Theobalt. Egoavatar: Egocentric view-driven and photorealistic full- body avatars. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 14

  5. [5]

    High-quality streamable free-viewpoint video

    Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (ToG), 34(4):1–13,

  6. [6]

    A volumetric method for building complex models from range images

    Brian Curless and Marc Levoy. A volumetric method for building complex models from range images. InProceedings of the 23rd annual conference on Computer graphics and interactive techniques, pages 303–312, 1996. 3

  7. [7]

    Performance capture from sparse multi-view video

    Edilson De Aguiar, Carsten Stoll, Christian Theobalt, Naveed Ahmed, Hans-Peter Seidel, and Sebastian Thrun. Performance capture from sparse multi-view video. In ACM SIGGRAPH 2008 papers, pages 1–10. 2008. 2

  8. [8]

    Fusion4d: Real-time performance capture of challeng- ing scenes

    Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Taylor, et al. Fusion4d: Real-time performance capture of challeng- ing scenes. ACM Transactions on Graphics (ToG), 35(4): 1–13, 2016. 2

Show all 94 references
  1. [9]

    Motion2fusion: Real-time volumetric performance capture

    Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, and Shahram Izadi. Motion2fusion: Real-time volumetric performance capture. ACM Transactions on Graphics (ToG), 36(6):1–16, 2017. 3

  2. [10]

    Fof: Learning fourier occupancy field for monocular real- time human reconstruction

    Qiao Feng, Yebin Liu, Yu-Kun Lai, Jingyu Yang, and Kun Li. Fof: Learning fourier occupancy field for monocular real- time human reconstruction. Advances in Neural Information Processing Systems, 35:7397–7409, 2022. 3

  3. [11]

    Capturing and animation of body and clothing from monocular video

    Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J Black, and Timo Bolkart. Capturing and animation of body and clothing from monocular video. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022. 14

  4. [12]

    Motion capture using joint skeleton tracking and surface estimation

    Juergen Gall, Carsten Stoll, Edilson De Aguiar, Christian Theobalt, Bodo Rosenhahn, and Hans-Peter Seidel. Motion capture using joint skeleton tracking and surface estimation. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 1746–1753. Ieee, 2009. 2

  5. [13]

    A latent implicit 3d shape model for multiple levels of detail

    Benoit Guillard, Marc Habermann, Christian Theobalt, and Pascal Fua. A latent implicit 3d shape model for multiple levels of detail. arXiv preprint arXiv:2409.06231, 2024. 6

  6. [14]

    Human performance capture from monocular video in the wild

    Chen Guo, Xu Chen, Jie Song, and Otmar Hilliges. Human performance capture from monocular video in the wild. In 2021 International Conference on 3D Vision (3DV), pages 889–898. IEEE, 2021. 2

  7. [15]

    Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition

    Chen Guo, Tianjian Jiang, Xu Chen, Jie Song, and Ot- mar Hilliges. Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12858–12868, 2023. 2

  8. [16]

    The re- lightables: V olumetric performance capture of humans with realistic relighting

    Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts- Escolano, Rohit Pandey, Jason Dourgarian, et al. The re- lightables: V olumetric performance capture of humans with realistic relighting. ACM Transactions on Graphics (To...

  9. [17]

    Livecap: Real-time human performance capture from monocular video

    Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger- ard Pons-Moll, and Christian Theobalt. Livecap: Real-time human performance capture from monocular video. ACM Transactions On Graphics (TOG), 38(2):1–17, 2019. 2, 3

  10. [18]

    Deepcap: Monoc- ular human performance capture using weak supervision

    Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger- ard Pons-Moll, and Christian Theobalt. Deepcap: Monoc- ular human performance capture using weak supervision. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020. 2, 3, 4

  11. [19]

    Real-time deep dynamic characters

    Marc Habermann, Lingjie Liu, Weipeng Xu, Michael Zoll- hoefer, Gerard Pons-Moll, and Christian Theobalt. Real-time deep dynamic characters. ACM Transactions on Graphics (ToG), 40(4):1–16, 2021. 2, 3, 4, 5, 7, 8, 15, 17

  12. [20]

    Hd- humans: A hybrid approach for high-fidelity digital hu- mans

    Marc Habermann, Lingjie Liu, Weipeng Xu, Gerard Pons- Moll, Michael Zollhoefer, and Christian Theobalt. Hd- humans: A hybrid approach for high-fidelity digital hu- mans. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 6(3):1–23, 2023. 2

  13. [21]

    Survey of texture mapping

    Paul S Heckbert. Survey of texture mapping. IEEE computer graphics and applications, 6(11):56–67, 1986. 4

  14. [22]

    Real- time 3d motion capture

    Thanarat Horprasert, Ismail Haritaoglu, Christopher Wren, David Harwood, Larry Davis, and Alex Pentland. Real- time 3d motion capture. In Second workshop on perceptual interfaces. Citeseer, 1998. 3

  15. [23]

    Tech: Text- guided reconstruction of lifelike clothed humans

    Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Jiaxiang Tang, Deng Cai, and Justus Thies. Tech: Text- guided reconstruction of lifelike clothed humans. In 2024 International Conference on 3D Vision (3DV), pages 1531–

  16. [24]

    Deep volumetric video from very sparse multi-view perfor- mance capture

    Zeng Huang, Tianye Li, Weikai Chen, Yajie Zhao, Jun Xing, Chloe LeGendre, Linjie Luo, Chongyang Ma, and Hao Li. Deep volumetric video from very sparse multi-view perfor- mance capture. In Proceedings of the European Conference on Computer Vision (ECCV), pages 336–354, 2018. 2

  17. [25]

    Neuralhofu- sion: Neural volumetric rendering under human-object in- teractions

    Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kai- wen Guo, Minye Wu, Jingyi Yu, and Lan Xu. Neuralhofu- sion: Neural volumetric rendering under human-object in- teractions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155– 616...

  18. [26]

    Self-intersection removal in triangular mesh offsetting

    Wonhyung Jung, Hayong Shin, and Byoung K Choi. Self-intersection removal in triangular mesh offsetting. Computer-Aided Design and Applications, 1(1-4):477–484,

  19. [27]

    Real- time animation of realistic virtual humans

    Prem Kalra, Nadia Magnenat-Thalmann, Laurent Moccozet, Gael Sannier, Amaury Aubel, and Daniel Thalmann. Real- time animation of realistic virtual humans. IEEE Computer Graphics and Applications, 18(5):42–56, 1998. 3

  20. [28]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,

  21. [29]

    Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. 3

  22. [30]

    Neural human performer: Learning generalizable ra- diance fields for human performance rendering

    Youngjoong Kwon, Dahun Kim, Duygu Ceylan, and Henry Fuchs. Neural human performer: Learning generalizable ra- diance fields for human performance rendering. Advances in Neural Information Processing Systems, 34:24741–24752,

  23. [31]

    Deliffas: Deformable light fields for fast avatar synthesis

    Youngjoong Kwon, Lingjie Liu, Henry Fuchs, Marc Haber- mann, and Christian Theobalt. Deliffas: Deformable light fields for fast avatar synthesis. Advances in Neural Information Processing Systems, 2023. 2

  24. [32]

    Gener- alizable human gaussians for sparse view synthesis

    Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang, Francisco Vicente Carrasco, Albert Mosella- Montoro, Jianjin Xu, Shingo Takagi, Daeil Kim, et al. Gener- alizable human gaussians for sparse view synthesis. ECCV,

  25. [33]

    Modular primi- tives for high-performance differentiable rendering

    Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primi- tives for high-performance differentiable rendering. ACM Transactions on Graphics, 39(6), 2020. 15

  26. [34]

    Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation

    JP Lewis, Matt Cordner, and Nickson Fong. Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 165–172, 2000. 2, 3, 7

  27. [35]

    Monocular real-time volumetric per- formance capture

    Ruilong Li, Yuliang Xiu, Shunsuke Saito, Zeng Huang, Kyle Olszewski, and Hao Li. Monocular real-time volumetric per- formance capture. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pages 49–67. Springer, 2020. 3

  28. [36]

    A survey of convolutional neural networks: analy- sis, applications, and prospects

    Zewen Li, Fan Liu, Wenjie Yang, Shouheng Peng, and Jun Zhou. A survey of convolutional neural networks: analy- sis, applications, and prospects. IEEE transactions on neural networks and learning systems, 33(12):6999–7019, 2021. 4

  29. [37]

    Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling

    Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Ani- matable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19711–19722, 2024. 5

  30. [38]

    Efficient neural radiance fields for interactive free-viewpoint video

    Haotong Lin, Sida Peng, Zhen Xu, Yunzhi Yan, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Efficient neural radiance fields for interactive free-viewpoint video. In SIGGRAPH Asia Conference Proceedings, 2022. 1, 6, 7, 15

  31. [39]

    Real-time high-resolution background matting

    Shanchuan Lin, Andrey Ryabtsev, Soumyadip Sengupta, Brian L Curless, Steven M Seitz, and Ira Kemelmacher- Shlizerman. Real-time high-resolution background matting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8762–8771, 2021. 3

  32. [40]

    Neural actor: Neural free-view synthesis of human actors with pose con- trol

    Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose con- trol. ACM Trans. Graph.(ACM SIGGRAPH Asia), 2021. 4, 17

  33. [41]

    Markerless motion capture of inter- acting characters using multi-view image segmentation

    Yebin Liu, Carsten Stoll, Juergen Gall, Hans-Peter Seidel, and Christian Theobalt. Markerless motion capture of inter- acting characters using multi-view image segmentation. In CVPR 2011, pages 1249–1256. Ieee, 2011. 2

  34. [42]

    Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 6

  35. [43]

    Decoupled weight decay regularization

    I Loshchilov. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5

  36. [44]

    Image-based visual hulls

    Wojciech Matusik, Chris Buehler, Ramesh Raskar, Steven J Gortler, and Leonard McMillan. Image-based visual hulls. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 369–374, 2000. 2

  37. [45]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 1

  38. [46]

    A real time anatomical converter for human motion capture

    Tom Molet, Ronan Boulic, and Daniel Thalmann. A real time anatomical converter for human motion capture. In Computer Animation and Simulation’96: Proceedings of the Eurographics Workshop in Poitiers, France, August 31–September 1, 1996, pages 79–94. Springer, 1996. 3

  39. [47]

    Holoportation: Virtual 3d teleportation in real-time

    Sergio Orts-Escolano, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip L Davidson, Sameh Khamis, Mingsong Dou, et al. Holoportation: Virtual 3d teleportation in real-time. In Proceedings of the 29th annual symposium on user interfa...

  40. [48]

    Ash: Animatable gaus- sian splats for efficient and photoreal human rendering

    Haokai Pang, Heming Zhu, Adam Kortylewski, Christian Theobalt, and Marc Habermann. Ash: Animatable gaus- sian splats for efficient and photoreal human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1165–1175,

  41. [49]

    Automatic differentiation in pytorch

    Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 15

  42. [50]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...

  43. [51]

    Shell maps

    Serban D Porumbescu, Brian Budge, Louis Feng, and Ken- neth I Joy. Shell maps. ACM Transactions on Graphics (TOG), 24(3):626–633, 2005. 3

  44. [52]

    Drivable volumetric avatars using texel-aligned features

    Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito, Chenglei Wu, Tomas Simon, Shih-En Wei, Kaiwen Guo, Zhe Cao, Fabian Prada, Jason Saragih, et al. Drivable volumetric avatars using texel-aligned features. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–9, 2022. 1, 2, 3, ...

  45. [53]

    Model-based outdoor perfor- mance capture

    Nadia Robertini, Dan Casas, Helge Rhodin, Hans-Peter Sei- del, and Christian Theobalt. Model-based outdoor perfor- mance capture. In Proceedings of the 2016 International Conference on 3D Vision (3DV 2016), 2016. 2

  46. [54]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...

  47. [55]

    Pifu: Pixel-aligned implicit function for high-resolution clothed human digi- tization

    Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Mor- ishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digi- tization. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2304–2314, 2019. 2, 3

  48. [56]

    Diffhu- man: Probabilistic photorealistic 3d reconstruction of hu- mans

    Akash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona, Andrei Zanfir, and Cristian Sminchisescu. Diffhu- man: Probabilistic photorealistic 3d reconstruction of hu- mans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page...

  49. [57]

    Floren: Real-time high-quality human performance rendering via appearance flow using sparse rgb cameras

    Ruizhi Shao, Liliang Chen, Zerong Zheng, Hongwen Zhang, Yuxiang Zhang, Han Huang, Yandong Guo, and Yebin Liu. Floren: Real-time high-quality human performance rendering via appearance flow using sparse rgb cameras. In SIGGRAPH Asia 2022 Conference Papers, pages 1–10,

  50. [58]

    Diffustereo: High quality human recon- struction via diffusion-based stereo using sparse cameras

    Ruizhi Shao, Zerong Zheng, Hongwen Zhang, Jingxiang Sun, and Yebin Liu. Diffustereo: High quality human recon- struction via diffusion-based stereo using sparse cameras. In European Conference on Computer Vision, pages 702–720. Springer, 2022. 2

  51. [59]

    Holo- ported characters: Real-time free-viewpoint rendering of humans from sparse rgb cameras

    Ashwath Shetty, Marc Habermann, Guoxing Sun, Diogo Lu- vizon, Vladislav Golyanik, and Christian Theobalt. Holo- ported characters: Real-time free-viewpoint rendering of humans from sparse rgb cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  52. [60]

    Surface capture for performance-based animation

    Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE computer graphics and applications, 27(3):21–31, 2007. 2

  53. [61]

    Fast articulated motion tracking us- ing a sums of gaussians body model

    Carsten Stoll, Nils Hasler, Juergen Gall, Hans-Peter Seidel, and Christian Theobalt. Fast articulated motion tracking us- ing a sums of gaussians body model. In 2011 International Conference on Computer Vision, pages 951–958. IEEE,

  54. [62]

    Embedded deformation for shape manipulation

    Robert W Sumner, Johannes Schmid, and Mark Pauly. Embedded deformation for shape manipulation. In ACM siggraph 2007 papers, pages 80–es. 2007. 4

  55. [63]

    Neural free-viewpoint performance rendering under complex human-object interactions

    Guoxing Sun, Xin Chen, Yizhang Chen, Anqi Pang, Pei Lin, Yuheng Jiang, Lan Xu, Jingyi Yu, and Jingya Wang. Neural free-viewpoint performance rendering under complex human-object interactions. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4651–4660,

  56. [64]

    Metacap: Meta-learning priors from multi-view imagery for sparse-view human per- formance capture and rendering

    Guoxing Sun, Rishabh Dabral, Pascal Fua, Christian Theobalt, and Marc Habermann. Metacap: Meta-learning priors from multi-view imagery for sparse-view human per- formance capture and rendering. In ECCV, 2024. 1, 2

  57. [65]

    3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos

    Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing. 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...

  58. [66]

    The Captury

    TheCaptury. The Captury. http://www.thecaptury. com/, 2023. 3, 14, 16

  59. [67]

    A parallel framework for silhouette-based human motion capture

    Christian Theobalt, Joel Carranza, Marcus A Magnor, and Hans-Peter Seidel. A parallel framework for silhouette-based human motion capture. In VMV, pages 207–214. Citeseer,

  60. [68]

    Combining 3d flow fields with silhouette- based human motion capture for immersive video.Graphical Models, 66(6):333–351, 2004

    Christian Theobalt, Joel Carranza, Marcus A Magnor, and Hans-Peter Seidel. Combining 3d flow fields with silhouette- based human motion capture for immersive video.Graphical Models, 66(6):333–351, 2004. 2

  61. [69]

    Neural-gif: Neural generalized implicit func- tions for animating people in clothing

    Garvita Tiwari, Nikolaos Sarafianos, Tony Tung, and Ger- ard Pons-Moll. Neural-gif: Neural generalized implicit func- tions for animating people in clothing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11708–11718, 2021. 2, 4

  62. [70]

    Articulated mesh animation from multi-view sil- houettes

    Daniel Vlasic, Ilya Baran, Wojciech Matusik, and Jovan Popovi´c. Articulated mesh animation from multi-view sil- houettes. In Acm Siggraph 2008 papers, pages 1–9. 2008. 2

  63. [71]

    Im- proved laplacian smoothing of noisy surface meshes

    J ¨org V ollmer, Robert Mencl, and Heinrich Mueller. Im- proved laplacian smoothing of noisy surface meshes. In Computer graphics forum, pages 131–138. Wiley Online Li- brary, 1999. 4

  64. [72]

    Metaavatar: Learning animatable clothed human models from few depth images

    Shaofei Wang, Marko Mihajlovic, Qianli Ma, Andreas Geiger, and Siyu Tang. Metaavatar: Learning animatable clothed human models from few depth images. Advances in Neural Information Processing Systems, 34:2810–2822,

  65. [73]

    Arah: Animatable volume rendering of articulated human sdfs

    Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu Tang. Arah: Animatable volume rendering of articulated human sdfs. In European Conference on Computer Vision,

  66. [74]

    Image inpainting via generative multi-column convolu- tional neural networks

    Yi Wang, Xin Tao, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Image inpainting via generative multi-column convolu- tional neural networks. In Advances in Neural Information Processing Systems, pages 331–340, 2018. 5, 14

  67. [75]

    Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction

    Yiming Wang, Qin Han, Marc Habermann, Kostas Dani- ilidis, Christian Theobalt, and Lingjie Liu. Neus2: Fast learning of neural implicit surfaces for multi-view recon- struction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 1, 4, 14

  68. [76]

    Image quality assessment: from error visibility to structural similarity

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 5, 6

  69. [77]

    Hu- mannerf: Free-viewpoint rendering of moving people from monocular video

    Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF conference on computer vision and pattern Recognition, pages 16210–...

  70. [78]

    Multi-view neural human rendering

    Minye Wu, Yuehao Wang, Qiang Hu, and Jingyi Yu. Multi-view neural human rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1682–1691, 2020. 1

  71. [79]

    Driv- able avatar clothing: Faithful full-body telepresence with dy- namic clothing driven by sparse rgb-d input

    Donglai Xiang, Fabian Prada, Zhe Cao, Kaiwen Guo, Chen- glei Wu, Jessica Hodgins, and Timur Bagautdinov. Driv- able avatar clothing: Faithful full-body telepresence with dy- namic clothing driven by sparse rgb-d input. In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023. 14, 17

  72. [80]

    Unstructuredfusion: realtime 4d geometry and texture recon- struction using commercial rgbd cameras

    Lan Xu, Zhuo Su, Lei Han, Tao Yu, Yebin Liu, and Lu Fang. Unstructuredfusion: realtime 4d geometry and texture recon- struction using commercial rgbd cameras. IEEE transactions on pattern analysis and machine intelligence, 42(10):2508– 2522, 2019. 2

  73. [81]

    Point- nerf: Point-based neural radiance fields

    Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point- nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5438–5448, 2022. 3

  74. [82]

    Monoperfcap: Human performance capture from monocular video

    Weipeng Xu, Avishek Chatterjee, Michael Zollh ¨ofer, Helge Rhodin, Dushyant Mehta, Hans-Peter Seidel, and Christian Theobalt. Monoperfcap: Human performance capture from monocular video. ACM Transactions on Graphics (ToG), 37 (2):1–15, 2018. 2

  75. [83]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024. 8

  76. [84]

    Generalizable neural voxels for fast human radiance fields

    Taoran Yi, Jiemin Fang, Xinggang Wang, and Wenyu Liu. Generalizable neural voxels for fast human radiance fields. arxiv:2303.15387, 2023. 1, 2

  77. [85]

    Body- fusion: Real-time capture of human motion and surface ge- ometry using a single depth camera

    Tao Yu, Kaiwen Guo, Feng Xu, Yuan Dong, Zhaoqi Su, Jian- hui Zhao, Jianguo Li, Qionghai Dai, and Yebin Liu. Body- fusion: Real-time capture of human motion and surface ge- ometry using a single depth camera. In Proceedings of the IEEE International Conference on Computer Visio...

  78. [86]

    Double- fusion: Real-time capture of human performances with in- ner body shapes from a single depth sensor

    Tao Yu, Zerong Zheng, Kaiwen Guo, Jianhui Zhao, Qiong- hai Dai, Hao Li, Gerard Pons-Moll, and Yebin Liu. Double- fusion: Real-time capture of human performances with in- ner body shapes from a single depth sensor. In Proceedings of the IEEE conference on computer vision and pa...

  79. [87]

    Function4d: Real-time human volumetric capture from very sparse consumer rgbd sen- sors

    Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qiong- hai Dai, and Yebin Liu. Function4d: Real-time human volumetric capture from very sparse consumer rgbd sen- sors. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR2021), 2021. 3

  80. [88]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 6

  81. [89]

    Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis

    Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  82. [90]

    Deepmulticap: Perfor- mance capture of multiple characters using sparse multiview cameras

    Yang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu, Zerong Zheng, Qionghai Dai, and Yebin Liu. Deepmulticap: Perfor- mance capture of multiple characters using sparse multiview cameras. In IEEE Conference on Computer Vision (ICCV 2021), 2021. 2

  83. [91]

    Structured local radiance fields for human avatar modeling

    Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yan- dong Guo, and Yebin Liu. Structured local radiance fields for human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 5, 6, 7

  84. [92]

    Trihuman: a real-time and controllable tri- plane representation for detailed human geometry and ap- pearance synthesis

    Heming Zhu, Fangneng Zhan, Christian Theobalt, and Marc Habermann. Trihuman: a real-time and controllable tri- plane representation for detailed human geometry and ap- pearance synthesis. ACM Transactions on Graphics, 44(1): 1–17, 2024. 2

  85. [93]

    Surface splatting

    Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Surface splatting. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 371–378, 2001. 3 Figure 9. Given four-view image streams and body motions from the dis...

  86. [2024]

    1, 2, 3, 4, 5, 6, 7, 13, 14, 15

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.