Pith. sign in

REVIEW 3 major objections 5 minor 68 references

Multi-Garment Net: Learning to Dress 3D People from Images

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that a few images of a person are enough to reconstruct the naked body and each garment as a separate, transferable 3D mesh.

desk verdict Valuable layered garment representation from images, but the quantitative evaluation needs clarification and the registration validation is missing. read the letter →

arxiv 1908.06903 v2 pith:IE2CJK5O submitted 2019-08-19 cs.CV

classification cs.CV
keywords 3DhumanreconstructiongarmentregistrationSMPLdigitalwardrobemulti-layerbodyrepresentationvirtualtry-onretargetingsemanticsegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims MGN is the first model that, given one to eight frames of a person turning in front of a camera, outputs the body shape and each garment (shirt, t-shirt, coat, short pants, long pants) as separate meshes rather than one fused surface. If true, clothing becomes a swappable asset: a garment extracted from one person can be dressed onto any other body and reposed, and texture can be transferred between garments of the same category. The method is trained on 356 real 3D scans, from which the authors register a digital wardrobe of 712 garment instances in vertex correspondence. On held-out scans, garment surfaces are reconstructed to within about 5.8 mm when the pose is known and 11.9 mm when the pose is predicted, at a small cost in raw accuracy relative to a single-mesh baseline that reports 5.72 mm.

What carries the argument

The machinery is the fixed-topology garment template attached to SMPL. For each garment category, one template mesh is placed in vertex correspondence with the body and registered to every scan instance; a Laplacian-boundary linear solve globally stretches the template to match clothing boundaries before non-rigid registration, and per-category PCA plus bounded residual displacements encode garment shape. This correspondence is what lets the network output separable layers and lets any garment be reposed or transferred.

What would settle it

Register a garment style that the five templates do not cover, such as a dress, a skirt, an open coat, or a shirt with sleeves of unusual cut, and compare MGN's predicted mesh against a high-resolution 3D scan of the same person. If the mean vertex-to-surface error on such items is far above the reported 5.78 mm with ground-truth pose or 11.90 mm with predicted pose, the claim that the method dresses people generally from images fails.

Watch

Extended reading notes

Core claim

The central claim is that clothing can be factored out of human shape reconstruction by learning per-category garment templates in correspondence with the SMPL body model, a skinned linear model of pose and shape. Registering one fixed-topology template per category to real scans yields a digital wardrobe; PCA on unposed garment vertices gives a pose-invariant low-dimensional shape space, and per-vertex displacements add high-frequency detail. A CNN consumes semantic segmentation images and 2D joint estimates, averages per-frame garment and shape codes, and predicts body shape, pose, and garment parameters, while a differentiable renderer with a per-garment segmentation loss forces each predicted layer to explain its own region in the image. The result is that body and garments are separate meshes that can be reposed, retargeted to new bodies, and re-textured.

Load-bearing premise

The load-bearing premise is that one fixed template mesh per garment category can be stretched into correspondence with every real garment of that category; if a style cannot be matched by the template, the registration error propagates into the PCA model and every network prediction.

Editorial extensions

If this is right

  • A few frames of a rotating person are enough to obtain a body mesh plus separate garment meshes, so virtual try-on can be driven by ordinary video rather than multi-camera capture.
  • Garments predicted from one subject can be dressed onto a different SMPL body in a different pose, making wardrobe transfer a direct operation rather than a physics simulation problem.
  • Because every garment in a category shares one topology and UV parameterization, texture can be mapped from any registered garment instance onto any other instance of the same category.
  • Training with per-garment 2D segmentation, not just whole-silhouette overlap, pushes each predicted layer to explain its own image region and yields cleaner garment boundaries than single-mesh displacement models.
  • The released digital wardrobe of 712 registered real garments lets anyone dress an SMPL body with real clothing rather than synthetic cloth simulations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same fixed-topology layer representation should generalize to dresses, skirts, and open coats only if new templates are added; the paper's five categories, not the method itself, bound the wardrobe.
  • Because the network averages per-frame garment codes before decoding, the quality of the input semantic segmentation likely sets the ceiling; a testable extension is to corrupt or drop segmentation regions and measure how garment errors grow.
  • The retargeting step associates each source garment vertex with its nearest body vertex, so loose or draped garments may interpenetrate on very different body shapes; a learned or physics-aware association would be the next step beyond the paper.
  • The 5.78 mm error is measured on held-out scans from the same capture setup used for training; the claim 'from images directly' would be tested harder on in-the-wild web photos, which the paper does not evaluate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces Multi-Garment Network (MGN), a method that predicts a layered representation of a person's body and separate garment meshes from a small set of RGB images (1–8 frames). The method builds on a digital wardrobe of 712 garment registrations obtained by registering fixed-topology garment templates to 356 real 3D scans, learns per-garment PCA shape spaces, and trains a CNN with a combination of 3D vertex losses, a 2D semantic segmentation loss, and intermediate losses. The paper also presents applications including garment retargeting across subjects and texture transfer between garments of the same category, and reports a mean vertex-to-surface garment error of 5.78 mm, comparing favorably with a re-trained version of Alldieck et al.

Significance. If the central claims hold, this is a meaningful step forward: it provides a representation that decouples body and clothing into separate meshes, enables garment retargeting and texture transfer from images, and releases a digital wardrobe and code that could be reused by the community. The paper's strengths include a concrete registration pipeline, a clear application-oriented evaluation, and a commitment to public release of assets and models. However, the quantitative support for the headline accuracy is internally inconsistent, and the evaluation lacks a validation of the registered garment meshes against the original scans, which is load-bearing because those registrations serve as both training supervision and evaluation reference.

major comments (3)
  1. [Sec. 4.1, Quantitative Comparison and GT vs Predicted pose] The paper reports a mean vertex-to-surface garment error of 5.78 mm with 8 frames as input, but then in the same section states a mean vertex-to-surface error of 5.78 mm with GT poses and 11.90 mm with predicted poses. These statements cannot both be correct unless the first number is conditioned on ground-truth pose. Please provide a consistent table that jointly reports the frame count, the pose condition, the mean error, and standard deviations, and clearly identify which number is the headline result for the method as described in the abstract.
  2. [Sec. 3.1 and Eq. (18)] The registered garment meshes produced by the pipeline of Sec. 3.1 are used both as supervision in Eqs. (12)–(13) and as the reference surfaces in the evaluation metric of Eq. (18). The paper reports no quantitative measure of registration quality against the original scan surfaces: no mean distance from registered templates to the scans, no per-category statistics, and no failure cases. Without such validation, the reported 5.78 mm error may reflect bias in the registrations rather than the accuracy of the predicted garments relative to the observed geometry. Please add per-category distance-to-scan statistics and discuss failure cases, especially for garment styles that deviate from the fixed-topology template.
  3. [Sec. 3.1, Garment Registration and Sec. 3.2, Garment Shape Space] The method assumes a single fixed-topology template per garment category can be registered to every instance of that category. This assumption is load-bearing because the Laplacian initialization of Eq. (6) followed by the 35-component PCA and capped high-frequency displacements will encode any registration bias as legitimate garment geometry, and the downstream evaluation in Eq. (18) cannot detect this. Please provide an explicit analysis of how template coverage limitations (e.g., dresses, skirts, open coats, unusual sleeves) affect the registration and the learned shape space, and quantify the proportion of scans for which the registration is expected to be accurate.
minor comments (5)
  1. [Sec. 3.3, Eq. (17)] The 2D segmentation loss is called a self-supervision loss, but it relies on semantic segmentation masks produced by a pre-trained network. To avoid confusion, please describe this as weak supervision or image-level supervision from automatically generated labels rather than self-supervision in the strict sense.
  2. [Sec. 4.1, Eq. (18)] The notation in Eq. (18) uses S_i^g both for a set of vertices and for a surface; please clarify the distinction, for example by denoting the surface as a mesh and the vertex set separately, so that the symmetric error is unambiguous.
  3. [Sec. 4.1, Quantitative Comparison] The comparison with Alldieck et al. [3] is based on a re-trained model by overlapping authors. Please specify the exact training split, hyperparameters, and number of frames used for the baseline, and report per-garment errors in the main text rather than only in the supplementary material.
  4. [Sec. 1, Introduction] The claim of being 'the first model capable of inferring human body and layered garments on top as separate meshes from images directly' should be softened or qualified in light of existing multi-layer garment models such as ClothCap, even though those are not image-based; a more precise statement about the specific novelty would avoid an overclaim.
  5. [Sec. 5, Conclusion] The main text states that limitations and future work are discussed in the supplementary material, but the substantative limitations of the registration and the evaluation are not summarized in the main text; please add a brief limitations paragraph to the main paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: MGN's image-to-garment predictions are learned from registered ground truth and evaluated on held-out scans; the registration-bias concern is a validation issue, not a circular derivation.

full rationale

MGN's derivation chain is self-contained in the sense required here. The network is trained with 3D vertex losses (Eqs. 15-16) that compare predictions against registered garment meshes produced by the Sec. 3.1 pipeline, and it is evaluated on 70 held-out scans (Sec. 4.1). The garment PCA space and high-frequency displacement branch (Eq. 12) are learned from those registrations, but the regression from image-derived latent codes to PCA coefficients and displacements is not algebraically forced by the PCA construction; it is a fitted mapping tested on data not used for fitting. The 2D segmentation loss (Eq. 17) does compare rendered masks to the same semantic segmentations used as network input, but the paper explicitly frames this as self-supervision and test-time refinement, not as an independent ground-truth source, so it does not make the headline error 'derived from itself.' The skeptic's concern that registered garment meshes are used both as training labels and as the evaluation reference (Eq. 18) without an explicit check against raw scan surfaces is a legitimate validity and accuracy limitation, but it is not a circularity of the kind where an equation reduces to its own input by construction. Self-citations to Alldieck et al. [3,5], ClothCap [47], and SMPL [40] are used for baselines, initialization, and standard body modeling; none is invoked as an unverified uniqueness theorem or as the sole justification of the central image-to-garment mapping. Therefore no pattern from the circularity taxonomy applies, and the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The paper's central result rests on a data-processing chain (scan segmentation, template registration, PCA, CNN) rather than on a small number of physical axioms. The free parameters are hyperparameters of that chain, and the main domain assumption is that the fixed-template registration produces correct correspondences for all five garment categories; the method inherits any registration or segmentation error.

free parameters (3)
  • Number of PCA components per garment (35)
    Chosen by hand in Sec. 3.2, Garment Shape Space; sets the capacity of the pose-invariant garment shape basis and limits detail, and the paper acknowledges a smoothness bias.
  • Interpenetration penalty weight w = 25
    Constant in Eq. 8, chosen in experiments; controls how strongly garment vertices are pushed outside the body during registration.
  • High-frequency displacement cap = 1 cm
    Sec. 3.4 caps per-vertex displacements to 1 cm so that PCA explains the overall shape; this cap is a hand-set hyperparameter.
assumptions (6)
  • domain assumption SMPL body model accurately represents the naked body shape under clothing
    MGN layers garments on SMPL and extracts body shape under clothing via SMPL registration; if SMPL cannot express the true body, garment offsets are biased (Sec. 3.1, Eqs. 1-5).
  • domain assumption The 356 Twindom scans are accurate, complete, and sufficiently diverse
    Training and evaluation use these scans; errors in scanning or limited diversity restrict generality (Sec. 4).
  • ad hoc to paper A single fixed-topology template per garment category can be registered to all instances in that category
    The entire digital wardrobe and PCA shape space depend on this correspondence; dresses, skirts, and open garments are excluded (Sec. 3.1).
  • domain assumption Input semantic segmentation from a pretrained network [20] is reliable enough
    MGN discards RGB and uses segmentations plus 2D joints as input; poor segmentation bounds downstream garment prediction (Sec. 3.2).
  • domain assumption PCA with 35 components plus at most 1 cm displacements captures garment geometry sufficiently
    The garment decoder is PCA-based; the paper notes it is biased toward smooth results and cannot represent high-frequency detail (Sec. 3.2, 4.1).
  • domain assumption Fixed camera and subject turn-around setting
    Assumes fixed camera parameters and a person rotating in front of the camera; there is no in-the-wild evaluation (Sec. 4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-Garment Net: Learning to Dress 3D People from Images." pith.science (2026). https://pith.science/paper/IE2CJK5O

@misc{pith2026190806903,
  author       = {Pith},
  title        = {Pith review of: Multi-Garment Net: Learning to Dress 3D People from Images},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IE2CJK5O}},
  note         = {Machine review of arXiv:1908.06903}
}
read the original abstract

We present Multi-Garment Network (MGN), a method to predict body shape and clothing, layered on top of the SMPL model from a few frames (1-8) of a video. Several experiments demonstrate that this representation allows higher level of control when compared to single mesh or voxel representations of shape. Our model allows to predict garment geometry, relate it to the body shape, and transfer it to new body shapes and poses. To train MGN, we leverage a digital wardrobe containing 712 digital garments in correspondence, obtained with a novel method to register a set of clothing templates to a dataset of real 3D scans of people in different clothing and poses. Garments from the digital wardrobe, or predicted by MGN, can be used to dress any body shape in arbitrary poses. We will make publicly available the digital wardrobe, the MGN model, and code to dress SMPL with the garments.

Figures

Figures reproduced from arXiv: 1908.06903 by the authors.

Figure 1
Figure 1. Garment re-targeting with Multi-Garment Network [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of our approach. Given a small number of RGB frames (currently 8), we pre-compute semantically segmented images [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Digital 3D wardrobe. We use our proposed multi-mesh registration approach to register garments present in the scans (left) to [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Left to right: Scan, segmentation with MRF and CNN [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Dressing SMPL with just images. We use MGN to extract garments from the images of a source subject (middle) and use the [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison with Alldieck et al.[3]. In each set we visualize 3D predictions from [3](left) and our method (right) for five test subjects. Since our approach explicitly models garment geometry, it preserves more garment details, as is evident from minimal di…
Figure 7
Figure 7. Figure 7: Texture transfer. We model each garment class as a mesh with fixed topology and surface parameterization. This enables us to [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Garment re-targeting by MGN using 8 RGB images. In each of the three sets we show the source subject, target subject and [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 65 canonical work pages

  1. [3]

    Learning to re- construct people in clothing from a single RGB camera

    Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to re- construct people in clothing from a single RGB camera. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 1, 2, 5, 6, 7

  2. [1]

    https://virtualhumans.mpi-inf.mpg.de/mgn/. 1

  3. [2]

    An Efficient V olumetric Framework for Shape Track- ing

    Benjamin Allain, Jean-S ´ebastien Franco, and Edmond Boyer. An Efficient V olumetric Framework for Shape Track- ing. In IEEE Conf. on Computer Vision and Pattern Recog- nition, pages 268–276, Boston, United States, 2015. IEEE. 2

  4. [4]

    Detailed human avatars from monocular video

    Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Detailed human avatars from monocular video. In International Conf. on 3D Vision, sep 2018. 1

  5. [5]

    Video based reconstruction of 3D people models

    Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3D people models. InIEEE Conf. on Computer Vision and Pattern Recognition, 2018. 1, 2, 6

  6. [6]

    Tex2shape: Detailed full human body geometry from a single image

    Thiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, and Marcus Magnor. Tex2shape: Detailed full human body geometry from a single image. In IEEE International Con- ference on Computer Vision (ICCV). IEEE, oct 2019. 2

  7. [7]

    SCAPE: shape completion and animation of people

    Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Se- bastian Thrun, Jim Rodgers, and James Davis. SCAPE: shape completion and animation of people. InACM Transac- tions on Graphics, volume 24, pages 408–416. ACM, 2005. 2

  8. [8]

    Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image

    Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Pe- ter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conf. on Computer Vision. Springer International Publishing, 2016. 2

Show all 68 references
  1. [9]

    Deepgarment: 3d garment shape estimation from a single image

    R c, Endri Dibra, C ¨Oztireli, Remo Ziegler, and Markus Gross. Deepgarment: 3d garment shape estimation from a single image. In Computer Graphics Forum, volume 36, pages 269–280. Wiley Online Library, 2017. 2

  2. [10]

    Proba- bilistic deformable surface tracking from multiple videos

    Cedric Cagniart, Edmond Boyer, and Slobodan Ilic. Proba- bilistic deformable surface tracking from multiple videos. In Kostas Daniilidis, Petros Maragos, and Nikos Paragios, ed- itors, European Conf. on Computer Vision, volume 6314 of Lecture Notes in Computer Science, pages 3...

  3. [11]

    Free-viewpoint video of human actors

    Joel Carranza, Christian Theobalt, Marcus A Magnor, and Hans-Peter Seidel. Free-viewpoint video of human actors. In ACM Transactions on Graphics, volume 22, pages 569–

  4. [12]

    Deformable model for estimating clothed and naked hu- man shapes from a single image

    Xiaowu Chen, Yu Guo, Bin Zhou, and Qinping Zhao. Deformable model for estimating clothed and naked hu- man shapes from a single image. The Visual Computer , 29(11):1187–1196, 2013. 2

  5. [13]

    Garment modeling with a depth camera.ACM Transactions on Graphics, 34(6):203, 2015

    Xiaowu Chen, Bin Zhou, Feixiang Lu, Lin Wang, Lang Bi, and Ping Tan. Garment modeling with a depth camera.ACM Transactions on Graphics, 34(6):203, 2015. 2

  6. [14]

    High-quality streamable free-viewpoint video

    Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics, 34(4):69, 2015. 2

  7. [15]

    Geodesics in heat: A new approach to computing distance based on heat flow

    Keenan Crane, Clarisse Weischedel, and Max Wardetzky. Geodesics in heat: A new approach to computing distance based on heat flow. ACM Transactions on Graphics (TOG), 32(5):152, 2013. 3

  8. [16]

    Kinectavatar: fully automatic body capture using a single kinect

    Yan Cui, Will Chang, Tobias N ¨oll, and Didier Stricker. Kinectavatar: fully automatic body capture using a single kinect. In Asian Conf. on Computer Vision, pages 133–147,

  9. [17]

    Edilson de Aguiar, Leonid Sigal, Adrien Treuille, and Jes- sica K. Hodgins. Stable spaces for real-time clothing. ACM Trans. Graph., 29(4):106:1–106:9, July 2010. 3

  10. [18]

    Performance capture from sparse multi-view video

    Edilson De Aguiar, Carsten Stoll, Christian Theobalt, Naveed Ahmed, Hans-Peter Seidel, and Sebastian Thrun. Performance capture from sparse multi-view video. In ACM Transactions on Graphics, page 98, 2008. 2

  11. [19]

    Fusion4d: Real-time performance capture of chal- lenging scenes

    Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Tay- lor, et al. Fusion4d: Real-time performance capture of chal- lenging scenes. ACM Transactions on Graphics, 35(4):114,

  12. [20]

    Instance-level human parsing via part grouping network

    Ke Gong, Xiaodan Liang, Yicheng Li, Yimin Chen, Ming Yang, and Liang Lin. Instance-level human parsing via part grouping network. In European Conf. on Computer Vision,

  13. [21]

    P. Guan, L. Reiss, D. Hirshberg, A. Weiss, and M. J. Black. DRAPE: DRessing Any PErson. ACM Trans. on Graphics (Proc. SIGGRAPH), 31(4):35:1–35:10, July 2012. 3

  14. [22]

    Estimating human shape and pose from a single image

    Peng Guan, Alexander Weiss, Alexandru O B ˘alan, and Michael J Black. Estimating human shape and pose from a single image. In IEEE International Conf. on Computer Vision, pages 1381–1388. IEEE, 2009. 2

  15. [23]

    Garnet: A two-stream network for fast and accurate 3d cloth draping

    Erhan Gundogdu, Victor Constantin, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, and Pascal Fua. Garnet: A two-stream network for fast and accurate 3d cloth draping. arXiv preprint arXiv:1811.10983, 2018. 3

  16. [24]

    Clothed and naked human shapes estimation from a single image

    Yu Guo, Xiaowu Chen, Bin Zhou, and Qinping Zhao. Clothed and naked human shapes estimation from a single image. Computational Visual Media, pages 43–50, 2012. 2

  17. [25]

    Livecap: Real-time human performance capture from monocular video

    Marc Habermann, Weipeng Xu, , Michael Zollhoefer, Ger- ard Pons-Moll, and Christian Theobalt. Livecap: Real-time human performance capture from monocular video. ACM Transactions on Graphics, (Proc. SIGGRAPH), jul 2019. 1, 2

  18. [26]

    A statistical model of human pose and body shape

    Nils Hasler, Carsten Stoll, Martin Sunkel, Bodo Rosenhahn, and H-P Seidel. A statistical model of human pose and body shape. In Computer Graphics Forum, volume 28, pages 337– 346, 2009. 2

  19. [27]

    Learning to generate and reconstruct 3d meshes with only 2d supervision

    Paul Henderson and Vittorio Ferrari. Learning to generate and reconstruct 3d meshes with only 2d supervision. In British Machine Vision Conference (BMVC), 2018. 5

  20. [28]

    V olumetric 3d tracking by detection

    Chun-Hao Huang, Benjamin Allain, Jean-S ´ebastien Franco, Nassir Navab, Slobodan Ilic, and Edmond Boyer. V olumetric 3d tracking by detection. In IEEE Conf. on Computer Vision and Pattern Recognition, pages 3862–3870, 2016. 2 9

  21. [29]

    V olumedeform: Real-time volumetric non-rigid reconstruction

    Matthias Innmann, Michael Zollh ¨ofer, Matthias Nießner, Christian Theobalt, and Marc Stamminger. V olumedeform: Real-time volumetric non-rigid reconstruction. In European Conf. on Computer Vision, 2016. 2

  22. [30]

    Kinectfusion: real-time 3d reconstruction and inter- action using a moving depth camera

    Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Dustin Freeman, Andrew Davison, et al. Kinectfusion: real-time 3d reconstruction and inter- action using a moving depth camera. In ACM symposium on User in...

  23. [31]

    Moviereshape: Tracking and reshaping of humans in videos

    Arjun Jain, Thorsten Thorm ¨ahlen, Hans-Peter Seidel, and Christian Theobalt. Moviereshape: Tracking and reshaping of humans in videos. In ACM Transactions on Graphics , volume 29, page 148. ACM, 2010. 2

  24. [32]

    Total capture: A 3d deformation model for tracking faces, hands, and bod- ies

    Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3d deformation model for tracking faces, hands, and bod- ies. In IEEE Conf. on Computer Vision and Pattern Recog- nition, pages 8320–8329, 2018. 2

  25. [33]

    Black, David W

    Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In IEEE Conf. on Computer Vision and Pattern Recog- nition. IEEE Computer Society, 2018. 2

  26. [34]

    Doyub Kim, Woojong Koh, Rahul Narain, Kayvon Fa- tahalian, Adrien Treuille, and James F. O’Brien. Near- exhaustive precomputation of secondary cloth effects. ACM Transactions on Graphics , 32(4):87:1–7, July 2013. Pro- ceedings of ACM SIGGRAPH 2013, Anaheim. 3

  27. [35]

    Deepwrin- kles: Accurate and realistic clothing modeling

    Zorah Lahner, Daniel Cremers, and Tony Tung. Deepwrin- kles: Accurate and realistic clothing modeling. In Pro- ceedings of the European Conference on Computer Vision (ECCV), pages 667–684, 2018. 3

  28. [36]

    Christoph Lassner, Gerard Pons-Moll, and Peter V . Gehler. A generative model of people in clothing. In Proceedings IEEE International Conference on Computer Vision (ICCV), Piscataway, NJ, USA, oct 2017. IEEE. 2

  29. [37]

    Multi-View Dynamic Shape Refinement Using Local Tem- poral Integration

    Vincent Leroy, Jean-S ´ebastien Franco, and Edmond Boyer. Multi-View Dynamic Shape Refinement Using Local Tem- poral Integration. In IEEE International Conf. on Computer Vision, Venice, Italy, 2017. 2

  30. [38]

    3d self-portraits

    Hao Li, Etienne V ouga, Anton Gudym, Linjie Luo, Jonathan T Barron, and Gleb Gusev. 3d self-portraits. ACM Transactions on Graphics, 32(6):187, 2013. 2

  31. [39]

    An intriguing failing of convolutional neural networks and the coordconv solution

    Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski. An intriguing failing of convolutional neural networks and the coordconv solution. In Proceedings of the 32Nd Interna- tional Conference on Neural Information Processing...

  32. [40]

    SMPL: A skinned multi-person linear model

    Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J Black. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics, 34(6):248:1–248:16, 2015. 1, 2

  33. [41]

    Sic- lope: Silhouette-based clothed people

    Ryota Natsume, Shunsuke Saito, Zeng Huang, Weikai Chen, Chongyang Ma, Hao Li, and Shigeo Morishima. Sic- lope: Silhouette-based clothed people. arXiv preprint arXiv:1901.00049, 2018. 1

  34. [42]

    A layered model of human body and garment deformation

    Alexandros Neophytou and Adrian Hilton. A layered model of human body and garment deformation. In International Conference on 3D Vision, 2014. 3

  35. [43]

    Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time

    Richard A Newcombe, Dieter Fox, and Steven M Seitz. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. In IEEE Conf. on Computer Vision and Pattern Recognition, pages 343–352, 2015. 2

  36. [44]

    Neural body fitting: Unifying deep learning and model based human pose and shape esti- mation

    Mohamed Omran, Christop Lassner, Gerard Pons-Moll, Pe- ter Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape esti- mation. In International Conf. on 3D Vision, 2018. 2

  37. [45]

    Holoportation: Virtual 3d teleportation in real-time

    Sergio Orts-Escolano, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip L Davidson, Sameh Khamis, Mingsong Dou, et al. Holoportation: Virtual 3d teleportation in real-time. In Sym- posium on User Interface Software and Technology , ...

  38. [46]

    Learning to estimate 3D human pose and shape from a single color image

    Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou, and Kostas Daniilidis. Learning to estimate 3D human pose and shape from a single color image. InIEEE Conf. on Computer Vision and Pattern Recognition, 2018. 2

  39. [47]

    ClothCap: Seamless 4D clothing capture and retar- geting

    Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael Black. ClothCap: Seamless 4D clothing capture and retar- geting. ACM Transactions on Graphics, 36(4), 2017. 3, 4, 5

  40. [48]

    Dyna: a model of dynamic human shape in motion

    Gerard Pons-Moll, Javier Romero, Naureen Mahmood, and Michael J Black. Dyna: a model of dynamic human shape in motion. ACM Transactions on Graphics, 34:120, 2015. 2

  41. [49]

    Context-aware garment modeling from sketches

    Cody Robson, Ron Maharik, Alla Sheffer, and Nathan Carr. Context-aware garment modeling from sketches. Computers & Graphics, 35(3):604–613, 2011. 2

  42. [50]

    Garment replacement in monoc- ular video sequences

    Lorenz Rogge, Felix Klose, Michael Stengel, Martin Eise- mann, and Marcus Magnor. Garment replacement in monoc- ular video sequences. ACM Transactions on Graphics , 34(1):6, 2014. 2

  43. [51]

    Pifu: Pixel-aligned implicit function for high-resolution clothed human digitiza- tion

    Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Mor- ishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitiza- tion. arXiv preprint arXiv:1905.05172, 2019. 1

  44. [52]

    Otaduy, and Dan Casas

    Igor Santesteban, Miguel A. Otaduy, and Dan Casas. Learning-based animation of clothing for virtual try-on. In Computer Graphics Forum (Proc. Eurographics) , vol- ume 32, pages 1–8, 2019. 3

  45. [53]

    Rapid avatar capture and simulation using commodity depth sensors

    Ari Shapiro, Andrew Feng, Ruizhe Wang, Hao Li, Mark Bo- las, Gerard Medioni, and Evan Suma. Rapid avatar capture and simulation using commodity depth sensors. Computer Animation and Virtual Worlds, 25(3-4):201–211, 2014. 2

  46. [54]

    A perceptual control space for garment simulation

    Leonid Sigal, Moshe Mahler, Spencer Diaz, Kyna McIntosh, Elizabeth Carter, Timothy Richards, and Jessica Hodgins. A perceptual control space for garment simulation. In ACM Transactions on Graphics, 2015. 3

  47. [55]

    Killingfusion: Non-rigid 3d reconstruc- tion without correspondences

    Miroslava Slavcheva, Maximilian Baust, Daniel Cremers, and Slobodan Ilic. Killingfusion: Non-rigid 3d reconstruc- tion without correspondences. In IEEE Conf. on Computer Vision and Pattern Recognition, volume 3, page 7, 2017. 2

  48. [56]

    Laplacian mesh processing

    Olga Sorkine. Laplacian mesh processing. In Eurographics (STARs), pages 53–70, 2005. 4 10

  49. [57]

    Surface capture for performance-based animation

    Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE Computer Graphics and Applications, 27(3), 2007. 2

  50. [58]

    Pons-Moll, and Yebin Liu

    Yu Tao, Zerong Zheng, Kaiwen Guo, Jianhui Zhao, Dai Quionhai, Hao Li, G. Pons-Moll, and Yebin Liu. Double- fusion: Real-time capture of human performance with inner body shape from a depth sensor. InIEEE Conf. on Computer Vision and Pattern Recognition, 2018. 2

  51. [59]

    Simulcap : Single-view human performance capture with cloth simula- tion

    Yu Tao, Zerong Zheng, Yuan Zhong, Jianhui Zhao, Dai Quionhai, Gerard Pons-Moll, and Yebin Liu. Simulcap : Single-view human performance capture with cloth simula- tion. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), jun 2019. 2

  52. [60]

    Estimation of human body shape in motion with wide clothing

    Jinlong Yang, Jean-S ´ebastien Franco, Franck H ´etroy- Wheeler, and Stefanie Wuhrer. Estimation of human body shape in motion with wide clothing. In European Confer- ence on Computer Vision, 2016. 3

  53. [61]

    Analyzing clothing layer de- formation statistics of 3d human motions

    Jinlong Yang, Jean-S ´ebastien Franco, Franck H ´etroy- Wheeler, and Stefanie Wuhrer. Analyzing clothing layer de- formation statistics of 3d human motions. InEuropean Conf. on Computer Vision, pages 237–253, 2018. 3

  54. [62]

    Human appearance transfer

    Mihai Zanfir, Alin-Ionut Popa, Andrei Zanfir, and Cristian Sminchisescu. Human appearance transfer. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5391–5399, 2018. 2

  55. [63]

    Templateless quasi-rigid shape modeling with implicit loop- closure

    Ming Zeng, Jiaxiang Zheng, Xuan Cheng, and Xinguo Liu. Templateless quasi-rigid shape modeling with implicit loop- closure. In IEEE Conf. on Computer Vision and Pattern Recognition, pages 145–152, 2013. 2

  56. [64]

    Detailed, accurate, human shape estimation from clothed 3D scan sequences

    Chao Zhang, Sergi Pujades, Michael Black, and Gerard Pons-Moll. Detailed, accurate, human shape estimation from clothed 3D scan sequences. In IEEE Conf. on Computer Vi- sion and Pattern Recognition (CVPR), 2017. 3, 5

  57. [65]

    Garment modeling from a single image

    Bin Zhou, Xiaowu Chen, Qiang Fu, Kan Guo, and Ping Tan. Garment modeling from a single image. InComputer graph- ics forum, volume 32, pages 85–91. Wiley Online Library,

  58. [66]

    Color map optimization for 3d reconstruction with consumer depth cameras

    Qian-Yi Zhou and Vladlen Koltun. Color map optimization for 3d reconstruction with consumer depth cameras. ACM Transactions on Graphics, 33(4):155, 2014. 2

  59. [67]

    Parametric reshaping of human bodies in images

    Shizhe Zhou, Hongbo Fu, Ligang Liu, Daniel Cohen-Or, and Xiaoguang Han. Parametric reshaping of human bodies in images. In ACM Transactions on Graphics, volume 29, page

  60. [68]

    The stitched puppet: A graphical model of 3d human shape and pose

    Silvia Zuffi and Michael J Black. The stitched puppet: A graphical model of 3d human shape and pose. In IEEE Conf. on Computer Vision and Pattern Recognition , pages 3537–

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.