Pith. sign in

REVIEW 3 major objections 4 minor 78 references

360-Degree Textures of People in Clothing from a Single Image

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A single photograph can be turned into a complete, reposable 3D avatar with texture, clothing segmentation, and clothing geometry.

desk verdict A practical single-image-to-3D-avatar pipeline built on a strong registered-scan dataset, but the lack of quantitative evaluation and the unmeasured DensePose test-time risk keep it conditional. read the letter →

arxiv 1908.07117 v1 pith:53XHTTOQ submitted 2019-08-20 cs.CV cs.GR

classification cs.CVcs.GR
keywords 3Dhumanreconstructionsingle-imageavatartexturecompletionSMPLUVspaceimage-to-imagetranslationgarmentsegmentationdisplacementmappredictionclothingediting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that a single ordinary photo of a person contains enough information to build a full 3D avatar that can be rotated, reposed, and edited. It predicts the appearance, clothing layout, and clothing geometry of the unseen side of the body by working on a flat map of a standard body model rather than directly in 3D. The authors' trick is to register thousands of 3D scans of clothed people to that body model, so that all the needed information becomes ordinary 2D images that image-to-image networks can complete. If the paper is right, one smartphone photo could produce an animatable avatar for games, virtual try-on, or augmented reality.

What carries the argument

The machinery is the SMPL model's fixed UV parameterization: SMPL is a skinned human body model with a fixed mesh topology, so every point on the body surface maps to a fixed pixel location in a 2D texture atlas. The paper non-rigidly registers SMPL to 4,541 clothed 3D scans, which turns texture, clothing labels, and geometry offsets into aligned 2D images. Three separate image-to-image translation networks then operate on those images: one completes the partial texture, one completes the partial clothing segmentation, and one predicts a displacement map from the completed segmentation. This last map, applied as vertex offsets to SMPL, gives the clothing geometry.

What would settle it

Take people photographed from a single frontal view while also capturing full 3D scans, then measure the error of the predicted texture, segmentation, and displacement maps on the occluded back half; if the back-half predictions are no better than a blurred average of training textures, the central claim fails.

Watch

Extended reading notes

Core claim

The central claim is that a complete, animatable 3D human model can be computed from one image by predicting three aligned maps on the SMPL body model's UV atlas (a fixed 2D layout of the body surface): a full texture map, a full clothing segmentation map, and a displacement map encoding how far the clothing sticks out from the body. A non-rigid registration step warps SMPL to 4,541 scans of clothed people, producing training images in exact correspondence; at inference, DensePose (a method that maps image pixels to body-surface locations) extracts a partial texture and partial segmentation from the input view, and three image-to-image translation networks complete them. The completed displacement map is applied as vertex offsets to SMPL to give clothing geometry, and the completed texture is draped over the result. Experiments on rendered scans and on real images from DeepFashion and People Snapshot show plausible avatars, and the segmentation map enables garment swapping and garment-length editing.

Load-bearing premise

The method assumes the partial texture and segmentation maps extracted by DensePose from the input photo are aligned well enough for the completion networks, which were trained only on clean registered scan textures, to still work.

Editorial extensions

If this is right

  • A single photograph suffices to produce a fully textured, reposable 3D avatar, removing the need for multi-view capture or video.
  • Because texture, geometry, and segmentation are predicted in a fixed UV space, the avatar can be reposed and reshaped with the underlying body model, avoiding the pose generalization problems of image-based reposing.
  • The completed segmentation map gives hands-on control: garment textures can be swapped, sleeve and pant lengths edited, and garments transferred between subjects.
  • The pipeline is limited by DensePose's accuracy: when DensePose misaligns on hair or clothing, the partial input maps are distorted and reconstruction quality drops.
  • The recovered 3D model supports clothing-aware editing that image-based methods cannot do coherently.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Implication the authors leave implicit: because the three maps are predicted separately, each factor can be improved or retrained with different data (geometry scans, internet photos) without rebuilding the whole pipeline.
  • Testable extension: train the completion networks on DensePose-derived partial maps as well as on clean registered textures, which would directly test whether the train/test alignment gap is the main source of real-image error.
  • Neighboring application: the same UV-space completion scheme should transfer to any articulated object that has a canonical template and registered scans, such as animals, provided a DensePose-like correspondence is available.
  • Boundary the paper names but does not pursue: the fixed SMPL topology cannot represent garments with different topology (skirts, dresses); implicit-function reconstructions could cover those cases, but would lose the direct editing control that the segmentation map gives.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a method to predict a complete 3D avatar of a person from a single RGB image. The approach first registers thousands of 3D scans of people in clothing to the SMPL model, encoding appearance, garment layout, and geometry as images in a common UV space. At test time, DensePose and a garment segmenter produce partial texture and segmentation maps, which are completed by separate image-to-image translation networks. A third network predicts a displacement map from the completed segmentation. The predicted maps are applied to SMPL to obtain a fully textured, reposable, and editable avatar. The paper describes the dataset, registration procedure, network objectives, and qualitative results on People Snapshot and DeepFashion, and demonstrates applications such as garment swapping and garment-length editing.

Significance. If the results hold, the paper offers a practical and elegant reduction of single-image 3D human avatar reconstruction to image-to-image translation in UV space. The main strengths are the new registered scan dataset (360° People Textures), the clean decomposition into texture, segmentation, and displacement maps, and the qualitative demonstrations of reposing, garment transfer, and garment editing. The paper also states that source code will be released. The central technical risk is the reliance on DensePose alignment at test time and the lack of quantitative evaluation, which limits confidence in the generalization claim and in the reproducibility of the registration pipeline.

major comments (3)
  1. [§4 (Experiments)] The paper states that quantitative evaluation was omitted because 'the typical metrics for evaluating GAN models known in the literature did not correspond to the human-perceived quality of the generated texture.' This is not sufficient for a paper whose central claim is that a single photo yields a complete, reposable avatar. The authors have 10 held-out registered scans and a separate subset of 2056 scans with ground-truth segmentation and displacement maps; they should report per-pixel texture error (L1, SSIM), segmentation IoU, and vertex/displacement error on these held-out scans, even if these numbers are not the primary quality measure. A small user study would also substantiate the qualitative claims.
  2. [§3.1.1, Eq. (7)] The vertex weights w_i in the registration objective are introduced only qualitatively ('penalize deviations from the model more heavily for the vertices on the hands and the feet'), and the prior weights λθ, λβ in Eqs. (5)–(7) are not specified anywhere. Since the registrations produced by Eq. (7) define the ground-truth texture, segmentation, and displacement maps used for training, the choice of these weights directly influences every result in the paper. Please specify the values (or a table of them) and provide a sensitivity analysis, or at least state that results are robust to a reasonable range.
  3. [§2 and §3.2.1] The method's test-time behavior depends on DensePose and the garment segmenter producing partial UV maps that are well aligned with the scan-derived training distribution. The authors themselves note that DensePose 'was not designed to accommodate for hair and clothing deviating from the body' and that 'significant miss-alignments might occur.' Because the training partial maps come from rendered scans rather than unconstrained photographs, the magnitude of test-time DensePose errors is unknown. This is a load-bearing assumption for the claimed generalization to real images; the paper should quantify it, for example by perturbing DensePose correspondences at test time and measuring degradation, or by collecting a small annotated set of real images with manual UV correspondences.
minor comments (4)
  1. [Various] Minor typographical and grammatical errors: 'We train and our method on our newly created dataset' (§1) should read 'We train our method on our newly created dataset'; 'it’s garment segmentation' (§3.2) should be 'its'; 'miss-alignments' (§2) should be 'misalignments'; 'suplementary' (§4) should be 'supplementary'.
  2. [Eq. (7)] The Geman-McClure cost ρ is named but its explicit form is not given; please provide the formula or a precise reference so that the registration objective is fully reproducible.
  3. [§4.1] The statement that GAN metrics 'did not correspond to the human-perceived quality' would be more convincing if the paper listed which metrics were tried and in what way they failed, particularly because the lack of quantitative results is a salient departure from standard evaluation practice.
  4. [Figure 4] The qualitative results are the main evidence for the central claim, but the figure images are small; larger crops or side-by-side zooms would help the reader verify texture completion, segmentation boundaries, and garment geometry.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: texture, segmentation, and displacement predictions are supervised by independent 3D-scan registrations, and the cited self-works are standard body-model/registration references rather than definitional inputs.

full rationale

The paper's derivation chain is: register SMPL to 4541 scans to obtain complete texture, segmentation, and displacement ground truth; render scan views and run DensePose to create partial maps; train image-to-image networks to map partial maps to complete maps; at test time apply the same DensePose extraction and run the trained networks. No output is used to define its own input. The targets are independent: 'we use DensePose solely to create the partial texture map from the input view, but train from high quality aligned and complete texture maps, which are obtained by registering SMPL to 3D scans.' The reconstruction and adversarial losses (Eqs. 8-14) compare network outputs to registered ground-truth maps, so the 'predictions' are not fitted parameters renamed as outputs. Self-citations to SMPL [41] and earlier registration work [16,53] are standard model usage that does not smuggle in the target result; SMPL is an external body model, and the registration equations (5)-(7) are supervised by scan geometry rather than by the final avatar. The statement that DensePose 'was not designed to accommodate for hair and clothing deviating from the body' is a limitation for test-time generalization, not a circular step. The paper's lack of quantitative evaluation is a rigor concern, but it does not make the derivation definitionally equivalent to its inputs.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central contribution is a learned pipeline. The free parameters are the loss and registration weights that shape the outputs, none of which are pinned down in the main text. The main domain assumptions are the quality of the UV-space registration, DensePose at test time, and the sufficiency of segmentation for geometry. No new physical entities are introduced.

free parameters (3)
  • Loss weights λ1-λ4 for texture, segmentation, and displacement objectives = not specified in main text
    The losses in Equations 8, 13, 14 are weighted sums with λ1..λ4; the values are not given in the paper and are presumably tuned in supplementary material. The exact balance of GAN, reconstruction, perceptual, and SSIM terms affects the output quality.
  • Pose and shape prior weights λθ, λβ = not specified
    Used in the registration objective Equations 5-7; values are not provided in the main text, though the paper says the priors regularize pose and shape.
  • Registration vertex weights wi = not specified
    In Equation 7, wi weights penalize deviations from the model more heavily for hands and feet; exact values are not given, so a replication would need to choose them.
assumptions (5)
  • standard math SMPL body model provides a common UV parametrization of the human body
    The entire learning is done in the UV-space of SMPL (Section 3.1.1), relying on the model's shape and pose space from Loper et al. 2015.
  • domain assumption Non-rigid registration of SMPL to 4541 scans produces consistent correspondences across clothing and hair
    The training data quality depends on the registration objective in Equation 7 producing well-aligned textures and displacement maps across diverse people; no quantitative verification of registration accuracy is provided.
  • domain assumption DensePose maps test images to SMPL surface with sufficient accuracy for partial map extraction
    DensePose is used to build partial texture and segmentation maps at test time (Section 3.2.1); the paper acknowledges DensePose can miss-align for hair/clothing deviating from the body.
  • domain assumption Clothing segmentation from [25] provides reliable ground-truth labels
    Full segmentation ground truth is obtained by running the method of Gong et al. [25] and stitching projections; errors in that method propagate to training labels.
  • domain assumption Segmentation alone is sufficient to predict plausible clothing geometry
    Displacement prediction in Section 3.2.3 is conditioned only on the completed segmentation; the paper admits ambiguity between tight and loose clothing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of 360-Degree Textures of People in Clothing from a Single Image." pith.science (2026). https://pith.science/paper/53XHTTOQ

@misc{pith2026190807117,
  author       = {Pith},
  title        = {Pith review of: 360-Degree Textures of People in Clothing from a Single Image},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/53XHTTOQ}},
  note         = {Machine review of arXiv:1908.07117}
}
read the original abstract

In this paper we predict a full 3D avatar of a person from a single image. We infer texture and geometry in the UV-space of the SMPL model using an image-to-image translation method. Given partial texture and segmentation layout maps derived from the input view, our model predicts the complete segmentation map, the complete texture map, and a displacement map. The predicted maps can be applied to the SMPL model in order to naturally generalize to novel poses, shapes, and even new clothing. In order to learn our model in a common UV-space, we non-rigidly register the SMPL model to thousands of 3D scans, effectively encoding textures and geometries as images in correspondence. This turns a difficult 3D inference task into a simpler image-to-image translation one. Results on rendered scans of people and images from the DeepFashion dataset demonstrate that our method can reconstruct plausible 3D avatars from a single image. We further use our model to digitally change pose, shape, swap garments between people and edit clothing. To encourage research in this direction we will make the source code available for research purpose.

Figures

Figures reproduced from arXiv: 1908.07117 by the authors.

Figure 1
Figure 1. Given a single view of a person we predict a complete texture map in the UV space, complete clothing segmentation [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. 360◦ People Textures: Example registrations and texture maps used to train our model. After registering SMPL to the scans their appearances are encoded as texture maps in a common UV-space. a description [78]. Other works demonstrated models ca￾pable of swapping the appearances between two different subjects [54, 75]. Multi-view texture generation. Texture generation is chal￾lenging even in the case of multi-view im… view at source ↗
Figure 3
Figure 3. Registration: we bring all scan appearances into [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Generalization of our method to real images. Left to right: real image, segmented image, complete texture map, [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Comparison of our single view method (bottom row) to the monocular video method of [ [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Garment re-targeting results: Four subjects are [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 9
Figure 9. Figure 9: Image-based reposing methods cannot handle [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 8
Figure 8. Figure 8: Model editing results: Models reconstructed from [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 71 canonical work pages

  1. [1]

    https://renderpeople.com/. 5

  2. [2]

    https://secure.axyz-design.com/. 5

  3. [3]

    https://web.twindom.com/. 5

  4. [4]

    https://www.treedys.com/. 5

  5. [5]

    http://virtualhumans.mpi-inf.mpg.de/360tex. 1

  6. [6]

    Alldieck, M

    T. Alldieck, M. Magnor, B. L. Bhatnagar, C. Theobalt, and G. Pons-Moll. Learning to reconstruct people in clothing from a single RGB camera. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), jun 2019. 3, 5, 7

  7. [7]

    Alldieck, M

    T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons- Moll. Detailed human avatars from monocular video. In International Conference on 3D Vision (3DV), sep 2018. 3

  8. [8]

    Alldieck, M

    T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons- Moll. Video based reconstruction of 3d people models. In IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2018. 3, 5

Show all 78 references
  1. [9]

    Alldieck, G

    T. Alldieck, G. Pons-Moll, C. Theobalt, and M. Magnor. Tex2shape: Detailed full human body geometry from a sin- gle image. arXiv preprint arXiv:1904.08645, 2019. 5

  2. [10]

    Arjovsky, S

    M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein gan. arXiv preprint arXiv:1701.07875, 2017. 4

  3. [11]

    Balakrishnan, A

    G. Balakrishnan, A. Zhao, A. V . Dalca, F. Durand, and J. Guttag. Synthesizing images of humans in unseen poses. arXiv preprint arXiv:1804.07739, 2018. 2

  4. [12]

    Baumberg

    A. Baumberg. Blending images for texturing 3d models. In British Machine Vision Conference, volume 3, page 5. Cite- seer, 2002. 3

  5. [13]

    Bernardini, I

    F. Bernardini, I. M. Martin, and H. Rushmeier. High-quality texture reconstruction from multiple scans. IEEE Transac- tions on Visualization and Computer Graphics , 7(4):318– 332, 2001. 3

  6. [14]

    S. Bi, N. K. Kalantari, and R. Ramamoorthi. Patch-based optimization for image-based texture mapping. ACM Trans- actions on Graphics, 36(4), 2017. 3

  7. [15]

    F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Euro- pean Conference on Computer Vision. Springer International Publishing, Oct. 2016. 5

  8. [16]

    F. Bogo, J. Romero, G. Pons-Moll, and M. J. Black. Dy- namic FAUST: Registering human bodies in motion. InIEEE Conf. on Computer Vision and Pattern Recognition, 2017. 3

  9. [17]

    Z. Cao, T. Simon, S.-E. Wei, and Y . Sheikh. Realtime multi- person 2d pose estimation using part affinity fields. In IEEE Conf. on Computer Vision and Pattern Recognition, 2017. 3

  10. [18]

    C. Chan, S. Ginosar, T. Zhou, and A. A. Efros. Everybody dance now. arXiv preprint arXiv:1808.07371, 2018. 2, 8

  11. [19]

    Y . Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo. Stargan: Unified generative adversarial networks for multi- domain image-to-image translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 8789–8797, 2018. 4

  12. [20]

    P. E. Debevec, C. J. Taylor, and J. Malik. Modeling and ren- dering architecture from photographs: A hybrid geometry- and image-based approach. In Annual Conf. on Computer Graphics and Interactive Techniques , pages 11–20. ACM,

  13. [21]

    Eisemann, B

    M. Eisemann, B. De Decker, M. Magnor, P. Bekaert, E. De Aguiar, N. Ahmed, C. Theobalt, and A. Sellent. Float- ing textures. Computer Graphics Forum , 27(2):409–418,

  14. [22]

    Esser, E

    P. Esser, E. Sutter, and B. Ommer. A variational u-net for conditional appearance and shape generation. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8857–8866, 2018. 2

  15. [23]

    Y . Fu, Q. Yan, L. Yang, J. Liao, and C. Xiao. Texture map- ping for 3d reconstruction with rgb-d sensor. In IEEE Conf. on Computer Vision and Pattern Recognition, 2018. 3

  16. [24]

    L. A. Gatys, A. S. Ecker, and M. Bethge. Image style trans- fer using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 2414–2423, 2016. 8

  17. [25]

    K. Gong, X. Liang, Y . Li, Y . Chen, M. Yang, and L. Lin. Instance-level human parsing via part grouping network. In Proceedings of the European Conference on Computer Vi- sion (ECCV), pages 770–785, 2018. 5

  18. [26]

    Grigorev, A

    A. Grigorev, A. Sevastopolsky, A. Vakhitov, and V . Lem- pitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. 2, 3

  19. [27]

    R. A. G ¨uler, N. Neverova, and I. Kokkinos. Densepose: Dense human pose estimation in the wild. In IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  20. [28]

    Gulrajani, F

    I. Gulrajani, F. Ahmed, M. Arjovsky, V . Dumoulin, and A. C. Courville. Improved training of wasserstein gans. In Advances in Neural Information Processing Systems , pages 5767–5777, 2017. 4

  21. [29]

    Hamada, K

    K. Hamada, K. Tachibana, T. Li, H. Honda, and Y . Uchida. Full-body high-resolution anime generation with progressive structure-conditional generative adversarial networks. arXiv preprint arXiv:1809.01890, 2018. 2

  22. [30]

    X. Han, Z. Wu, Z. Wu, R. Yu, and L. S. Davis. Viton: An image-based virtual try-on network. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2018. 2, 8

  23. [31]

    Huang, T

    Z. Huang, T. Li, W. Chen, Y . Zhao, J. Xing, C. LeGendre, L. Luo, C. Ma, and H. Li. Deep volumetric video from very sparse multi-view performance capture. In Proceedings of the European Conference on Computer Vision (ECCV) , pages 336–354, 2018. 8

  24. [32]

    Huynh, W

    L. Huynh, W. Chen, S. Saito, J. Xing, K. Nagano, A. Jones, P. Debevec, and H. Li. Mesoscopic facial geometry infer- ence using deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, pages 8407–8416, 2018. 3

  25. [33]

    Isola, J.-Y

    P. Isola, J.-Y . Zhu, T. Zhou, and A. A. Efros. Image-to- image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 2, 4

  26. [34]

    Johnson, A

    J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision , pages 694–711. Springer,

  27. [35]

    Kanazawa, M

    A. Kanazawa, M. J. Black, D. W. Jacobs, and J. Malik. End- to-end recovery of human shape and pose. In Computer Vi- sion and Pattern Regognition (CVPR), 2018. 5

  28. [36]

    Karras, T

    T. Karras, T. Aila, S. Laine, and J. Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2017. 2

  29. [37]

    Lassner, G

    C. Lassner, G. Pons-Moll, and P. V . Gehler. A generative model of people in clothing. In Proceedings IEEE Interna- tional Conference on Computer Vision (ICCV) , Piscataway, NJ, USA, oct 2017. IEEE. 2

  30. [38]

    Lempitsky and D

    V . Lempitsky and D. Ivanov. Seamless mosaicing of image- based texture maps. In IEEE Conf. on Computer Vision and Pattern Recognition, pages 1–6. IEEE, 2007. 3

  31. [39]

    H. P. Lensch, W. Heidrich, and H.-P. Seidel. A silhouette- based algorithm for texture registration and stitching.Graph- ical Models, 63(4):245–262, 2001. 3

  32. [40]

    Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang. Deepfash- ion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1096–1104,

  33. [41]

    M. M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics, 34(6):248:1–248:16, Oct. 2015. 1, 2, 4

  34. [42]

    Lorenz, L

    D. Lorenz, L. Bereska, T. Milbich, and B. Ommer. Unsuper- vised part-based disentangling of object shape and appear- ance. arXiv preprint arXiv:1903.06946, 2019. 2

  35. [43]

    L. Ma, X. Jia, Q. Sun, B. Schiele, T. Tuytelaars, and L. Van Gool. Pose guided person image generation. In Advances in Neural Information Processing Systems , pages 406–416, 2017. 2, 8

  36. [44]

    L. Ma, Q. Sun, S. Georgoulis, L. Van Gool, B. Schiele, and M. Fritz. Disentangled person image generation. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 99–108, 2018. 2, 8

  37. [45]

    Martin-Brualla, R

    R. Martin-Brualla, R. Pandey, S. Yang, P. Pidlypenskyi, J. Taylor, J. Valentin, S. Khamis, P. Davidson, A. Tkach, P. Lincoln, et al. Lookingood: enhancing performance capture with real-time neural re-rendering. arXiv preprint arXiv:1811.05029, 2018. 8

  38. [46]

    Mescheder, M

    L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger. Occupancy networks: Learning 3d reconstruction in function space. arXiv preprint arXiv:1812.03828, 2018. 8

  39. [47]

    Nagano, J

    K. Nagano, J. Seo, J. Xing, L. Wei, Z. Li, S. Saito, A. Agar- wal, J. Fursund, H. Li, R. Roberts, et al. pagan: real-time avatars using dynamic textures. In SIGGRAPH Asia 2018 Technical Papers, page 258. ACM, 2018. 3

  40. [48]

    Natsume, S

    R. Natsume, S. Saito, Z. Huang, W. Chen, C. Ma, H. Li, and S. Morishima. Siclope: Silhouette-based clothed people. arXiv preprint arXiv:1901.00049, 2018. 3, 8

  41. [49]

    Neverova, R

    N. Neverova, R. A. G¨uler, and I. Kokkinos. Dense pose trans- fer. arXiv preprint arXiv:1809.01995, 2018. 2, 3

  42. [50]

    Niem and J

    W. Niem and J. Wingbermuhle. Automatic reconstruction of 3d objects using a mobile monoscopic camera. In Proc. Inter. Conf. on Recent Advances in 3-D Digital Imaging and Modeling, pages 173–180. IEEE, 1997. 3

  43. [51]

    E. Ofek, E. Shilat, A. Rappoport, and M. Werman. Mul- tiresolution textures from image sequences. IEEE Computer Graphics and Applications, 17(2):18–29, Mar. 1997. 3

  44. [52]

    Omran, C

    M. Omran, C. Lassner, G. Pons-Moll, P. Gehler, and B. Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape estimation. In 2018 In- ternational Conference on 3D Vision (3DV), pages 484–494. IEEE, 2018. 5

  45. [53]

    Pons-Moll, J

    G. Pons-Moll, J. Romero, N. Mahmood, and M. J. Black. Dyna: a model of dynamic human shape in motion. ACM Transactions on Graphics, 34:120, 2015. 3

  46. [54]

    A. Raj, P. Sangkloy, H. Chang, J. Hays, D. Ceylan, and J. Lu. Swapnet: Image based garment transfer. In European Conference on Computer Vision , pages 679–695. Springer, Cham, 2018. 3

  47. [55]

    Rocchini, P

    C. Rocchini, P. Cignoni, C. Montani, and R. Scopigno. Mul- tiple textures stitching and blending on 3d objects. In Ren- dering Techniques 99, pages 119–130. Springer, 1999. 3

  48. [56]

    Saito, Z

    S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. arXiv preprint arXiv:1905.05172, 2019. 3, 8

  49. [57]

    M. Sela, E. Richardson, and R. Kimmel. Unrestricted fa- cial geometry reconstruction using image-to-image transla- tion. In Proceedings of the IEEE International Conference on Computer Vision, pages 1576–1585, 2017. 3

  50. [58]

    Shysheya, E

    A. Shysheya, E. Zakharov, K.-A. Aliev, R. Bashirov, E. Burkov, K. Iskakov, A. Ivakhnenko, Y . Malkov, I. Pasech- nik, D. Ulyanov, et al. Textured neural avatars. arXiv preprint arXiv:1905.08776, 2019. 2

  51. [59]

    C. Si, W. Wang, L. Wang, and T. Tan. Multistage adversarial losses for pose-based human image synthesis. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 118–126, 2018. 2

  52. [60]

    Siarohin, E

    A. Siarohin, E. Sangineto, S. Lathuili `ere, and N. Sebe. Deformable gans for pose-based human image generation. In CVPR 2018-Computer Vision and Pattern Recognition ,

  53. [61]

    Starck and A

    J. Starck and A. Hilton. Model-based multiple view recon- struction of people. In IEEE international conference on computer vision, pages 915–922, 2003. 3

  54. [62]

    Tewari, F

    A. Tewari, F. Bernard, P. Garrido, G. Bharaj, M. Elgharib, H.-P. Seidel, P. P ´erez, M. Zollh ¨ofer, and C. Theobalt. Fml: Face model learning from videos. arXiv preprint arXiv:1812.07603, 2018. 3

  55. [63]

    Tewari, M

    A. Tewari, M. Zollh ¨ofer, P. Garrido, F. Bernard, H. Kim, P. P´erez, and C. Theobalt. Self-supervised multi-level face model learning for monocular reconstruction at over 250 hz. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2549–2559...

  56. [64]

    Varol, J

    G. Varol, J. Romero, X. Martin, N. Mahmood, M. J. Black, I. Laptev, and C. Schmid. Learning from synthetic humans. In CVPR, 2017. 5

  57. [65]

    Waechter, N

    M. Waechter, N. Moehrle, and M. Goesele. Let there be color! large-scale texturing of 3d reconstructions. In Euro- pean Conf. on Computer Vision , pages 836–850. Springer,

  58. [66]

    B. Wang, H. Zheng, X. Liang, Y . Chen, L. Lin, and M. Yang. Toward characteristic-preserving image-based virtual try-on network. In ECCV, 2018. 2, 8

  59. [67]

    Wang, M.-Y

    T.-C. Wang, M.-Y . Liu, J.-Y . Zhu, A. Tao, J. Kautz, and B. Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 8798–8807, 2018. 4

  60. [68]

    T. Y . Wang, D. Cylan, J. Popovic, and N. J. Mitra. Learning a shared shape space for multimodal garment design. arXiv preprint arXiv:1806.11335, 2018. 3

  61. [69]

    Z. Wang, A. C. Bovik, H. R. Sheikh, E. P. Simoncelli, et al. Image quality assessment: from error visibility to struc- tural similarity. IEEE transactions on image processing , 13(4):600–612, 2004. 4

  62. [70]

    Z. Wang, E. P. Simoncelli, and A. C. Bovik. Multiscale struc- tural similarity for image quality assessment. In The Thrity- Seventh Asilomar Conference on Signals, Systems & Com- puters, 2003, volume 2, pages 1398–1402. Ieee, 2003. 4

  63. [71]

    Q. Xu, W. Wang, D. Ceylan, R. Mech, and U. Neumann. Disn: Deep implicit surface network for high-quality single- view 3d reconstruction. arXiv preprint arXiv:1905.10711 ,

  64. [72]

    Yamaguchi, S

    S. Yamaguchi, S. Saito, K. Nagano, Y . Zhao, W. Chen, K. Olszewski, S. Morishima, and H. Li. High-fidelity facial reflectance and geometry inference from an unconstrained image. ACM Transactions on Graphics (TOG) , 37(4):162,

  65. [73]

    C. Yang, Z. Wang, X. Zhu, C. Huang, J. Shi, and D. Lin. Pose guided human video generation. In ECCV, 2018. 2

  66. [74]

    Yang, J.-S

    J. Yang, J.-S. Franco, F. H´etroy-Wheeler, and S. Wuhrer. An- alyzing clothing layer deformation statistics of 3d human motions. In Proceedings of the European Conference on Computer Vision (ECCV), pages 237–253, 2018. 3

  67. [75]

    Zanfir, A.-I

    M. Zanfir, A.-I. Popa, A. Zanfir, and C. Sminchisescu. Hu- man appearance transfer. In The IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), June 2018. 3

  68. [76]

    Zhou and V

    Q.-Y . Zhou and V . Koltun. Color map optimization for 3d reconstruction with consumer depth cameras. ACM Trans- actions on Graphics, 33(4):155, 2014. 3

  69. [77]

    J.-Y . Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image- to-image translation using cycle-consistent adversarial net- works. In Proceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017. 4

  70. [78]

    S. Zhu, S. Fidler, R. Urtasun, D. Lin, and C. Change Loy. Be your own prada: Fashion synthesis with structural co- herence. In International Conference on Computer Vision (ICCV), 2017. 2

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.