REVIEW 3 major objections 4 minor 78 references
360-Degree Textures of People in Clothing from a Single Image
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A single photograph can be turned into a complete, reposable 3D avatar with texture, clothing segmentation, and clothing geometry.
desk verdict A practical single-image-to-3D-avatar pipeline built on a strong registered-scan dataset, but the lack of quantitative evaluation and the unmeasured DensePose test-time risk keep it conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the SMPL model's fixed UV parameterization: SMPL is a skinned human body model with a fixed mesh topology, so every point on the body surface maps to a fixed pixel location in a 2D texture atlas. The paper non-rigidly registers SMPL to 4,541 clothed 3D scans, which turns texture, clothing labels, and geometry offsets into aligned 2D images. Three separate image-to-image translation networks then operate on those images: one completes the partial texture, one completes the partial clothing segmentation, and one predicts a displacement map from the completed segmentation. This last map, applied as vertex offsets to SMPL, gives the clothing geometry.
What would settle it
Take people photographed from a single frontal view while also capturing full 3D scans, then measure the error of the predicted texture, segmentation, and displacement maps on the occluded back half; if the back-half predictions are no better than a blurred average of training textures, the central claim fails.
Extended reading notes
Core claim
The central claim is that a complete, animatable 3D human model can be computed from one image by predicting three aligned maps on the SMPL body model's UV atlas (a fixed 2D layout of the body surface): a full texture map, a full clothing segmentation map, and a displacement map encoding how far the clothing sticks out from the body. A non-rigid registration step warps SMPL to 4,541 scans of clothed people, producing training images in exact correspondence; at inference, DensePose (a method that maps image pixels to body-surface locations) extracts a partial texture and partial segmentation from the input view, and three image-to-image translation networks complete them. The completed displacement map is applied as vertex offsets to SMPL to give clothing geometry, and the completed texture is draped over the result. Experiments on rendered scans and on real images from DeepFashion and People Snapshot show plausible avatars, and the segmentation map enables garment swapping and garment-length editing.
Load-bearing premise
The method assumes the partial texture and segmentation maps extracted by DensePose from the input photo are aligned well enough for the completion networks, which were trained only on clean registered scan textures, to still work.
Editorial extensions
If this is right
- A single photograph suffices to produce a fully textured, reposable 3D avatar, removing the need for multi-view capture or video.
- Because texture, geometry, and segmentation are predicted in a fixed UV space, the avatar can be reposed and reshaped with the underlying body model, avoiding the pose generalization problems of image-based reposing.
- The completed segmentation map gives hands-on control: garment textures can be swapped, sleeve and pant lengths edited, and garments transferred between subjects.
- The pipeline is limited by DensePose's accuracy: when DensePose misaligns on hair or clothing, the partial input maps are distorted and reconstruction quality drops.
- The recovered 3D model supports clothing-aware editing that image-based methods cannot do coherently.
Reading between the lines
- Implication the authors leave implicit: because the three maps are predicted separately, each factor can be improved or retrained with different data (geometry scans, internet photos) without rebuilding the whole pipeline.
- Testable extension: train the completion networks on DensePose-derived partial maps as well as on clean registered textures, which would directly test whether the train/test alignment gap is the main source of real-image error.
- Neighboring application: the same UV-space completion scheme should transfer to any articulated object that has a canonical template and registered scans, such as animals, provided a DensePose-like correspondence is available.
- Boundary the paper names but does not pursue: the fixed SMPL topology cannot represent garments with different topology (skirts, dresses); implicit-function reconstructions could cover those cases, but would lose the direct editing control that the segmentation map gives.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a method to predict a complete 3D avatar of a person from a single RGB image. The approach first registers thousands of 3D scans of people in clothing to the SMPL model, encoding appearance, garment layout, and geometry as images in a common UV space. At test time, DensePose and a garment segmenter produce partial texture and segmentation maps, which are completed by separate image-to-image translation networks. A third network predicts a displacement map from the completed segmentation. The predicted maps are applied to SMPL to obtain a fully textured, reposable, and editable avatar. The paper describes the dataset, registration procedure, network objectives, and qualitative results on People Snapshot and DeepFashion, and demonstrates applications such as garment swapping and garment-length editing.
Significance. If the results hold, the paper offers a practical and elegant reduction of single-image 3D human avatar reconstruction to image-to-image translation in UV space. The main strengths are the new registered scan dataset (360° People Textures), the clean decomposition into texture, segmentation, and displacement maps, and the qualitative demonstrations of reposing, garment transfer, and garment editing. The paper also states that source code will be released. The central technical risk is the reliance on DensePose alignment at test time and the lack of quantitative evaluation, which limits confidence in the generalization claim and in the reproducibility of the registration pipeline.
major comments (3)
- [§4 (Experiments)] The paper states that quantitative evaluation was omitted because 'the typical metrics for evaluating GAN models known in the literature did not correspond to the human-perceived quality of the generated texture.' This is not sufficient for a paper whose central claim is that a single photo yields a complete, reposable avatar. The authors have 10 held-out registered scans and a separate subset of 2056 scans with ground-truth segmentation and displacement maps; they should report per-pixel texture error (L1, SSIM), segmentation IoU, and vertex/displacement error on these held-out scans, even if these numbers are not the primary quality measure. A small user study would also substantiate the qualitative claims.
- [§3.1.1, Eq. (7)] The vertex weights w_i in the registration objective are introduced only qualitatively ('penalize deviations from the model more heavily for the vertices on the hands and the feet'), and the prior weights λθ, λβ in Eqs. (5)–(7) are not specified anywhere. Since the registrations produced by Eq. (7) define the ground-truth texture, segmentation, and displacement maps used for training, the choice of these weights directly influences every result in the paper. Please specify the values (or a table of them) and provide a sensitivity analysis, or at least state that results are robust to a reasonable range.
- [§2 and §3.2.1] The method's test-time behavior depends on DensePose and the garment segmenter producing partial UV maps that are well aligned with the scan-derived training distribution. The authors themselves note that DensePose 'was not designed to accommodate for hair and clothing deviating from the body' and that 'significant miss-alignments might occur.' Because the training partial maps come from rendered scans rather than unconstrained photographs, the magnitude of test-time DensePose errors is unknown. This is a load-bearing assumption for the claimed generalization to real images; the paper should quantify it, for example by perturbing DensePose correspondences at test time and measuring degradation, or by collecting a small annotated set of real images with manual UV correspondences.
minor comments (4)
- [Various] Minor typographical and grammatical errors: 'We train and our method on our newly created dataset' (§1) should read 'We train our method on our newly created dataset'; 'it’s garment segmentation' (§3.2) should be 'its'; 'miss-alignments' (§2) should be 'misalignments'; 'suplementary' (§4) should be 'supplementary'.
- [Eq. (7)] The Geman-McClure cost ρ is named but its explicit form is not given; please provide the formula or a precise reference so that the registration objective is fully reproducible.
- [§4.1] The statement that GAN metrics 'did not correspond to the human-perceived quality' would be more convincing if the paper listed which metrics were tried and in what way they failed, particularly because the lack of quantitative results is a salient departure from standard evaluation practice.
- [Figure 4] The qualitative results are the main evidence for the central claim, but the figure images are small; larger crops or side-by-side zooms would help the reader verify texture completion, segmentation boundaries, and garment geometry.
Circularity Check
No significant circularity: texture, segmentation, and displacement predictions are supervised by independent 3D-scan registrations, and the cited self-works are standard body-model/registration references rather than definitional inputs.
full rationale
The paper's derivation chain is: register SMPL to 4541 scans to obtain complete texture, segmentation, and displacement ground truth; render scan views and run DensePose to create partial maps; train image-to-image networks to map partial maps to complete maps; at test time apply the same DensePose extraction and run the trained networks. No output is used to define its own input. The targets are independent: 'we use DensePose solely to create the partial texture map from the input view, but train from high quality aligned and complete texture maps, which are obtained by registering SMPL to 3D scans.' The reconstruction and adversarial losses (Eqs. 8-14) compare network outputs to registered ground-truth maps, so the 'predictions' are not fitted parameters renamed as outputs. Self-citations to SMPL [41] and earlier registration work [16,53] are standard model usage that does not smuggle in the target result; SMPL is an external body model, and the registration equations (5)-(7) are supervised by scan geometry rather than by the final avatar. The statement that DensePose 'was not designed to accommodate for hair and clothing deviating from the body' is a limitation for test-time generalization, not a circular step. The paper's lack of quantitative evaluation is a rigor concern, but it does not make the derivation definitionally equivalent to its inputs.
Assumptions & free parameters
free parameters (3)
- Loss weights λ1-λ4 for texture, segmentation, and displacement objectives =
not specified in main text
- Pose and shape prior weights λθ, λβ =
not specified
- Registration vertex weights wi =
not specified
assumptions (5)
- standard math SMPL body model provides a common UV parametrization of the human body
- domain assumption Non-rigid registration of SMPL to 4541 scans produces consistent correspondences across clothing and hair
- domain assumption DensePose maps test images to SMPL surface with sufficient accuracy for partial map extraction
- domain assumption Clothing segmentation from [25] provides reliable ground-truth labels
- domain assumption Segmentation alone is sufficient to predict plausible clothing geometry
Cite this review
Pith. "Pith review of 360-Degree Textures of People in Clothing from a Single Image." pith.science (2026). https://pith.science/paper/53XHTTOQ
@misc{pith2026190807117,
author = {Pith},
title = {Pith review of: 360-Degree Textures of People in Clothing from a Single Image},
year = {2026},
howpublished = {\url{https://pith.science/paper/53XHTTOQ}},
note = {Machine review of arXiv:1908.07117}
}
read the original abstract
In this paper we predict a full 3D avatar of a person from a single image. We infer texture and geometry in the UV-space of the SMPL model using an image-to-image translation method. Given partial texture and segmentation layout maps derived from the input view, our model predicts the complete segmentation map, the complete texture map, and a displacement map. The predicted maps can be applied to the SMPL model in order to naturally generalize to novel poses, shapes, and even new clothing. In order to learn our model in a common UV-space, we non-rigidly register the SMPL model to thousands of 3D scans, effectively encoding textures and geometries as images in correspondence. This turns a difficult 3D inference task into a simpler image-to-image translation one. Results on rendered scans of people and images from the DeepFashion dataset demonstrate that our method can reconstruct plausible 3D avatars from a single image. We further use our model to digitally change pose, shape, swap garments between people and edit clothing. To encourage research in this direction we will make the source code available for research purpose.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
https://renderpeople.com/. 5
-
[2]
https://secure.axyz-design.com/. 5
-
[3]
https://web.twindom.com/. 5
-
[4]
https://www.treedys.com/. 5
-
[5]
http://virtualhumans.mpi-inf.mpg.de/360tex. 1
-
[6]
T. Alldieck, M. Magnor, B. L. Bhatnagar, C. Theobalt, and G. Pons-Moll. Learning to reconstruct people in clothing from a single RGB camera. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), jun 2019. 3, 5, 7
work page 2019
-
[7]
T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons- Moll. Detailed human avatars from monocular video. In International Conference on 3D Vision (3DV), sep 2018. 3
work page 2018
-
[8]
T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons- Moll. Video based reconstruction of 3d people models. In IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2018. 3, 5
work page 2018
Show all 78 references
-
[9]
Alldieck, G
T. Alldieck, G. Pons-Moll, C. Theobalt, and M. Magnor. Tex2shape: Detailed full human body geometry from a sin- gle image. arXiv preprint arXiv:1904.08645, 2019. 5
1904 arXiv
-
[10]
Arjovsky, S
M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein gan. arXiv preprint arXiv:1701.07875, 2017. 4
2017 arXiv
-
[11]
Balakrishnan, A
G. Balakrishnan, A. Zhao, A. V . Dalca, F. Durand, and J. Guttag. Synthesizing images of humans in unseen poses. arXiv preprint arXiv:1804.07739, 2018. 2
2018 arXiv
-
[12]
Baumberg
A. Baumberg. Blending images for texturing 3d models. In British Machine Vision Conference, volume 3, page 5. Cite- seer, 2002. 3
2002
-
[13]
Bernardini, I
F. Bernardini, I. M. Martin, and H. Rushmeier. High-quality texture reconstruction from multiple scans. IEEE Transac- tions on Visualization and Computer Graphics , 7(4):318– 332, 2001. 3
2001
-
[14]
S. Bi, N. K. Kalantari, and R. Ramamoorthi. Patch-based optimization for image-based texture mapping. ACM Trans- actions on Graphics, 36(4), 2017. 3
2017
-
[15]
F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Euro- pean Conference on Computer Vision. Springer International Publishing, Oct. 2016. 5
2016
-
[16]
F. Bogo, J. Romero, G. Pons-Moll, and M. J. Black. Dy- namic FAUST: Registering human bodies in motion. InIEEE Conf. on Computer Vision and Pattern Recognition, 2017. 3
2017
-
[17]
Z. Cao, T. Simon, S.-E. Wei, and Y . Sheikh. Realtime multi- person 2d pose estimation using part affinity fields. In IEEE Conf. on Computer Vision and Pattern Recognition, 2017. 3
2017
-
[18]
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros. Everybody dance now. arXiv preprint arXiv:1808.07371, 2018. 2, 8
2018 arXiv
-
[19]
Y . Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo. Stargan: Unified generative adversarial networks for multi- domain image-to-image translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 8789–8797, 2018. 4
2018
-
[20]
P. E. Debevec, C. J. Taylor, and J. Malik. Modeling and ren- dering architecture from photographs: A hybrid geometry- and image-based approach. In Annual Conf. on Computer Graphics and Interactive Techniques , pages 11–20. ACM,
-
[21]
Eisemann, B
M. Eisemann, B. De Decker, M. Magnor, P. Bekaert, E. De Aguiar, N. Ahmed, C. Theobalt, and A. Sellent. Float- ing textures. Computer Graphics Forum , 27(2):409–418,
-
[22]
Esser, E
P. Esser, E. Sutter, and B. Ommer. A variational u-net for conditional appearance and shape generation. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8857–8866, 2018. 2
2018
-
[23]
Y . Fu, Q. Yan, L. Yang, J. Liao, and C. Xiao. Texture map- ping for 3d reconstruction with rgb-d sensor. In IEEE Conf. on Computer Vision and Pattern Recognition, 2018. 3
2018
-
[24]
L. A. Gatys, A. S. Ecker, and M. Bethge. Image style trans- fer using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 2414–2423, 2016. 8
2016
-
[25]
K. Gong, X. Liang, Y . Li, Y . Chen, M. Yang, and L. Lin. Instance-level human parsing via part grouping network. In Proceedings of the European Conference on Computer Vi- sion (ECCV), pages 770–785, 2018. 5
2018
-
[26]
Grigorev, A
A. Grigorev, A. Sevastopolsky, A. Vakhitov, and V . Lem- pitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. 2, 3
2018 arXiv
-
[27]
R. A. G ¨uler, N. Neverova, and I. Kokkinos. Densepose: Dense human pose estimation in the wild. In IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[28]
Gulrajani, F
I. Gulrajani, F. Ahmed, M. Arjovsky, V . Dumoulin, and A. C. Courville. Improved training of wasserstein gans. In Advances in Neural Information Processing Systems , pages 5767–5777, 2017. 4
2017
-
[29]
Hamada, K
K. Hamada, K. Tachibana, T. Li, H. Honda, and Y . Uchida. Full-body high-resolution anime generation with progressive structure-conditional generative adversarial networks. arXiv preprint arXiv:1809.01890, 2018. 2
2018 arXiv
-
[30]
X. Han, Z. Wu, Z. Wu, R. Yu, and L. S. Davis. Viton: An image-based virtual try-on network. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2018. 2, 8
2018
-
[31]
Huang, T
Z. Huang, T. Li, W. Chen, Y . Zhao, J. Xing, C. LeGendre, L. Luo, C. Ma, and H. Li. Deep volumetric video from very sparse multi-view performance capture. In Proceedings of the European Conference on Computer Vision (ECCV) , pages 336–354, 2018. 8
2018
-
[32]
Huynh, W
L. Huynh, W. Chen, S. Saito, J. Xing, K. Nagano, A. Jones, P. Debevec, and H. Li. Mesoscopic facial geometry infer- ence using deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, pages 8407–8416, 2018. 3
2018
-
[33]
Isola, J.-Y
P. Isola, J.-Y . Zhu, T. Zhou, and A. A. Efros. Image-to- image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 2, 4
2017
-
[34]
Johnson, A
J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision , pages 694–711. Springer,
-
[35]
Kanazawa, M
A. Kanazawa, M. J. Black, D. W. Jacobs, and J. Malik. End- to-end recovery of human shape and pose. In Computer Vi- sion and Pattern Regognition (CVPR), 2018. 5
2018
-
[36]
Karras, T
T. Karras, T. Aila, S. Laine, and J. Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2017. 2
2017
-
[37]
Lassner, G
C. Lassner, G. Pons-Moll, and P. V . Gehler. A generative model of people in clothing. In Proceedings IEEE Interna- tional Conference on Computer Vision (ICCV) , Piscataway, NJ, USA, oct 2017. IEEE. 2
2017
-
[38]
Lempitsky and D
V . Lempitsky and D. Ivanov. Seamless mosaicing of image- based texture maps. In IEEE Conf. on Computer Vision and Pattern Recognition, pages 1–6. IEEE, 2007. 3
2007
-
[39]
H. P. Lensch, W. Heidrich, and H.-P. Seidel. A silhouette- based algorithm for texture registration and stitching.Graph- ical Models, 63(4):245–262, 2001. 3
2001
-
[40]
Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang. Deepfash- ion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1096–1104,
-
[41]
M. M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics, 34(6):248:1–248:16, Oct. 2015. 1, 2, 4
2015
-
[42]
Lorenz, L
D. Lorenz, L. Bereska, T. Milbich, and B. Ommer. Unsuper- vised part-based disentangling of object shape and appear- ance. arXiv preprint arXiv:1903.06946, 2019. 2
1903 arXiv
-
[43]
L. Ma, X. Jia, Q. Sun, B. Schiele, T. Tuytelaars, and L. Van Gool. Pose guided person image generation. In Advances in Neural Information Processing Systems , pages 406–416, 2017. 2, 8
2017
-
[44]
L. Ma, Q. Sun, S. Georgoulis, L. Van Gool, B. Schiele, and M. Fritz. Disentangled person image generation. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 99–108, 2018. 2, 8
2018
-
[45]
Martin-Brualla, R
R. Martin-Brualla, R. Pandey, S. Yang, P. Pidlypenskyi, J. Taylor, J. Valentin, S. Khamis, P. Davidson, A. Tkach, P. Lincoln, et al. Lookingood: enhancing performance capture with real-time neural re-rendering. arXiv preprint arXiv:1811.05029, 2018. 8
2018 arXiv
-
[46]
Mescheder, M
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger. Occupancy networks: Learning 3d reconstruction in function space. arXiv preprint arXiv:1812.03828, 2018. 8
2018 arXiv
-
[47]
Nagano, J
K. Nagano, J. Seo, J. Xing, L. Wei, Z. Li, S. Saito, A. Agar- wal, J. Fursund, H. Li, R. Roberts, et al. pagan: real-time avatars using dynamic textures. In SIGGRAPH Asia 2018 Technical Papers, page 258. ACM, 2018. 3
2018
-
[48]
Natsume, S
R. Natsume, S. Saito, Z. Huang, W. Chen, C. Ma, H. Li, and S. Morishima. Siclope: Silhouette-based clothed people. arXiv preprint arXiv:1901.00049, 2018. 3, 8
1901 arXiv
-
[49]
Neverova, R
N. Neverova, R. A. G¨uler, and I. Kokkinos. Dense pose trans- fer. arXiv preprint arXiv:1809.01995, 2018. 2, 3
2018 arXiv
-
[50]
Niem and J
W. Niem and J. Wingbermuhle. Automatic reconstruction of 3d objects using a mobile monoscopic camera. In Proc. Inter. Conf. on Recent Advances in 3-D Digital Imaging and Modeling, pages 173–180. IEEE, 1997. 3
1997
-
[51]
E. Ofek, E. Shilat, A. Rappoport, and M. Werman. Mul- tiresolution textures from image sequences. IEEE Computer Graphics and Applications, 17(2):18–29, Mar. 1997. 3
1997
-
[52]
Omran, C
M. Omran, C. Lassner, G. Pons-Moll, P. Gehler, and B. Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape estimation. In 2018 In- ternational Conference on 3D Vision (3DV), pages 484–494. IEEE, 2018. 5
2018
-
[53]
Pons-Moll, J
G. Pons-Moll, J. Romero, N. Mahmood, and M. J. Black. Dyna: a model of dynamic human shape in motion. ACM Transactions on Graphics, 34:120, 2015. 3
2015
-
[54]
A. Raj, P. Sangkloy, H. Chang, J. Hays, D. Ceylan, and J. Lu. Swapnet: Image based garment transfer. In European Conference on Computer Vision , pages 679–695. Springer, Cham, 2018. 3
2018
-
[55]
Rocchini, P
C. Rocchini, P. Cignoni, C. Montani, and R. Scopigno. Mul- tiple textures stitching and blending on 3d objects. In Ren- dering Techniques 99, pages 119–130. Springer, 1999. 3
1999
-
[56]
Saito, Z
S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. arXiv preprint arXiv:1905.05172, 2019. 3, 8
1905 arXiv
-
[57]
M. Sela, E. Richardson, and R. Kimmel. Unrestricted fa- cial geometry reconstruction using image-to-image transla- tion. In Proceedings of the IEEE International Conference on Computer Vision, pages 1576–1585, 2017. 3
2017
-
[58]
Shysheya, E
A. Shysheya, E. Zakharov, K.-A. Aliev, R. Bashirov, E. Burkov, K. Iskakov, A. Ivakhnenko, Y . Malkov, I. Pasech- nik, D. Ulyanov, et al. Textured neural avatars. arXiv preprint arXiv:1905.08776, 2019. 2
1905 arXiv
-
[59]
C. Si, W. Wang, L. Wang, and T. Tan. Multistage adversarial losses for pose-based human image synthesis. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 118–126, 2018. 2
2018
-
[60]
Siarohin, E
A. Siarohin, E. Sangineto, S. Lathuili `ere, and N. Sebe. Deformable gans for pose-based human image generation. In CVPR 2018-Computer Vision and Pattern Recognition ,
2018
-
[61]
Starck and A
J. Starck and A. Hilton. Model-based multiple view recon- struction of people. In IEEE international conference on computer vision, pages 915–922, 2003. 3
2003
-
[62]
Tewari, F
A. Tewari, F. Bernard, P. Garrido, G. Bharaj, M. Elgharib, H.-P. Seidel, P. P ´erez, M. Zollh ¨ofer, and C. Theobalt. Fml: Face model learning from videos. arXiv preprint arXiv:1812.07603, 2018. 3
2018 arXiv
-
[63]
Tewari, M
A. Tewari, M. Zollh ¨ofer, P. Garrido, F. Bernard, H. Kim, P. P´erez, and C. Theobalt. Self-supervised multi-level face model learning for monocular reconstruction at over 250 hz. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2549–2559...
2018
-
[64]
Varol, J
G. Varol, J. Romero, X. Martin, N. Mahmood, M. J. Black, I. Laptev, and C. Schmid. Learning from synthetic humans. In CVPR, 2017. 5
2017
-
[65]
Waechter, N
M. Waechter, N. Moehrle, and M. Goesele. Let there be color! large-scale texturing of 3d reconstructions. In Euro- pean Conf. on Computer Vision , pages 836–850. Springer,
-
[66]
B. Wang, H. Zheng, X. Liang, Y . Chen, L. Lin, and M. Yang. Toward characteristic-preserving image-based virtual try-on network. In ECCV, 2018. 2, 8
2018
-
[67]
Wang, M.-Y
T.-C. Wang, M.-Y . Liu, J.-Y . Zhu, A. Tao, J. Kautz, and B. Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 8798–8807, 2018. 4
2018
-
[68]
T. Y . Wang, D. Cylan, J. Popovic, and N. J. Mitra. Learning a shared shape space for multimodal garment design. arXiv preprint arXiv:1806.11335, 2018. 3
2018 arXiv
-
[69]
Z. Wang, A. C. Bovik, H. R. Sheikh, E. P. Simoncelli, et al. Image quality assessment: from error visibility to struc- tural similarity. IEEE transactions on image processing , 13(4):600–612, 2004. 4
2004
-
[70]
Z. Wang, E. P. Simoncelli, and A. C. Bovik. Multiscale struc- tural similarity for image quality assessment. In The Thrity- Seventh Asilomar Conference on Signals, Systems & Com- puters, 2003, volume 2, pages 1398–1402. Ieee, 2003. 4
2003
-
[71]
Q. Xu, W. Wang, D. Ceylan, R. Mech, and U. Neumann. Disn: Deep implicit surface network for high-quality single- view 3d reconstruction. arXiv preprint arXiv:1905.10711 ,
1905 arXiv
-
[72]
Yamaguchi, S
S. Yamaguchi, S. Saito, K. Nagano, Y . Zhao, W. Chen, K. Olszewski, S. Morishima, and H. Li. High-fidelity facial reflectance and geometry inference from an unconstrained image. ACM Transactions on Graphics (TOG) , 37(4):162,
-
[73]
C. Yang, Z. Wang, X. Zhu, C. Huang, J. Shi, and D. Lin. Pose guided human video generation. In ECCV, 2018. 2
2018
-
[74]
Yang, J.-S
J. Yang, J.-S. Franco, F. H´etroy-Wheeler, and S. Wuhrer. An- alyzing clothing layer deformation statistics of 3d human motions. In Proceedings of the European Conference on Computer Vision (ECCV), pages 237–253, 2018. 3
2018
-
[75]
Zanfir, A.-I
M. Zanfir, A.-I. Popa, A. Zanfir, and C. Sminchisescu. Hu- man appearance transfer. In The IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), June 2018. 3
2018
-
[76]
Zhou and V
Q.-Y . Zhou and V . Koltun. Color map optimization for 3d reconstruction with consumer depth cameras. ACM Trans- actions on Graphics, 33(4):155, 2014. 3
2014
-
[77]
J.-Y . Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image- to-image translation using cycle-consistent adversarial net- works. In Proceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017. 4
2017
-
[78]
S. Zhu, S. Fidler, R. Urtasun, D. Lin, and C. Change Loy. Be your own prada: Fashion synthesis with structural co- herence. In International Conference on Computer Vision (ICCV), 2017. 2
2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.