REVIEW 4 major objections 5 minor 93 references
Disentangled Clothed Avatar Generation with Layered Representation
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper introduces LayerAvatar, a feed-forward diffusion method that generates clothed 3D avatars whose body, hair, shoes, and top/bottom clothing are stored in separate layers of a Gaussian-based UV feature plane, so components can be…
desk verdict Layered UV feature plane is a real step forward for feed-forward avatar generation, but the headline FID is unproven until the authors report a held-out training/evaluation split. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The layered UV feature plane carries the argument. Instead of scattering Gaussians in an unstructured 3D field, the method initializes Gaussian primitives on per-component SMPL-X templates and writes their attributes (position offset, opacity, rotation delta, scale delta, color) as local features in a shared 2D UV plane divided into three layers with semantic labels. Two lightweight shared MLP decoders, one for geometry and one for texture, convert the plane into attribute maps, and each Gaussian samples its own attributes by bilinear interpolation. Because overlapping components occupy different layers, they get independent feature spaces, which is what lets the diffusion model separate body from clothing; because the templates carry SMPL-X skinning weights, the same representation deforms into novel poses and body shapes by linear blend skinning.
What would settle it
Take a group of subjects in tight, skin-colored, or patterned clothing, render their multi-view images, run the same segmentation and SMPL-X fitting, and compare the isolated body layer against the true body surface obtained from an under-clothing scan or body-fitting method. If the body layer systematically adopts clothing color or geometry wherever segmentation confuses fabric with skin, the claimed disentanglement fails. A cheaper quantitative version is to measure L-PSNR between body and clothing layers on such subjects; if it falls toward the single-layer baseline value, the layered representation only separates components when the supervision is clean.
Extended reading notes
Core claim
The central claim is that component disentanglement can be baked into the generative representation rather than imposed afterward. LayerAvatar represents a clothed avatar as a set of Gaussian primitives $G_{\text{avatar}} = \{G_{\text{body}}, G_{\text{top}}, G_{\text{bottom}}, G_{\text{hair}}, G_{\text{shoes}}\}$, each component anchored to an SMPL-X-based template and mapped into a three-layer UV feature plane with semantic labels: the innermost layer holds the body, the second holds hair and shoes, and the third holds top and bottom clothing. A single-stage diffusion model is trained on these layered planes, simultaneously fitting the planes from multi-view images and learning the denoising prior, with per-component rendering losses, a semantic segmentation loss, and two occlusion-specific constraints (an inner-body mask term and a skin-color prior from the hands) to keep the heavily occluded body layer plausible. The paper reports state-of-the-art FID on THuman2.0, layer-wise L-PSNR above 40 on Tightcap, decomposition without identity shifting, and successful component transfer across subjects.
Load-bearing premise
The load-bearing premise, which the paper itself flags in Section 8.1(1), is that the semantic segmentation maps used as ground truth correctly separate body, hair, top, bottom, and shoes, and that the fitted SMPL-X body template lies inside the clothing; if either is wrong, the layers receive incorrect supervision and the disentanglement collapses.
Editorial extensions
If this is right
- A single avatar is generated in about two seconds, because the diffusion model produces the whole layered representation in one pass and no per-subject optimization is needed.
- Components transfer directly: the paper demonstrates upper clothes, pants, hair, and shoes being moved between avatars of different body shapes while keeping detail.
- Generated avatars support animation: because each layer is attached to SMPL-X templates, novel gestures, facial expressions, and body shapes are obtained by warping and linear blend skinning.
- Decomposition is stable: rendering any single layer does not cause identity shifting, which the paper uses as evidence that the layers truly separate the avatar rather than sharing entangled features.
- The representation also generalizes to multi-layer outfits and dress/skirt types, and to single-image reconstruction of unseen subjects, suggesting the same pipeline will scale to larger, more diverse training sets.
Reading between the lines
- A testable extension is to add more layers for accessories such as glasses or bags; the architecture treats layers as independent UV planes, so capacity and disentanglement should scale with layer count rather than requiring a redesign.
- Because generation happens in a latent UV space, a conditioning mechanism (text prompt, pose, or identity image) could be attached to the denoiser to make the same representation support controlled generation; the paper does not explore this.
- The fidelity of the body layer will probably track the quality of the segmentation supervision; comparing results on datasets with true component meshes, such as the Tightcap split used here, against the Sapiens-mask supervision on THuman2.0 would isolate this dependence.
- The paper leaves implicit that a collision-avoidance post-process on the body and clothing Gaussians could turn the generated layers into physically plausible assets, which matters for animation and simulation downstream.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LayerAvatar, a feed-forward diffusion-based method for generating clothed avatars with disentangled body, hair, shoes, top, and bottom components. The key representation is a layered UV feature plane storing Gaussian attributes for three component groups (body; hair+shoes; top+bottom), with semantic labels used to separate components at render time. The method trains a single-stage diffusion model on multi-view images with reconstruction, segmentation, and occlusion constraints, and claims high-quality generation in seconds, support for animation and novel-view synthesis, and component transfer. Experiments compare holistic generation quality on THuman2.0, layer-wise quality on Tightcap, and component quality against optimization-based methods, with ablations studying the layered representation and single-stage training.
Significance. If the quantitative claims are established, this is a meaningful step: a feed-forward generator that outputs disentangled, animatable avatars with component transfer is practically useful, and the layered UV feature plane is a well-motivated representation that combines the editability of layers with the quality of Gaussian splatting. The paper also provides ablations that support the layered representation over single-layer alternatives, and it ships a reproducible-looking pipeline with explicit loss terms and hyperparameters. However, the current evaluation protocol leaves the headline claims of 'high-quality' and 'superior performance' insufficiently supported, so the significance depends on closing the measurement gaps below.
major comments (4)
- [Sec. 4 (Dataset) and Sec. 4.1 (Evaluation of Generation Quality)] The FID evaluation on THuman2.0 does not specify any held-out split. Section 4 describes sampling 500 THuman2.0 scans, rendering 54 views each, and using them for training, and then describes a composite training set of 1954 selected scans from THuman2.0, THuman2.1, and CustomHuman. Section 4.1 then evaluates 'on THuman2.0 dataset' and reports FID 12.50, but it is never stated that the scans used to compute the FID reference set are disjoint from the scans used to fit the layered UV feature planes and train the diffusion model. If the reference images are drawn from the training scans, the FID can be artificially low due to memorization. This directly affects the abstract's claim of 'high-quality' generation. Please state the exact split and, if necessary, recompute the FID on a held-out subset of THuman2.0 subjects.
- [Table 1a and footnotes] The baseline FID values for EVA3D, StructLDM, and E3Gen are adopted from their respective papers (footnotes *, †, ⋆), which may use different reference image sets, rendering protocols, and evaluation subsets. A lower FID under a different protocol is not evidence of superiority. The claim in Sec. 4.1 that 'Our method outperforms all baselines on THuman2.0' is therefore not supported by a controlled comparison. Please re-run the baselines under the same evaluation protocol used for LayerAvatar, or at least clearly restrict the claim to the adopted numbers and discuss protocol differences.
- [Table 1b and Sec. 4.1] The L-PSNR value for the proposed method is reported as '>40' rather than as an exact number, with the explanation that the two masked layers are nearly identical. This is not a precise metric value and it makes the disentanglement comparison in Table 1b difficult to interpret; the text states the method 'surpasses other methods' in L-PSNR, but a threshold without a distribution cannot support a quantitative superiority claim. Please report the actual mean L-PSNR (and ideally its standard deviation or per-sample distribution) for the Tightcap layer-wise evaluation.
- [Sec. 3.4 and Sec. 7.3] The component supervision for THuman2.0, THuman2.1, and CustomHuman derives entirely from Sapiens-predicted semantic segmentation, as stated in Sec. 3.4 ('The ground truth of silhouette masks is estimated based on the semantic segmentation results predicted by Sapiens'). The authors acknowledge this in Sec. 8.1(1). Since the same kind of predicted masks are also used to verify disentanglement in the THuman2.0-based results, the paper would be strengthened by a quantitative sensitivity analysis of the segmentation noise, beyond the qualitative Figure J. For example, report FID or component-mask IoU when training with clean masks (e.g., on Tightcap) versus predicted masks, or evaluate the final generation quality on a subset with manually verified masks.
minor comments (5)
- [Sec. 4.1] The sentence 'The elimination of identity shifting demonstrates that our method achieves full disentanglement' overstates the implication: absence of identity shifting is one proxy for disentanglement, not a proof of full disentanglement. Consider softening the wording.
- [Table 1] No error bars, confidence intervals, or number of random seeds are reported for any FID, KID, L-PSNR, or user-study result. Given the variability of generative model metrics, at least a note on single-seed reporting or a variance estimate would improve reproducibility.
- [Sec. 7.2 and Figure 2] The label 'Clothed AvatarHuman BodyExteriorComponentsSegmentation maskConstraint Loss' in Figure 2 appears to be a formatting artifact; please fix the figure layout so that each label is legible and attached to the correct part.
- [Sec. 8.1(2)] There is a typo in '3G Gaussians' which should read '3D Gaussians'.
- [Sec. 3.2 and Sec. 7.1] The main text says the UV feature plane is split 'channel-wise', but the supplementary says the three layers are concatenated width-wise into a tensor of size 12 × 128 × 384. Please clarify the exact layout (channel dimension and width dimension) in the main text for consistency.
Circularity Check
No circular derivation: generation and disentanglement claims rest on supervised diffusion training and independent benchmarks, not on self-referential inputs.
full rationale
The paper's central claim is an empirical generative architecture, not a derivation that takes its conclusion as an input. The pipeline fits layered UV feature planes to multi-view images under reconstruction, mask, perceptual, and segmentation losses (Eqs. 7-10), trains a denoising UNet on the planes (Eqs. 11-12), and samples noise to produce new avatars; no fitted parameter is renamed as a prediction and no evaluation quantity is used as a training target. The Sapiens-derived component masks in Sec. 3.4 are training supervision, and Sec. 8.1(1) explicitly lists segmentation/SMPL-X accuracy as an external premise rather than as part of the claim. The layer-wise disentanglement evaluation in Tab. 1b is run on Tightcap, whose per-component masks come from separate 3D meshes (Sec. 4), so the metric does not simply re-measure alignment with the Sapiens oracle. The same-first-author citation [87] is used for the UV-space Gaussian attribute encoding and is a design predecessor/baseline, not a load-bearing self-citation or an imported uniqueness theorem. The THuman2.0 FID protocol does not state a held-out split, which is a legitimate evaluation-transparency caveat, but on the quoted evidence it is not a reduction of the derivation to its own inputs and therefore is not scored as circularity.
Assumptions & free parameters
free parameters (3)
- Loss weights in Eq. 7-10 =
lambda_color=18, lambda_mask=9, lambda_per=0.05, lambda_seg=9, lambda_maskin=5, lambda_skin=0.5, lambda_offset=5…
- UV feature plane size =
12 channels x 128 x 384
- Training hyperparameters =
lr UNet/decoder 1e-4, lr UV plane 0.04, batch 4 scenes per GPU, 2 views per scene
assumptions (5)
- domain assumption SMPL-X body model with shape, pose, and expression blend shapes is an adequate prior for human body geometry and skinning.
- domain assumption Sapiens semantic segmentation provides accurate per-component masks for body, hair, top, bottom, and shoes.
- ad hoc to paper The fixed three-layer grouping (body; hair+shoes; top+bottom) is sufficient to represent relevant avatar components and their occlusion order.
- ad hoc to paper Occluded body skin color can be estimated from the average hand color.
- standard math 3D Gaussian splatting with the described rendering equation is differentiable and adequate for high-quality avatar rendering.
Cite this review
Pith. "Pith review of Disentangled Clothed Avatar Generation with Layered Representation." pith.science (2026). https://pith.science/paper/JEG2M7IB
@misc{pith2026250104631,
author = {Pith},
title = {Pith review of: Disentangled Clothed Avatar Generation with Layered Representation},
year = {2026},
howpublished = {\url{https://pith.science/paper/JEG2M7IB}},
note = {Machine review of arXiv:2501.04631}
}
read the original abstract
Clothed avatar generation has wide applications in virtual and augmented reality, filmmaking, and more. Previous methods have achieved success in generating diverse digital avatars, however, generating avatars with disentangled components (\eg, body, hair, and clothes) has long been a challenge. In this paper, we propose LayerAvatar, the first feed-forward diffusion-based method for generating component-disentangled clothed avatars. To achieve this, we first propose a layered UV feature plane representation, where components are distributed in different layers of the Gaussian-based UV feature plane with corresponding semantic labels. This representation supports high-resolution and real-time rendering, as well as expressive animation including controllable gestures and facial expressions. Based on the well-designed representation, we train a single-stage diffusion model and introduce constrain terms to address the severe occlusion problem of the innermost human body layer. Extensive experiments demonstrate the impressive performances of our method in generating disentangled clothed avatars, and we further explore its applications in component transfer. The project page is available at: https://olivia23333.github.io/LayerAvatar/
Figures
Reference graph
Works this paper leans on
-
[1]
Gaussian shell maps for efficient 3d human generation, 2023
Rameen Abdal, Wang Yifan, Zifan Shi, Yinghao Xu, Ryan Po, Zhengfei Kuang, Qifeng Chen, Dit-Yan Yeung, and Gor- don Wetzstein. Gaussian shell maps for efficient 3d human generation, 2023. 3
2023
-
[2]
Driving-signal aware full-body avatars
Timur Bagautdinov, Chenglei Wu, Tomas Simon, Fabi ´an Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, and Jason Saragih. Driving-signal aware full-body avatars. ACM Trans. Graph., 40(4), 2021. 1
2021
-
[3]
Bergman, Petr Kellnhofer, Wang Yifan, Eric R
Alexander W. Bergman, Petr Kellnhofer, Wang Yifan, Eric R. Chan, David B. Lindell, and Gordon Wetzstein. Gen- erative neural articulated radiance fields. In NeurIPS, 2022. 2
2022
-
[4]
pi-gan: Periodic implicit generative ad- versarial networks for 3d-aware image synthesis
Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative ad- versarial networks for 3d-aware image synthesis. In Proc. CVPR, 2021. 2
2021
-
[5]
Chan, Connor Z
Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In CVPR, 2022. 2, 6
2022
-
[6]
Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction
Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction. In ICCV, 2023. 2, 5, 6
2023
-
[7]
Neural-abc: Neural parametric models for articulated body with clothes
Honghu Chen, Yuxin Yao, and Juyong Zhang. Neural-abc: Neural parametric models for articulated body with clothes. IEEE Transactions on Visualization and Computer Graphics,
-
[8]
Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation
Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22246–22256, 2023. 2
2023
Show all 93 references
-
[9]
Tightcap: 3d human shape capture with clothing tightness field
Xin Chen, Anqi Pang, Yang Wei, Wang Peihao, Lan Xu, and Jingyi Yu. Tightcap: 3d human shape capture with clothing tightness field. ACM Transactions on Graphics (Presented at ACM SIGGRAPH), 2021. 2, 6, 7
2021
-
[10]
Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes
Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes. In Interna- tional Conference on Computer Vision (ICCV), 2021. 3
2021
-
[11]
Black, and Otmar Hilliges
Xu Chen, Tianjian Jiang, Jie Song, Max Rietmann, Andreas Geiger, Michael J. Black, and Otmar Hilliges. Fast-snarf: A fast deformer for articulated neural fields. Pattern Analysis and Machine Intelligence (PAMI), 2023. 3
2023
-
[12]
Primdiffusion: V olumet- ric primitives diffusion for 3d human generation
Zhaoxi Chen, Fangzhou Hong, Haiyi Mei, Guangcong Wang, Lei Yang, and Ziwei Liu. Primdiffusion: V olumet- ric primitives diffusion for 3d human generation. In Thirty- seventh Conference on Neural Information Processing Sys- tems, 2023. 2
2023
-
[13]
Expressive telepresence via mod- ular codec avatars
Hang Chu, Shugao Ma, Fernando De la Torre, Sanja Fi- dler, and Yaser Sheikh. Expressive telepresence via mod- ular codec avatars. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XII 16, pages 330–345. Springer, 2020. 1
2020
-
[14]
Smplicit: Topology-aware generative model for clothed people
Enric Corona, Albert Pumarola, Guillem Aleny `a, Ger- ard Pons-Moll, and Francesc Moreno-Noguer. Smplicit: Topology-aware generative model for clothed people. In CVPR, 2021. 2
2021
-
[15]
Tela: Text to layer-wise 3d clothed human generation
Junting Dong, Qi Fang, Zehuan Huang, Xudong Xu, Jingbo Wang, Sida Peng, and Bo Dai. Tela: Text to layer-wise 3d clothed human generation. arXiv preprint arXiv:2404.16748, 2024. 2, 3, 6, 7
2024 arXiv
-
[16]
AG3D: Learning to Gen- erate 3D Avatars from 2D Image Collections
Zijian Dong, Xu Chen, Jinlong Yang, Michael J Black, Ot- mar Hilliges, and Andreas Geiger. AG3D: Learning to Gen- erate 3D Avatars from 2D Image Collections. In Interna- tional Conference on Computer Vision (ICCV) , 2023. 2, 3, 5
2023
-
[17]
Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning. Neural networks, 107:3–11,
-
[18]
Black, and Timo Bolkart
Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J. Black, and Timo Bolkart. Capturing and animation of body and clothing from monocular video. In SIGGRAPH Asia 2022 Conference Papers, 2022. 3, 5
2022
-
[19]
Laga: Layered 3d avatar genera- tion and customization via gaussian splatting
Jia Gong, Shenyu Ji, Lin Geng Foo, Kang Chen, Hossein Rahmani, and Jun Liu. Laga: Layered 3d avatar genera- tion and customization via gaussian splatting. arXiv preprint arXiv:2405.12663, 2024. 2, 3, 6, 7
2024 arXiv
-
[20]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 6
2017
-
[21]
Learn- ing locally editable virtual humans
Hsuan-I Ho, Lixin Xue, Jie Song, and Otmar Hilliges. Learn- ing locally editable virtual humans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21024–21035, 2023. 2, 6
2023
-
[22]
EV A3d: Compositional 3d human generation from 2d image collections
Fangzhou Hong, Zhaoxi Chen, Yushi LAN, Liang Pan, and Ziwei Liu. EV A3d: Compositional 3d human generation from 2d image collections. In International Conference on Learning Representations, 2023. 2, 3, 6, 7
2023
-
[23]
3dtopia: Large text-to-3d genera- tion model with hybrid diffusion priors
Fangzhou Hong, Jiaxiang Tang, Ziang Cao, Min Shi, Tong Wu, Zhaoxi Chen, Shuai Yang, Tengfei Wang, Liang Pan, Dahua Lin, et al. 3dtopia: Large text-to-3d genera- tion model with hybrid diffusion priors. arXiv preprint arXiv:2403.02234, 2024. 2
2024 arXiv
-
[24]
Humanliff: Layer-wise 3d human generation with diffusion model.arXiv preprint, 2023
Shoukang Hu, Fangzhou Hong, Tao Hu, Liang Pan, Haiyi Mei, Weiye Xiao, Lei Yang, and Ziwei Liu. Humanliff: Layer-wise 3d human generation with diffusion model.arXiv preprint, 2023. 2, 3, 6, 7
2023
-
[25]
Gauhuman: Articu- lated gaussian splatting from monocular human videos
Shoukang Hu, Tao Hu, and Ziwei Liu. Gauhuman: Articu- lated gaussian splatting from monocular human videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 20418–20431, 2024. 4
2024
-
[26]
Structldm: Struc- tured latent diffusion for 3d human generation, 2024
Tao Hu, Fangzhou Hong, and Ziwei Liu. Structldm: Struc- tured latent diffusion for 3d human generation, 2024. 2, 6, 7
2024
-
[27]
Robust estimation of a location parameter
Peter J Huber. Robust estimation of a location parameter. In Breakthroughs in statistics: Methodology and distribution , pages 492–518. Springer, 1992. 5
1992
-
[28]
Multiply: Re- construction of multiple people from monocular video in the wild
Zeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang, Julien Valentin, Otmar Hilliges, and Jie Song. Multiply: Re- construction of multiple people from monocular video in the wild. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[29]
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision , pages 694–711. Springer, 2016. 5
2016
-
[30]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 2, 4, 1
2023
-
[31]
Sapiens: Foundation for human vision mod- els
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vision mod- els. arXiv preprint arXiv:2408.12569, 2024. 5, 6, 2
2024 arXiv
-
[32]
Chupa: Carving 3d clothed humans from skinned shape priors us- ing 2d diffusion probabilistic models
Byungjun Kim, Patrick Kwon, Kwangho Lee, Myunggi Lee, Sookwan Han, Daesik Kim, and Hanbyul Joo. Chupa: Carving 3d clothed humans from skinned shape priors us- ing 2d diffusion probabilistic models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICC...
2023
-
[33]
Adam: A method for stochastic opti- mization
Diederik P Kingma. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980, 2014. 1
2014 arXiv
-
[34]
Instant 3d human avatar generation using image diffusion models
Nikos Kolotouros, Thiemo Alldieck, Enric Corona, Ed- uard Gabriel Bazavan, and Cristian Sminchisescu. Instant 3d human avatar generation using image diffusion models. arXiv preprint arXiv:2406.07516, 2024. 3
2024 arXiv
-
[35]
Neuraltailor: Recon- structing sewing pattern structures from 3d point clouds of garments
Maria Korosteleva and Sung-Hee Lee. Neuraltailor: Recon- structing sewing pattern structures from 3d point clouds of garments. ACM Transactions on Graphics (TOG), 41(4):1– 16, 2022. 2
2022
-
[36]
Garma- genet: A multimodal generative framework for sewing pat- tern design and generic garment modeling
Siran Li, Chen Liu, Ruiyang Liu, Zhendong Wang, Gaofeng He, Yong-Lu Li, Xiaogang Jin, and Huamin Wang. Garma- genet: A multimodal generative framework for sewing pat- tern design and generic garment modeling. arXiv preprint arXiv:2504.01483, 2025. 2
2025
-
[37]
Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching
Yixun Liang, Xin Yang, Jiantao Lin, Haodong Li, Xiao- gang Xu, and Yingcong Chen. Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 6517–6526, 2024. 2
2024
-
[38]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[39]
Layga: Layered gaussian avatars for animatable clothing transfer
Siyou Lin, Zhe Li, Zhaoqi Su, Zerong Zheng, Hongwen Zhang, and Yebin Liu. Layga: Layered gaussian avatars for animatable clothing transfer. InACM SIGGRAPH 2024 Con- ference Papers, pages 1–11, 2024. 3
2024
-
[40]
Zero-1-to- 3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to- 3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9298–9309, 2023. 2
2023
-
[41]
Clothedreamer: Text- guided garment generation with 3d gaussians.arXiv preprint arXiv:2406.16815, 2024
Yufei Liu, Junshu Tang, Chu Zheng, Shijie Zhang, Jinkun Hao, Junwei Zhu, and Dongjin Huang. Clothedreamer: Text- guided garment generation with 3d gaussians.arXiv preprint arXiv:2406.16815, 2024. 5
2024 arXiv
-
[42]
Black, Derek Nowrouzezahrai, Liam Paull, and Weiyang Liu
Zhen Liu, Yao Feng, Michael J. Black, Derek Nowrouzezahrai, Liam Paull, and Weiyang Liu. Meshd- iffusion: Score-based generative 3d mesh modeling. In International Conference on Learning Representations ,
-
[43]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi- person linear model. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 2
2015
-
[44]
Coap: Compositional articulated occupancy of people
Marko Mihajlovic, Shunsuke Saito, Aayush Bansal, Michael Zollhoefer, and Siyu Tang. Coap: Compositional articulated occupancy of people. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13201–13210, 2022. 3
2022
-
[45]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 2
2021
-
[46]
Expressive whole-body 3d gaussian avatar
Gyeongsik Moon, Takaaki Shiratori, and Shunsuke Saito. Expressive whole-body 3d gaussian avatar. arXiv preprint arXiv:2407.21686, 2024. 3
2024 arXiv
-
[47]
Diffrf: Rendering-guided 3d radiance field diffusion
Norman M ¨uller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. Diffrf: Rendering-guided 3d radiance field diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4328–4338, 2023. 2
2023
-
[48]
Point-e: A system for generat- ing 3d point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generat- ing 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2
2022 arXiv
-
[49]
Giraffe: Represent- ing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger. Giraffe: Represent- ing scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recog- nition (CVPR), 2021. 2
2021
-
[50]
Unsupervised learning of efficient geometry-aware neural articulated representations
Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. Unsupervised learning of efficient geometry-aware neural articulated representations. In European Conference on Computer Vision, 2022. 2, 3
2022
-
[51]
Spams: Structured implicit parametric models
Pablo Palafox, Nikolaos Sarafianos, Tony Tung, and Angela Dai. Spams: Structured implicit parametric models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12851–12860, 2022. 3
2022
-
[52]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 3, 2
2019
-
[53]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,
-
[54]
Pica: Physics-integrated clothed avatar
Bo Peng, Yunfan Tao, Haoyu Zhan, Yudong Guo, and Juy- ong Zhang. Pica: Physics-integrated clothed avatar. arXiv preprint arXiv:2407.05324, 2024. 3
2024 arXiv
-
[55]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 2, 5
2022 arXiv
-
[56]
Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to- 3d
Lingteng Qiu, Guanying Chen, Xiaodong Gu, Qi Zuo, Mu- tian Xu, Yushuang Wu, Weihao Yuan, Zilong Dong, Liefeng Bo, and Xiaoguang Han. Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to- 3d. In Proceedings of the IEEE/CVF Conference on Com- ...
-
[57]
Xcube: Large-scale 3d generative modeling using sparse voxel hierarchies
Xuanchi Ren, Jiahui Huang, Xiaohui Zeng, Ken Museth, Sanja Fidler, and Francis Williams. Xcube: Large-scale 3d generative modeling using sparse voxel hierarchies. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2
2024
-
[58]
High-resolution image syn- thesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models, 2021. 2, 1
2021
-
[59]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 6
2022
-
[60]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...
2015
-
[61]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Confer- ence on Learning Representations, 2022. 6
2022
-
[62]
Deep marching tetrahedra: a hybrid represen- tation for high-resolution 3d shape synthesis
Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep marching tetrahedra: a hybrid represen- tation for high-resolution 3d shape synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021. 2
2021
-
[63]
Mvdream: Multi-view diffusion for 3d gen- eration
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d gen- eration. arXiv preprint arXiv:2308.16512, 2023. 2
2023 arXiv
-
[64]
3d neural field generation using triplane diffusion
J Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein. 3d neural field generation using triplane diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 20875–20886, 2023. 2
2023
-
[65]
Danbo: Disentangled articulated neural body representations via graph neural networks
Shih-Yang Su, Timur Bagautdinov, and Helge Rhodin. Danbo: Disentangled articulated neural body representations via graph neural networks. InEuropean Conference on Com- puter Vision, 2022. 3
2022
-
[66]
Barbie: Text to barbie-style 3d avatars
Xiaokun Sun, Zhenyu Zhang, Ying Tai, Qian Wang, Hao Tang, Zili Yi, and Jian Yang. Barbie: Text to barbie-style 3d avatars. arXiv preprint arXiv:2408.09126, 2024. 2
2024
-
[67]
Splatter image: Ultra-fast single-view 3d recon- struction
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 4
2024
-
[68]
Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation. arXiv preprint arXiv:2309.16653,
-
[69]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation.arXiv preprint arXiv:2402.05054, 2024
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation.arXiv preprint arXiv:2402.05054, 2024. 4
2024 arXiv
-
[70]
Disentangled clothed avatar generation from text de- scriptions, 2023
Jionghao Wang, Yuan Liu, Zhiyang Dou, Zhengming Yu, Yongqing Liang, Xin Li, Wenping Wang, Rong Xie, and Li Song. Disentangled clothed avatar generation from text de- scriptions, 2023. 2, 6, 7
2023
-
[71]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS, 2021. 2
2021
-
[72]
Rodin: A generative model for sculpting 3d digital avatars using diffusion
Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In Proceedings of the IEEE/CVF conference on computer vision and...
2023
-
[73]
Humancoser: Layered 3d human genera- tion via semantic-aware diffusion model
Yi Wang, Jian Ma, Ruizhi Shao, Qiao Feng, Yu-Kun Lai, and Kun Li. Humancoser: Layered 3d human genera- tion via semantic-aware diffusion model. arXiv preprint arXiv:2408.11357, 2024. 2
2024 arXiv
-
[74]
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion. Advances in Neural Information Processing Systems , 36, 2024. 2
2024
-
[75]
Hu- mannerf: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern Recognition , pages 162...
2022
-
[76]
Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer
Shuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng, Jingxi Xu, Philip Torr, Xun Cao, and Yao Yao. Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer. arXiv preprint arXiv:2405.14832, 2024. 2
2024 arXiv
-
[77]
Dressing avatars: Deep photorealistic appearance for physically sim- ulated clothing
Donglai Xiang, Timur Bagautdinov, Tuur Stuyck, Fabian Prada, Javier Romero, Weipeng Xu, Shunsuke Saito, Jing- fan Guo, Breannan Smith, Takaaki Shiratori, et al. Dressing avatars: Deep photorealistic appearance for physically sim- ulated clothing. ACM Transactions on Graphics (...
2022
-
[78]
Efficient 3d articu- lated human generation with layered surface volumes
Yinghao Xu, Wang Yifan, Alexander W Bergman, Menglei Chai, Bolei Zhou, and Gordon Wetzstein. Efficient 3d articu- lated human generation with layered surface volumes. arXiv preprint arXiv:2307.05462, 2023. 3
2023 arXiv
-
[79]
Xingguang Yan, Han-Hung Lee, Ziyu Wan, and Angel X. Chang. An object is worth 64x64 pixels: Generating 3d ob- ject via image diffusion, 2024. 2
2024
-
[80]
Dialoguenerf: Towards realistic avatar face- to-face conversation video generation
Yichao Yan, Zanwei Zhou, Zi Wang, Jingnan Gao, and Xi- aokang Yang. Dialoguenerf: Towards realistic avatar face- to-face conversation video generation. Visual Intelligence, 2 (1):24, 2024. 3
2024
-
[81]
Function4d: Real-time human vol- umetric capture from very sparse consumer rgbd sensors
Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qiong- hai Dai, and Yebin Liu. Function4d: Real-time human vol- umetric capture from very sparse consumer rgbd sensors. In IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR2021), 2021. 2, 6, 7
2021
-
[82]
Surf-d: High-quality surface generation for arbitrary topologies using diffusion models
Zhengming Yu, Zhiyang Dou, Xiaoxiao Long, Cheng Lin, Zekun Li, Yuan Liu, Norman M ¨uller, Taku Komura, Marc Habermann, Christian Theobalt, et al. Surf-d: High-quality surface generation for arbitrary topologies using diffusion models. arXiv preprint arXiv:2311.17050, 2023. 2
2023 arXiv
-
[83]
Lion: Latent point diffusion models for 3d shape generation
Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 2
2022
-
[84]
Avatargen: A 3d generative model for ani- matable human avatars
Jianfeng Zhang, Zihang Jiang, Dingdong Yang, Hongyi Xu, Yichun Shi, Guoxian Song, Zhongcong Xu, Xinchao Wang, and Jiashi Feng. Avatargen: A 3d generative model for ani- matable human avatars. In Arxiv, 2022. 2
2022
-
[85]
Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets
Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets. ACM Transactions on Graphics (TOG), 43(4):1–20, 2024. 2
2024
-
[86]
Joint2human: High-quality 3d human genera- tion via compact spherical embedding of 3d joints
Muxin Zhang, Qiao Feng, Zhuo Su, Chao Wen, Zhou Xue, and Kun Li. Joint2human: High-quality 3d human genera- tion via compact spherical embedding of 3d joints. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[87]
e3gen: Efficient, expressive and ed- itable avatars generation
Weitian Zhang, Yichao Yan, Yunhui Liu, Xingdong Sheng, and Xiaokang Yang. e3gen: Efficient, expressive and ed- itable avatars generation. arXiv preprint arXiv:2405.19203,
-
[88]
Getavatar: Generative textured meshes for animatable human avatars
Xuanmeng Zhang, Jianfeng Zhang, Chacko Rohan, Hongyi Xu, Guoxian Song, Yi Yang, and Jiashi Feng. Getavatar: Generative textured meshes for animatable human avatars. In ICCV, 2023. 2
2023
-
[89]
Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation
Zibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng, Rui Wang, Pei Cheng, BIN FU, Tao Chen, Gang YU, and Shenghua Gao. Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation. In Thirty- seventh Conference on Neural Information Processing ...
2023
-
[90]
Driv- able 3d gaussian avatars
Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollh ¨ofer, Justus Thies, and Javier Romero. Driv- able 3d gaussian avatars. In International Conference on 3D Vision (3DV), 2025. 3 Disentangled Clothed Avatar Generation with Layered Representation Supplementary ...
2025
-
[91]
Supplementary Video We provide a supplementary video for quick understanding of our method. The video includes: • A brief introduction of our method; • Results of unconditional generation and decomposition; • Results of novel pose animation; • Results of component transfer
-
[92]
Network Architecture Layered UV Feature Plane and Shared Decoders
Implementation Details 7.1. Network Architecture Layered UV Feature Plane and Shared Decoders. The size of layered UV feature plane is 12 × 128 × 384, where we concatenate the three-layer Gaussian-based UV feature plane width-wise instead of channel-wise following [72]. Two sh...
-
[93]
Limitations and Discussions 8.1. Limitations (1) Due to the segmentation-map-based supervision and SMPLX-based templates, the performance of our method is affected by the accuracy of the estimated segmentation map and SMPLX parameters. Eliminating inaccurate segmenta- tion res...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.