REVIEW 4 major objections 5 minor 54 references
GGAvatar: Reconstructing Garment-Separated 3D Gaussian Splatting Avatars from Monocular Video
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read GGAvatar reconstructs a clothed human avatar from a monocular video as separate body and garment Gaussian layers, enabling clothing transfer and color editing.
desk verdict A promising system with a genuinely new garment-template initialization, but the central separation claim is never measured and the template chain is unvalidated; worth a careful revision, not a rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key object is a pair of parametric templates in a shared canonical space: the SMPL body mesh and per-garment meshes generated by the ISP model from a front-view SCHP segmentation. Vertices extracted from these templates become the initial means of 3D Gaussian components. The mechanism that carries the argument is the phased training and deformation: garment and body Gaussians are optimized separately first, then jointly, with an as-isometric-as-possible deformation regularizer and a collision loss that keep the layers from merging or intersecting. Rendering uses 3D Gaussian Splatting, which makes the whole pipeline differentiable, fast to train, and fast to render.
What would settle it
Take a monocular video of a person wearing loose or heavily wrinkled clothing, run the front-view template extraction, and compare the resulting avatar's garment geometry against a high-fidelity multi-view reconstruction of the same person. If the garment layer visibly deviates from the ground truth or the body and garment Gaussians intersect when reposed, the central claim of robust separation fails. A simpler test is to rotate the choice of front-view frame within the same video and measure how much the final avatar's garment shape and transfer quality change; large variation would weaken the method's reliability.
Extended reading notes
Core claim
The central claim is that by constructing garment templates from a single front-view image—using SCHP human parsing and the ISP implicit sewing pattern model—and aligning them with an SMPL body in canonical space, GGAvatar can initialize garment Gaussians separately from body Gaussians. A two-stage training strategy (isolation training on each component, then joint training with collision and isometry regularizers) prevents the point sets from intersecting while preserving fine texture. The deformation field, driven by learnable skinning weights on a shared skeleton, moves both body and garment Gaussians into arbitrary poses while maintaining their separation. The paper demonstrates that this yields a thoroughly decoupled avatar where individual garments can be transferred to other bodies or recolored, and that the reconstruction quality and speed exceed that of existing decoupled and non-decoupled monocular avatar models.
Load-bearing premise
The load-bearing premise is that a single front-view image, processed by pre-trained SCHP parsing and the ISP model, yields a garment template that is geometrically and topologically correct and aligned with the SMPL body in canonical space; if this template is wrong, the Gaussian initialization, separation, and all downstream editing and deformation results inherit the error.
Editorial extensions
If this is right
- A single monocular video suffices to produce a high-quality, garment-separated 3D avatar in about 20 minutes on one GPU.
- Individual garments, not just the whole outfit, can be transferred to another person's avatar, enabling fine-grained virtual try-on.
- Per-garment color editing can be done by simply specifying RGB values, without manual spherical harmonic manipulation.
- The avatar can be reposed or animated through a shared skeleton while keeping body and clothing layers distinct, which supports realistic novel-pose synthesis.
- Reconstruction quality is competitive with state-of-the-art 3DGS avatar models and significantly faster than NeRF-based alternatives.
Reading between the lines
- The robustness of the entire pipeline hinges on the front-view garment template: if the SCHP segmentation or ISP reconstruction is inaccurate for loose or occluded clothing, the error would propagate to the final avatar, but the paper does not measure this sensitivity.
- The method might extend to multi-frame template aggregation—averaging or refining the garment mesh over many video frames instead of a single view—which could improve geometric accuracy for difficult clothing.
- Because the body and garments are separate Gaussian sets, future work could apply physics-based cloth simulation or per-garment animation without re-training, a step the paper does not explore.
- The same template-initialization and phased-training recipe could generalize to other articulated objects with layered components, not just human bodies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GGAvatar, a monocular-video avatar reconstruction method based on 3D Gaussian Splatting with explicitly separated body and garment components. The method initializes garment Gaussians from a single-view implicit sewing pattern (ISP) template aligned to SMPL, then uses a two-phase training scheme (isolation and joint) with mask, S3IM, isometry, and collision losses to keep the components separated while optimizing reconstruction quality. The authors evaluate against Neural Body, InstantAvatar, GaussianAvatar, GART, and SCARF on the People Snapshot and ZJU-MoCap datasets, and demonstrate applications in animation, clothing transfer, and colour editing. The central claims are (i) state-of-the-art or comparable reconstruction quality with much faster training than NeRF-based alternatives, and (ii) a 'thorough separation' of distinct garments from monocular video, which the authors describe as potentially a first in the field.
Significance. If the central claims hold, GGAvatar would be a practical contribution: it produces an editable, garment-separated avatar from a single monocular video in about 20 minutes, with code released. The use of existing parametric templates (SMPL, ISP) to seed per-garment Gaussian sets is a reasonable design choice that avoids expensive 3D garment ground truth at training time. The reported rendering speed and training time are strong practical advantages over NeRF-based decoupling methods, and the applications shown (per-garment transfer and colour editing) are visually compelling. However, the significance as a scientific contribution is currently limited because the load-bearing claim of garment separation is not quantitatively evaluated, and the reconstruction-quality advantage over strong 3DGS baselines is not consistent across metrics and subjects.
major comments (4)
- [§5.2, Table 1] The paper states that GGAvatar achieves 'superior quality' and outperforms other methods, but the quantitative evidence in Table 1 is inconsistent with this claim. GGAvatar has higher (worse) LPIPS than Neural Body on three of the four People-Snapshot subjects (e.g., 0.0349 vs. 0.0326 for male-3-casual, 0.0509 vs. 0.0423 for male-4-casual) and lower PSNR than GaussianAvatar on female-3-casual (27.96 vs. 29.55) and female-4-casual (30.29 vs. 30.84). The conclusion that the model is 'on par with' or 'superior to' 3DGS baselines therefore depends on which metric and subject one emphasizes. The authors should either report statistical significance across multiple runs/subjects, provide a per-subject discussion, or soften the claim.
- [§5 (Garment Reconstruction Quality) and Fig. 4] The paper's core contribution is 'thorough separation' of garments, stated in the Introduction and Conclusion, yet no quantitative metric measures separation quality. Table 1 reports only whole-body PSNR/SSIM/LPIPS, and the garment evaluation is limited to qualitative side-by-side comparisons with SCARF in Fig. 4. A direct measurement is needed, such as per-garment IoU against hand-labeled or SMPL-X/SCHP-derived ground-truth garment regions, Chamfer distance between extracted garment Gaussians and a 3D garment scan, or a user study on garment-boundary accuracy. Without such evidence, the central claim of the paper remains untested.
- [§4.1 (Garment Templates Estimation)] The entire garment-separation pipeline depends on a single front-view image being processed by SCHP parsing and the ISP model to produce the canonical garment template T(c)_can. The paper states that the front view 'must be selected and aligned' with the SCHP segmentation, but gives no criterion for selecting this frame, no accuracy measure for the SCHP+ISP output, and no robustness study. If the selected frame is not truly frontal, or if SCHP mislabels garment boundaries, or if ISP returns an incorrect stitch order or layer topology, the errors propagate into the Gaussian initialization, the mask-supervised isolation loss of Eq. (10), and all downstream editing results, and the optimization cannot reassign a Gaussian to a different garment (the assignment is fixed after initialization). The authors should validate the template-estimation step, e.g., by comparing the ISP mesh against a multi-view reconstruction or a manual garment annotation, and report sensitivity to the choice of front-view frame.
- [§4.4, Eq. (10)] The isolation loss L(c)_mask is supervised by 2D SCHP segmentation masks, which do not necessarily correspond to a consistent 3D layer decomposition of the garment set, especially under self-occlusion and loose clothing. Because the separation is imposed by these masks and by the fixed initialization rather than measured as an output property, the reported disentanglement may be substantially inherited from the external supervision. The paper should either provide evidence that the final separated Gaussians are physically correct in 3D (e.g., by checking interpenetration between garment and body layers, or by inspecting cross-sectional slices), or at least discuss this limitation explicitly. This would help the reader judge how much of the claimed 'thorough separation' emerges from the model and how much is prescribed by the input templates.
minor comments (5)
- [Abstract and Introduction] The phrase 'potentially a first in this field, to my knowledge' is hedged but still a strong novelty claim; it would be more appropriate to cite the closest prior decoupling work (e.g., SCARF, DELTA, LayGA) and state precisely what is new beyond them, rather than rely on a field-wide search claim.
- [§5.2, Table 1] The table legend says 'female-4-causal' while the text says 'female-4-casual'; the spelling should be consistent. Also, the 'time' column mixes units (days, minutes, hours) and does not list hardware for all methods, which makes the speed comparison harder to interpret.
- [§4.4, Eq. (9)] The symbols d(·,·) are used for both the L2 distance between positions and the L2 distance between covariance matrices, but the two terms have different units and scales. The authors should clarify how the two distance terms are normalized or weighted when combined in one sum.
- [§5.3, Table 2] In Table 2, the row 'w/o L_iso' reports LPIPS 0.0347, which is slightly lower (better) than the full model's 0.0343? Actually 0.0347 is higher, so that is consistent, but the text says omission leads to a decline in rendering quality while LPIPS worsens; please verify the direction of the metric in the text to avoid confusion.
- [Fig. 2] The framework figure is dense and the flow lines (red, orange, blue) are hard to follow in the printed version; labeling each module with a letter or number and referencing them in the text would improve readability.
Circularity Check
No significant circularity: the pipeline is self-contained and externally benchmarked, though the separation claim rests on input templates and is not quantitatively validated.
full rationale
GGAvatar's derivation chain is not circular. The garment templates are seeded by external SCHP segmentation and ISP reconstruction (Sec ́4.1), and the per-garment Gaussian sets are then optimized under mask supervision from the same SCHP labels (Eq. 10). This is supervised initialization rather than a self-referential prediction: the output is a rendered avatar evaluated against held-out views and poses (Table 1, Fig. 3) with PSNR/SSIM/LPIPS, and the garment transfer/editing (Fig. 6) are applications of the fixed template partition, not quantities defined as the inputs. There are no load-bearing self-citations: the cited prior works (SCHP, ISP, SCARF, etc.) are external and independently published. The absence of a quantitative separation metric and the sensitivity of the pipeline to the single-view template are validation and robustness gaps, not evidence that any equation reduces to its own input or that a fitted constant is renamed as a prediction. Hence no circular step is exhibited.
Assumptions & free parameters
free parameters (5)
- Loss weights lambda_2 to lambda_6 =
Not stated in main text; deferred to missing supplementary
- SMPL shape beta and pose theta per subject =
Estimated with FrankMocap from the input video
- Per-subject body vertex offsets O =
Learned during training
- Per-Gaussian skinning weight offsets Delta w_j =
Learned during training
- Gaussian attributes (mean, rotation, scale, opacity, SH coefficients) =
Learned by gradient descent
assumptions (5)
- domain assumption SMPL is an adequate parametric model for both body and garment deformation.
- domain assumption Pre-trained ISP can produce a correct 3D garment mesh from a single front-view image.
- domain assumption SCHP human parsing masks correctly identify garment regions in every frame.
- ad hoc to paper The two-phase training schedule keeps Gaussian sets separated without intersection.
- standard math Volume rendering via 3DGS alpha compositing is differentiable and accurate enough for the task.
Cite this review
Pith. "Pith review of GGAvatar: Reconstructing Garment-Separated 3D Gaussian Splatting Avatars from Monocular Video." pith.science (2026). https://pith.science/paper/THIMKD7K
@misc{pith2026241109952,
author = {Pith},
title = {Pith review of: GGAvatar: Reconstructing Garment-Separated 3D Gaussian Splatting Avatars from Monocular Video},
year = {2026},
howpublished = {\url{https://pith.science/paper/THIMKD7K}},
note = {Machine review of arXiv:2411.09952}
}
read the original abstract
Avatar modelling has broad applications in human animation and virtual try-ons. Recent advancements in this field have focused on high-quality and comprehensive human reconstruction but often overlook the separation of clothing from the body. To bridge this gap, this paper introduces GGAvatar (Garment-separated 3D Gaussian Splatting Avatar), which relies on monocular videos. Through advanced parameterized templates and unique phased training, this model effectively achieves decoupled, editable, and realistic reconstruction of clothed humans. Comparative evaluations with other costly models confirm GGAvatar's superior quality and efficiency in modelling both clothed humans and separable garments. The paper also showcases applications in clothing editing, as illustrated in Figure 1, highlighting the model's benefits and the advantages of effective disentanglement. The code is available at https://github.com/J-X-Chen/GGAvatar/.
Figures
Reference graph
Works this paper leans on
-
[1]
Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Ger- ard Pons-Moll. 2018. Detailed human avatars from monocular video. In 2018 International Conference on 3D Vision (3DV) . IEEE, 98–109
2018
-
[2]
Katherine L Bouman, Bei Xiao, Peter Battaglia, and William T Freeman. 2013. Estimating the material properties of fabric from video. InProceedings of the IEEE international conference on computer vision . 1984–1991
work page 2013
-
[3]
Honghu Chen, Yuxin Yao, and Juyong Zhang. 2024. Neural-ABC: Neural Paramet- ric Models for Articulated Body With Clothes. IEEE Transactions on Visualization and Computer Graphics (2024)
work page 2024
-
[4]
Jianchuan Chen, Ying Zhang, Di Kang, Xuefei Zhe, Linchao Bao, Xu Jia, and Huchuan Lu. 2021. Animatable neural radiance fields from monocular rgb videos. arXiv preprint arXiv:2106.13629 (2021)
arXiv 2021
-
[5]
Xu Chen, Tianjian Jiang, Jie Song, Max Rietmann, Andreas Geiger, Michael J Black, and Otmar Hilliges. 2023. Fast-SNARF: A fast deformer for articulated neural fields. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
work page 2023
-
[6]
Enric Corona, Albert Pumarola, Guillem Alenya, Gerard Pons-Moll, and Francesc Moreno-Noguer. 2021. Smplicit: Topology-aware generative model for clothed people. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11875–11885
work page 2021
-
[7]
Luca De Luigi, Ren Li, Benoit Guillard, Mathieu Salzmann, and Pascal Fua. 2023. DrapeNet: Garment Generation and Self-Supervised Draping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
work page 2023
-
[8]
Yao Feng, Weiyang Liu, Timo Bolkart, Jinlong Yang, Marc Pollefeys, and Michael J Black. 2023. Learning disentangled avatars with hybrid 3d representations. arXiv preprint arXiv:2309.06441 (2023)
arXiv 2023
Show all 54 references
-
[9]
Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J Black, and Timo Bolkart
-
[10]
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. 2022. Plenoxels: Radiance fields without neural net- works. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5501–5510
2022
-
[11]
Chen Geng, Sida Peng, Zhen Xu, Hujun Bao, and Xiaowei Zhou. 2023. Learning neural volumetric representations of dynamic humans in minutes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8759– 8770
2023
-
[12]
Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Sheng- ping Zhang, and Liqiang Nie. 2024. Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians. In Proceedings of the IEEE/CVF Conference on Computer Vision a...
2024
-
[13]
Shoukang Hu, Tao Hu, and Ziwei Liu. 2024. Gauhuman: Articulated gaussian splatting from monocular human videos. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition . 20418–20431
2024
-
[14]
Boyi Jiang, Juyong Zhang, Yang Hong, Jinhao Luo, Ligang Liu, and Hujun Bao
-
[15]
Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. 2023. Instantavatar: Learning avatars from monocular video in 60 seconds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16922–16932
2023
-
[16]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis
-
[17]
Martin Kilian, Niloy J Mitra, and Helmut Pottmann. 2007. Geometric modelling in shape space. In ACM SIGGRAPH 2007 papers. 64–es
2007
-
[18]
Taeksoo Kim, Byungjun Kim, Shunsuke Saito, and Hanbyul Joo. 2024. GALA: Generating Animatable Layered Assets from a Single Scan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1535–1545
2024
-
[19]
Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. 2024. Hugs: Human gaussian splats. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 505–515
2024
-
[20]
Jiahui Lei, Yufu Wang, Georgios Pavlakos, Lingjie Liu, and Kostas Daniilidis
-
[21]
Peike Li, Yunqiu Xu, Yunchao Wei, and Yi Yang. 2020. Self-Correction for Human Parsing. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020). https://doi.org/10.1109/TPAMI.2020.3048039
2020
-
[22]
Ren Li, Benoît Guillard, and Pascal Fua. 2024. Isp: Multi-layered garment draping with implicit sewing patterns. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[23]
Ren Li, Benoit Guillard, Edoardo Remelli, and Pascal Fua. 2022. Dig: Draping implicit garment over the human body. In Proceedings of the Asian Conference on Computer Vision. 2780–2795
2022
-
[24]
Siyou Lin, Zhe Li, Zhaoqi Su, Zerong Zheng, Hongwen Zhang, and Yebin Liu
-
[25]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Transactions on Graphics 34, 6 (2015)
2015
-
[26]
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2024. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. In 2024 International Conference on 3D Vision (3DV) (2024)
2024
-
[27]
B Mildenhall, PP Srinivasan, M Tancik, JT Barron, R Ramamoorthi, and R Ng
-
[28]
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. In- stant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG) 41, 4 (2022), 1–15
2022
-
[29]
In ACM SIGGRAPH 2024 Conference Papers
LayGA: Layered Gaussian Avatars for Animatable Clothing Transfer. In ACM SIGGRAPH 2024 Conference Papers . 1–11
2024
-
[30]
Chaitanya Patel, Zhouyingcheng Liao, and Gerard Pons-Moll. 2020. Tailornet: Predicting clothing in 3d as a function of human pose, shape and garment style. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7365–7375
2020
-
[31]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. 2019. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern reco...
2019
-
[32]
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. 2021. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceedings of the IEEE/CVF Conference on Computer Visi...
2021
-
[33]
In European conference on computer vision
Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision
-
[34]
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang
-
[35]
Ahmed AA Osman, Timo Bolkart, and Michael J Black. 2020. Star: Sparse trained articulated human body regressor. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 598–613
2020
-
[36]
Steven M Seitz, Brian Curless, James Diebel, Daniel Scharstein, and Richard Szeliski. 2006. A comparison and evaluation of multi-view stereo reconstruction algorithms. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), Vol. 1. IEEE, 519–528
2006
-
[37]
Carsten Stoll, Juergen Gall, Edilson De Aguiar, Sebastian Thrun, and Christian Theobalt. 2010. Video-based reconstruction of animatable human characters. ACM Transactions on Graphics (TOG) 29, 6 (2010), 1–10
2010
-
[38]
Richard Szeliski, Steven Gortler, Radek Grzeszczuk, and Michael F Cohen. 1996. The lumigraph. In Proceedings of the 23rd annual conference on computer graphics and interactive techniques (SIGGRAPH 1996) . 43–54
1996
-
[39]
Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael J Black. 2017. ClothCap: Seamless 4D clothing capture and retargeting. ACM Transactions on Graphics (ToG) 36, 4 (2017), 1–15
2017
-
[40]
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. 2022. Humannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16210–16220
2022
-
[41]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5020–5030
-
[42]
Yu Rong, Takaaki Shiratori, and Hanbyul Joo. 2021. FrankMocap: A Monocular 3D Whole-Body Pose Estimation System via Regression and Integration. In IEEE International Conference on Computer Vision Workshops
2021
-
[43]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[44]
Yang Zheng, Qingqing Zhao, Guandao Yang, Wang Yifan, Donglai Xiang, Florian Dubost, Dmitry Lagun, Thabo Beeler, Federico Tombari, Leonidas Guibas, et al
-
[45]
Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. 2002. EWA splatting. IEEE Transactions on Visualization and Computer Graphics 8, 3 (2002), 223–238
2002
-
[46]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing 13, 4 (2004), 600–612
2004
-
[48]
Donglai Xiang, Fabian Prada, Timur Bagautdinov, Weipeng Xu, Yuan Dong, He Wen, Jessica Hodgins, and Chenglei Wu. 2021. Modeling clothing as a separate layer for an animatable human avatar. ACM Transactions on Graphics (TOG) 40, 6 (2021), 1–15
2021
-
[49]
Zeke Xie, Xindi Yang, Yujie Yang, Qi Sun, Yixiang Jiang, Haoran Wang, Yun- feng Cai, and Mingming Sun. 2023. S3im: Stochastic structural similarity and its unreasonable effectiveness for neural fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision ...
2023
-
[53]
arXiv preprint arXiv:2404.04421 (2024)
PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations. arXiv preprint arXiv:2404.04421 (2024)
2024 arXiv
-
[2018]
In Proceedings of the IEEE conference on computer vision and pattern recognition
The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition . 586–595
-
[2020]
In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16
Bcnet: Learning body and cloth shape from a single image. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16 . Springer, 18–35
2020
-
[2022]
In SIGGRAPH Asia 2022 Conference Papers
Capturing and animation of body and clothing from monocular video. In SIGGRAPH Asia 2022 Conference Papers . 1–9
2022
-
[2023]
ACM Transactions on Graphics 42, 4 (2023)
3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics 42, 4 (2023)
2023
-
[2024]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Gart: Gaussian articulated template models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19876–19887
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.