REVIEW 3 major objections 5 minor 78 references
A single photo is enough to build a photorealistic 3D head avatar that stays consistent from the side and back and can be driven by expression parameters in real time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 15:52 UTC pith:5AEVCPEH
load-bearing objection Solid systems paper: generate-first 3DGS + FLAME bind works, but quality is capped by frozen LGM and the gains are real but modest. the 3 major comments →
S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that a diffusion-guided feed-forward 3D Gaussian splat generator, followed by Chamfer-plus-landmark fitting of FLAME and an inverse-distance binding template (with optional density-aware scale adaptation), yields animatable head avatars from a single image whose novel-view and novel-expression renderings are more 3-D consistent and higher quality than prior one-shot mesh, GAN, and neural methods.
What carries the argument
The binding template B: a sparse matrix of normalized inverse-distance weights from each Gaussian splat to its nearest FLAME vertices. Multiplying B transpose by FLAME vertex displacements (and density ratios) deforms and rescales the canonical splats so expression and pose changes drive the avatar in real time.
Load-bearing premise
The whole pipeline assumes the diffusion model already produced a complete, multi-view-consistent set of initial Gaussian splats; if that first stage is wrong, later fitting and binding cannot recover a good avatar.
What would settle it
Take the same single input image, deliberately degrade or replace the initial splat stage (for example by removing the text prompts that suppress the Janus effect), run the identical fitting and binding steps, and check whether novel side/back views and expression transfers still match ground-truth NeRSemble frames at the reported LPIPS/PSNR levels.
If this is right
- Personal photorealistic head avatars become creatable from one ordinary photo rather than multi-view video or studio capture.
- Real-time VR/AR head rendering from unobserved angles (side, back) is feasible without per-subject neural optimization at inference.
- FLAME expression and pose parameters alone are sufficient to drive the avatar, so any FLAME tracker or HMD sensor pipeline can animate it.
- The modular split lets stronger future single-image Gaussian generators be swapped in without redesigning the animation stage.
Where Pith is reading between the lines
- If initial-splat quality is the bottleneck, the same binding idea could be stress-tested on non-head objects once a reliable single-image Gaussian generator exists for those categories.
- Density-aware splat scaling may transfer to other mesh-rigged Gaussian avatars (hands, bodies) where surface stretch under articulation is the main visual failure mode.
- Because generation is feed-forward and binding is sparse, on-device or edge reconstruction of personal avatars becomes a realistic systems target once the diffusion backbone is distilled.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. S-Avatar proposes a three-stage pipeline for one-shot animatable 3D head avatars: (1) feed-forward generation of a canonical 3D Gaussian splat set S from a single image via ImageDream multi-view diffusion and an LGM-style asymmetric U-Net; (2) gradient-based fitting of FLAME (shape, expression, pose, and rigid transform/scale) to S using a weighted Chamfer + landmark objective (Eqs. 7–9); and (3) construction of a sparse inverse-distance binding template B (COO) that deforms splat positions by nearest-FLAME-vertex displacements and adapts splat scales via relative mesh density ratios (Eqs. 13–17). The avatar is then driven in real time by novel FLAME parameters. On 45 NeRSemble subjects across 16 views, the method reports better LPIPS/PSNR/SSIM than ROME, Portrait4D-v1/v2, VOODOO, and GPAvatar under a per-method face-crop protocol, with ablations on N_closest, scale adaptation, loss weights, and COO sparsity, plus qualitative extreme-view, cross-reenactment, in-the-wild, and VR demo results.
Significance. One-shot photorealistic, expression-controllable head avatars with real-time 3DGS rendering remain practically important for VR/AR. The decoupled design—build a diffusion-guided canonical splat field first, then attach FLAME control via a fixed geometric binding—is a clear and reusable systems contribution relative to methods that optimize radiance fields or Gaussians directly from sparse views. Strengths include a working real-time path (COO binding, ~35 FPS with scale adaptation), honest limitations, ablations that isolate N_closest and density-based scale updates, and demonstrated VR/interactive tooling. If the extreme-view consistency and identity-preserving animation claims hold under stricter full-head evaluation, the work would be a useful advance in accessible avatar creation. Credit is due for modularity (LGM can be swapped) and for releasing a project page with demos.
major comments (3)
- [§3.2, Limitations, Appendix Fig. 11] Limitations and §3.2 state that final quality is permanently capped by the frozen initial splat set S from LGM/ImageDream; fitting (§3.3) only aligns FLAME and binding (§3.4) only applies fixed IDW displacements plus density-ratio scales—neither can invent missing side/back geometry or undo Janus/inconsistency. The central claims of extreme-view 3D consistency and superior novel-view realism therefore inherit unquantified upstream error. The manuscript shows prompt-related failures in Appendix Fig. 11 but reports no failure rate, multi-view consistency metric, or oracle upper bound of S on the 45-subject NeRSemble split. Without that characterization (or a quantitative full-head side/back metric), the extreme-view narrative rests mainly on qualitative Figs. 5–6 and cannot be independently verified.
- [§4.1–4.2, Table 1] Table 1 gains are measured after per-method face cropping “so that the comparison is performed within the same facial region” (§4.1–4.2). This protocol can mask full-head failures (ears, scalp, back) that the paper’s extreme-view claim specifically advertises. Absolute PSNR remains modest (~16.1). Please add (i) a fixed full-head or head-silhouette evaluation protocol shared across methods, and/or (ii) region-specific metrics for side/back views, plus error bars or paired significance tests over the 45 subjects so that the SOTA ranking is statistically supported rather than point-estimate only.
- [§3.4, Eqs. 10–17, Table 2] The binding model (Eqs. 10–17) assumes that local inverse-distance weighting of N_closest FLAME vertices plus a density-ratio scale update adequately approximates non-rigid facial deformation of Gaussians. Table 2 only sweeps N_closest for aggregate PSNR; there is no stress test on large expression magnitudes, jaw/neck extremes, or topological failure cases (lip contact, cheek stretch). Given that animation quality is a primary claim, please quantify expression-range fidelity (e.g., error vs. held-out extreme expressions, or comparison against a learned residual deformer) and discuss when IDW binding breaks.
minor comments (5)
- [§3.1] Eq. (2) writes Σ = R S S^T R^T but the surrounding text uses s for scale; keep notation consistent with Kerbl et al. (scale vector vs. matrix).
- [Fig. 2] Several figure glyphs and FLAME parameter symbols in Fig. 2 are corrupted or use non-ASCII lookalikes (e.g., β, ψ, θ render inconsistently), which hurts readability in the PDF.
- [§4.1] Implementation details give Adam β1/β2 and 60 epochs but not wall-clock breakdown for the diffusion multi-view stage vs. fitting on the same hardware as the baselines; a single end-to-end timing table would help.
- [§2.3] Related work cites many concurrent one-shot Gaussian head methods; a short explicit comparison table (input type, canonical vs. direct optimize, real-time, full-head) would clarify positioning versus GPAvatar, SEGA, LAM, etc.
- [Abstract, Acknowledgments, References] Typos/style: “S-A vatar” spacing artifacts in the abstract block; “we was supported” in Acknowledgments; arXiv IDs in the 2602.* range for self-citations look premature—double-check bibliographic entries.
Circularity Check
No circularity: pipeline is modular engineering; metrics come from external NeRSemble GT and independent baselines, not from quantities defined by the fit.
full rationale
S-Avatar’s derivation is a three-stage engineering pipeline (LGM/ImageDream feed-forward initial splats S → Chamfer+landmark FLAME fit → fixed sparse IDW binding template B with density-based scale). None of the load-bearing equations reduce a claimed prediction to its own inputs. Binding weights w_ij are computed from Euclidean distances between fixed S and fitted FLAME vertices (Eqs. 10–17); they are not fitted to LPIPS/PSNR/SSIM. Novel-view/expression scores in Table 1 are measured against held-out multi-view NeRSemble ground truth and compared to external baselines (ROME, Portrait4D, VOODOO, GPAvatar). Ablations (N_closest, scale adaptation, λ_c/λ_l, COO) vary internal hyperparameters and report the same external metrics; they do not redefine the target. Self-citations (OFERA, lab VR demos) support only application demos in the appendix and do not justify the central quantitative claims or uniqueness of the method. The paper’s own Limitations explicitly treat initial splat quality as an upstream dependency rather than claiming to derive it from the evaluation. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.
Axiom & Free-Parameter Ledger
free parameters (5)
- λ_c, λ_l (fitting loss weights) =
0.01 / 0.99
- N_closest =
10
- FLAME optimization learning rates and 60 epochs =
lr shape 0.005; expr/pose 0.0005; transform 0.01; 60 epochs
- Positive/negative text prompts for ImageDream/LGM =
Fixed English positive/negative strings in appendix
- ε in inverse-distance weights =
unspecified small constant
axioms (5)
- domain assumption Pretrained ImageDream+LGM produce view-consistent initial 3DGS adequate for heads from one RGB crop.
- domain assumption FLAME shape/pose/expression space plus rigid transform can align to and drive head geometry including regions FLAME does not model well (hair, ears detail).
- ad hoc to paper Local inverse-distance weighting of nearest FLAME vertices plus density-ratio scale updates approximates non-rigid facial deformation of Gaussians.
- ad hoc to paper Per-method face cropping yields a fair comparison across heterogeneous avatar outputs.
- standard math Standard 3DGS alpha-blending and FLAME LBS definitions hold as in Kerbl et al. and Li et al.
invented entities (2)
-
Binding template B (sparse NFLAME×Nsplat weight matrix in COO form)
no independent evidence
-
Relative-density Gaussian scale adaptation (σ from P and P')
no independent evidence
read the original abstract
We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation module and strategies for animating 3D Gaussian Splatting (3DGS). While single-image head avatar reconstruction is crucial for lifelike Virtual Reality (VR) applications, existing approaches often struggle to preserve 3D consistency under unseen viewpoints. S-Avatar addresses this limitation through a three-stage pipeline. First, a high-resolution 3DGS is synthesized directly from a single image using a diffusion-based Gaussian splat generation module. Next, the parametric head model FLAME is aligned with the generated 3DGS by optimizing its parameters and spatial transformations. Finally, to adapt the 3DGS to FLAME variations, we construct a binding template that encodes the spatial relationship between the initial splats and FLAME. The dynamic 3D head avatar can then be rendered in real time by deforming the 3DGS with the binding template. By combining diffusion-guided canonical 3DGS generation with FLAME-based control, our method achieves efficient and accurate reconstruction with enhanced 3D consistency. Evaluations on public datasets demonstrate that S-Avatar outperforms state-of-the-art methods in novel-view and expression generation, achieving superior realism and consistency. Consequently, our approach represents a significant advance in accessible avatar creation, applicable to a wide range of VR/AR applications. The project page is available at https://github.com/hailsong/savatar.
Figures
Reference graph
Works this paper leans on
-
[1]
Aliev, A
K.-A. Aliev, A. Sevastopolsky, M. Kolos, D. Ulyanov, and V . Lempit- sky. Neural point-based graphics. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Pro- ceedings, Part XXII 16, pp. 696–712. Springer, 2020. 2
2020
-
[2]
Alldieck, M
T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons-Moll. Video based reconstruction of 3d people models. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8387– 8397, 2018. 2
2018
-
[3]
Athar, Z
S. Athar, Z. Shu, and D. Samaras. Flame-in-nerf: Neural control of radiance fields for free view face animation. In2023 IEEE 17th In- ternational Conference on Automatic Face and Gesture Recognition (FG), pp. 1–8. IEEE, 2023. 2
2023
-
[4]
Blanz and T
V . Blanz and T. Vetter. A morphable model for the synthesis of 3d faces. InSeminal Graphics Papers: Pushing the Boundaries, V olume 2, pp. 157–164. 2023. 1, 2
2023
-
[5]
M. C. B ¨uhler, K. Sarkar, T. Shah, G. Li, D. Wang, L. Helminger, S. Orts-Escolano, D. Lagun, O. Hilliges, T. Beeler, et al. Preface: A data-driven volumetric prior for few-shot ultra high-resolution face synthesis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3402–3413, 2023. 2
2023
-
[6]
C. Cao, T. Simon, J. K. Kim, G. Schwartz, M. Zollhoefer, S.-S. Saito, S. Lombardi, S.-E. Wei, D. Belko, S.-I. Yu, et al. Authentic volumetric avatars from a phone scan.ACM Transactions on Graphics (TOG), 41(4):1–19, 2022. 2
2022
-
[7]
G. Chen and W. Wang. A survey on 3d gaussian splatting.arXiv preprint arXiv:2401.03890, 2024. 3
Pith/arXiv arXiv 2024
-
[8]
X. Chen, M. Mihajlovic, S. Wang, S. Prokudin, and S. Tang. Mor- phable diffusion: 3d-consistent diffusion for single-image avatar cre- ation. InProceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pp. 10359–10370, 2024. 2
2024
-
[9]
Y . Chen, L. Wang, Q. Li, H. Xiao, S. Zhang, H. Yao, and Y . Liu. Monogaussianavatar: Monocular gaussian point-based head avatar. arXiv preprint arXiv:2312.04558, 2023. 1, 2
Pith/arXiv arXiv 2023
-
[10]
Z. Chen, F. Wang, Y . Wang, and H. Liu. Text-to-3d using gaussian splatting. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 21401–21412, 2024. 2
2024
-
[11]
X. Chu, Y . Li, A. Zeng, T. Yang, L. Lin, Y . Liu, and T. Harada. Gpa- vatar: Generalizable and precise head avatar from image (s).arXiv preprint arXiv:2401.10215, 2024. 2, 6
Pith/arXiv arXiv 2024
-
[12]
Dan ˇeˇcek, M
R. Dan ˇeˇcek, M. J. Black, and T. Bolkart. Emoca: Emotion driven monocular face capture and animation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20311–20322, 2022. 1, 2
2022
-
[13]
Y . Deng, D. Wang, X. Ren, X. Chen, and B. Wang. Portrait4d: Learn- ing one-shot 4d head avatar synthesis using synthetic data. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 7119–7130, 2024. 2, 6
2024
-
[14]
Y . Deng, D. Wang, and B. Wang. Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer. InEuropean Conference on Computer Vision, pp. 316–333. Springer, 2024. 2, 6
2024
-
[15]
Y . Deng, J. Yang, S. Xu, D. Chen, Y . Jia, and X. Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pp. 0–0, 2019. 1, 2
2019
-
[16]
H. Fan, H. Su, and L. J. Guibas. A point set generation network for 3d object reconstruction from a single image. InProceedings of the IEEE conference on computer vision and pattern recognition, pp. 605–613,
-
[17]
Y . Feng, H. Feng, M. J. Black, and T. Bolkart. Learning an animatable detailed 3d face model from in-the-wild images.ACM Transactions on Graphics (ToG), 40(4):1–13, 2021. 1, 2
2021
-
[18]
Gafni, J
G. Gafni, J. Thies, M. Zollhofer, and M. Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 8649–8658, 2021. 2
2021
-
[19]
Garrido, M
P. Garrido, M. Zollh ¨ofer, D. Casas, L. Valgaerts, K. Varanasi, P. P´erez, and C. Theobalt. Reconstruction of personalized 3d face rigs from monocular video.ACM Transactions on Graphics (TOG), 35(3):1– 15, 2016. 1, 2
2016
-
[20]
Giebenhain, T
S. Giebenhain, T. Kirschstein, M. R ¨unz, L. Agapito, and M. Nießner. Npga: Neural parametric gaussian avatars. InSIGGRAPH Asia 2024 Conference Papers, pp. 1–11, 2024. 2
2024
-
[21]
C. Guo, Z. Su, J. Wang, S. Li, X. Chang, Z. Li, Y . Zhao, G. Wang, and R. Huang. Sega: Drivable 3d gaussian head avatar from a single image.arXiv preprint arXiv:2504.14373, 2025. 2
arXiv 2025
-
[22]
X. He, J. Chen, S. Peng, D. Huang, Y . Li, X. Huang, C. Yuan, W. Ouyang, and T. He. Gvgen: Text-to-3d generation with volumet- ric representation. InEuropean Conference on Computer Vision, pp. 463–479. Springer, 2024. 3
2024
-
[23]
Y . He, X. Gu, X. Ye, C. Xu, Z. Zhao, Y . Dong, W. Yuan, Z. Dong, and L. Bo. Lam: Large avatar model for one-shot animatable gaussian head.arXiv preprint arXiv:2502.17796, 2025. 2
Pith/arXiv arXiv 2025
-
[24]
C. A. R. Hoare. Algorithm 64: quicksort.Communications of the ACM, 4(7):321, 1961. 5
1961
-
[25]
M. R. Hugues and S. G. Petiton. Sparse matrix formats evaluation and optimization on a gpu. In2010 IEEE 12th International Conference on High Performance Computing and Communications (HPCC), pp. 122–129. IEEE, 2010. 6
2010
-
[26]
Huynh-Thu and M
Q. Huynh-Thu and M. Ghanbari. Scope of validity of psnr in im- age/video quality assessment.Electronics letters, 44(13):800–801,
-
[27]
Z. Ke, J. Sun, K. Li, Q. Yan, and R. W. Lau. Modnet: Real-time trimap-free portrait matting via objective decomposition. InAAAI,
-
[28]
Kerbl, G
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4):1–14, 2023. 1, 2, 3
2023
-
[29]
Khakhulin, V
T. Khakhulin, V . Sklyarova, V . Lempitsky, and E. Zakharov. Real- istic one-shot mesh-based head avatars. InEuropean Conference on Computer Vision, pp. 345–362. Springer, 2022. 2, 6
2022
-
[30]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 4
Pith/arXiv arXiv 2014
-
[31]
T. Kirschstein, S. Qian, S. Giebenhain, T. Walter, and M. Nießner. Nersemble: Multi-view radiance field reconstruction of human heads. ACM Trans. Graph., 42(4), jul 2023. doi:10.1145/35924556, 7
-
[32]
Y .-C. Lee, Y .-T. Chen, A. Wang, T.-H. Liao, B. Y . Feng, and J.-B. Huang. Vividdream: Generating 3d scene with ambient dynamics. arXiv preprint arXiv:2405.20334, 2024. 3
Pith/arXiv arXiv 2024
-
[33]
J. Li, C. Cao, G. Schwartz, R. Khirodkar, C. Richardt, T. Simon, Y . Sheikh, and S. Saito. Uravatar: Universal relightable gaussian codec avatars. InSIGGRAPH Asia 2024 Conference Papers, pp. 1–11,
2024
-
[34]
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero. Learning a model of facial shape and expression from 4d scans.ACM Trans. Graph., 36(6):194–1, 2017. 1, 2
2017
-
[35]
W. Li, L. Zhang, D. Wang, B. Zhao, Z. Wang, M. Chen, B. Zhang, Z. Wang, L. Bo, and X. Li. One-shot high-fidelity talking-head syn- thesis with deformable neural radiance field. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17969–17978, 2023. 2
2023
-
[36]
X. Li, S. De Mello, S. Liu, K. Nagano, U. Iqbal, and J. Kautz. General- izable one-shot 3d neural head avatar.Advances in Neural Information Processing Systems, 36:47239–47250, 2023. 2
2023
-
[37]
Loper, N
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. Smpl: A skinned multi-person linear model.ACM transactions on graphics (TOG), 34(6):1–16, 2015. 2
2015
-
[38]
W. Lyu, Y . Zhou, M.-H. Yang, and Z. Shu. Facelift: Single image to 3d head with view generation and gs-lrm, 2024. 9
2024
-
[39]
S. Ma, T. Simon, J. Saragih, D. Wang, Y . Li, F. De La Torre, and Y . Sheikh. Pixel codec avatars. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 64–73,
-
[40]
L. Melas-Kyriazi, I. Laina, C. Rupprecht, N. Neverova, A. Vedaldi, O. Gafni, and F. Kokkinos. Im-3d: Iterative multiview diffusion and reconstruction for high-quality 3d generation.arXiv preprint arXiv:2402.08682, 2024. 3
Pith/arXiv arXiv 2024
-
[41]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis.Communications of the ACM, 65(1):99–106, 2021. 1, 2
2021
-
[42]
S. Moon, H. M. Lew, S. Lee, J.-S. Kang, and G.-M. Park. Geoavatar: Adaptive geometrical gaussian splatting for 3d head avatar.arXiv preprint arXiv:2507.18155, 2025. 2
Pith/arXiv arXiv 2025
-
[43]
Y . Mu, X. Zuo, C. Guo, Y . Wang, J. Lu, X. Wu, S. Xu, P. Dai, Y . Yan, and L. Cheng. Gsd: View-guided gaussian splatting diffusion for 3d reconstruction. InEuropean Conference on Computer Vision, pp. 55–
-
[44]
Pavlakos, V
G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black. Expressive body capture: 3d hands, face, and body from a single image. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10975– 10985, 2019. 2
2019
-
[45]
S. Qian, T. Kirschstein, L. Schoneveld, D. Davoli, S. Giebenhain, and M. Nießner. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians.arXiv preprint arXiv:2312.02069, 2023. 1, 2
Pith/arXiv arXiv 2023
-
[46]
J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Mod- eling and capturing hands and bodies together.arXiv preprint arXiv:2201.02610, 2022. 2
Pith/arXiv arXiv 2022
-
[47]
Saito, J
S. Saito, J. Yang, Q. Ma, and M. J. Black. Scanimate: Weakly super- vised learning of skinned clothed avatar networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pp. 2886–2897, 2021. 2
2021
-
[48]
Sanyal, T
S. Sanyal, T. Bolkart, H. Feng, and M. J. Black. Learning to regress 3d face shape and expression from an image without 3d supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7763–7772, 2019. 1, 2
2019
-
[49]
Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y . Zhang, M. Fan, and Z. Wang. Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1606– 1616, 2024. 2
2024
-
[50]
Y . Shen, K. Zhou, H. Wang, Y . Yang, and T. Shao. High-fidelity 3d object generation from single image with rgbn-volume gaussian re- construction model.arXiv preprint arXiv:2504.01512, 2025. 9
Pith/arXiv arXiv 2025
-
[51]
Smailbegovic, G
F. Smailbegovic, G. N. Gaydadjiev, and S. Vassiliadis. Sparse ma- trix storage format. InProceedings of the 16th Annual Workshop on Circuits, Systems and Signal Processing, pp. 445–448, 2005. 6
2005
-
[52]
H. Song. Toward realistic 3d avatar generation with dynamic 3d gaus- sian splatting for ar/vr communication. In2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), pp. 869–870. IEEE, 2024. 2
2024
-
[53]
H. Song, B. Yoon, W. Cho, and W. Woo. Rc-smpl: Real-time cumu- lative smpl-based avatar body generation. In2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 89–98. IEEE, 2023. 2
2023
-
[54]
H. Song, B. Yoon, S. Yang, S. Kang, H. Kim, H. Metzmacher, and W. Woo. Vrgaussianavatar: Integrating 3d gaussian avatars into vr. arXiv preprint arXiv:2602.01674, 2026. 2
Pith/arXiv arXiv 2026
-
[55]
J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision, pp. 1–18. Springer, 2024. 3, 4
2024
-
[56]
J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation.arXiv preprint arXiv:2309.16653, 2023. 2
Pith/arXiv arXiv 2023
-
[57]
Tewari, F
A. Tewari, F. Bernard, P. Garrido, G. Bharaj, M. Elgharib, H.-P. Sei- del, P. P´erez, M. Zollhofer, and C. Theobalt. Fml: Face model learning from videos. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pp. 10812–10822, 2019. 1, 2
2019
-
[58]
Tewari, H.-P
A. Tewari, H.-P. Seidel, M. Elgharib, C. Theobalt, et al. Learning complete 3d morphable face models from images and videos. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 3361–3371, 2021. 1
2021
-
[59]
P. Tran, E. Zakharov, L.-N. Ho, A. T. Tran, L. Hu, and H. Li. V oodoo 3d: volumetric portrait disentanglement for one-shot 3d head reen- actment. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10336–10348, 2024. 2, 6
2024
-
[60]
P. Wang and Y . Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation.arXiv preprint arXiv:2312.02201, 2023. 4
Pith/arXiv arXiv 2023
-
[61]
S. Wang, M. Mihajlovic, Q. Ma, A. Geiger, and S. Tang. Metaavatar: Learning animatable clothed human models from few depth images. Advances in Neural Information Processing Systems, 34:2810–2822,
-
[62]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004. 6
2004
-
[63]
Xiang, X
J. Xiang, X. Gao, Y . Guo, and J. Zhang. Flashavatar: High-fidelity head avatar with efficient gaussian embedding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1802–1812, 2024. 2
2024
-
[64]
Y . Xu, B. Chen, Z. Li, H. Zhang, L. Wang, Z. Zheng, and Y . Liu. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 1931–1941, 2024. 1, 2
1931
-
[65]
Y . Xu, H. Zhang, L. Wang, X. Zhao, H. Huang, G. Qi, and Y . Liu. Latentavatar: Learning latent expression code for expressive neural head avatar. InACM SIGGRAPH 2023 Conference Proceedings, pp. 1–10, 2023. 1
2023
-
[66]
H. Yang, M. Zheng, C. Ma, Y .-K. Lai, P. Wan, and H. Huang. Vrmm: A volumetric relightable morphable head model. InACM SIGGRAPH 2024 Conference Papers, pp. 1–11, 2024. 2
2024
-
[67]
S. Yang, B. Yoon, S. Kang, H. Song, and W. Woo. Ofera: Blendshape- driven 3d gaussian control for occluded facial expression to realistic avatars in vr.arXiv preprint arXiv:2602.01748, 2026. 2, 8
arXiv 2026
-
[68]
T. Yi, J. Fang, J. Wang, G. Wu, L. Xie, X. Zhang, W. Liu, Q. Tian, and X. Wang. Gaussiandreamer: Fast generation from text to 3d gaus- sians by bridging 2d and 3d diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6796–6807, 2024. 2
2024
-
[69]
W. Yu, Y . Fan, Y . Zhang, X. Wang, F. Yin, Y . Bai, Y .-P. Cao, Y . Shan, Y . Wu, Z. Sun, et al. Nofa: Nerf-based one-shot facial avatar re- construction. InACM SIGGRAPH 2023 Conference Proceedings, pp. 1–12, 2023. 2
2023
-
[70]
Z. Yu, Z. Bai, A. Meka, F. Tan, Q. Xu, R. Pandey, S. Fanello, H. S. Park, and Y . Zhang. One2avatar: Generative implicit head avatar for few-shot user adaptation.arXiv preprint arXiv:2402.11909, 2024. 2
Pith/arXiv arXiv 2024
-
[71]
Zhang, Y
B. Zhang, Y . Cheng, J. Yang, C. Wang, F. Zhao, Y . Tang, D. Chen, and B. Guo. Gaussiancube: Structuring gaussian splatting using optimal transport for 3d generative modeling.arXiv e-prints, pp. arXiv–2403,
-
[72]
Zhang, Y
D. Zhang, Y . Liu, L. Lin, Y . Zhu, K. Chen, M. Qin, Y . Li, and H. Wang. Hravatar: High-quality and relightable gaussian head avatar. InPro- ceedings of the Computer Vision and Pattern Recognition Conference, pp. 26285–26296, 2025. 2
2025
-
[73]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018. 6
2018
-
[74]
Z. Zhao, Z. Bao, Q. Li, G. Qiu, and K. Liu. Psavatar: A point-based morphable shape model for real-time head avatar creation with 3d gaussian splatting.arXiv preprint arXiv:2401.12900, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[75]
X. Zheng, C. Wen, Z. Li, W. Zhang, Z. Su, X. Chang, Y . Zhao, Z. Lv, X. Zhang, Y . Zhang, et al. Headgap: Few-shot 3d head avatar via generalizable gaussian priors.arXiv preprint arXiv:2408.06019, 2024. 2
Pith/arXiv arXiv 2024
-
[76]
Zheng, V
Y . Zheng, V . F. Abrevaya, M. C. B ¨uhler, X. Chen, M. J. Black, and O. Hilliges. Im avatar: Implicit morphable head avatars from videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13545–13555, 2022. 1, 2
2022
-
[77]
Z. Zhou, F. Ma, H. Fan, and T.-S. Chua. Zero-1-to-a: Zero-shot one image to animatable head avatars using video diffusion. InProceed- ings of the Computer Vision and Pattern Recognition Conference, pp. 15941–15952, 2025. 2
2025
-
[78]
Zielonka, T
W. Zielonka, T. Bolkart, and J. Thies. Towards metrical reconstruction of human faces. InEuropean Conference on Computer Vision, pp. 250–269. Springer, 2022. 1, 2, 6 S-A vatar : Diffusion-Guided Gaussian Head A vatars from a Single Image Appendix. Figure 8: Reconstructed avatars rendered on Meta Quest 3. Our method supports photo-realistic and stereoscopi...
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.