Pith. sign in

REVIEW 3 major objections 5 minor 78 references

A single photo is enough to build a photorealistic 3D head avatar that stays consistent from the side and back and can be driven by expression parameters in real time.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 15:52 UTC pith:5AEVCPEH

load-bearing objection Solid systems paper: generate-first 3DGS + FLAME bind works, but quality is capped by frozen LGM and the gains are real but modest. the 3 major comments →

arxiv 2607.28164 v1 pith:5AEVCPEH submitted 2026-07-30 cs.CV cs.GR

S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

classification cs.CV cs.GR
keywords 3D Gaussian Splattinghead avatarsingle-image reconstructionFLAMEdiffusion modelsnovel-view synthesisVR/ARanimatable avatars
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

S-Avatar claims that one-shot head avatars no longer have to collapse under unseen viewpoints if you first let a diffusion model invent a full canonical 3D Gaussian head, then lock that head to the FLAME parametric model with a sparse binding template. Prior single-image methods either stay mesh-bound and lose fine detail or optimize neural fields from too few views and break on side and back angles. By separating geometry creation from animation control, the method produces avatars that render novel expressions and extreme camera angles while remaining real-time. On the NeRSemble benchmark it beats recent one-shot baselines on perceptual and pixel metrics, and the same pipeline works on casual smartphone photos and inside a VR headset. The practical stake is simple: personal VR/AR avatars become something anyone can make from one snapshot instead of a multi-view capture session.

Core claim

The paper establishes that a diffusion-guided feed-forward 3D Gaussian splat generator, followed by Chamfer-plus-landmark fitting of FLAME and an inverse-distance binding template (with optional density-aware scale adaptation), yields animatable head avatars from a single image whose novel-view and novel-expression renderings are more 3-D consistent and higher quality than prior one-shot mesh, GAN, and neural methods.

What carries the argument

The binding template B: a sparse matrix of normalized inverse-distance weights from each Gaussian splat to its nearest FLAME vertices. Multiplying B transpose by FLAME vertex displacements (and density ratios) deforms and rescales the canonical splats so expression and pose changes drive the avatar in real time.

Load-bearing premise

The whole pipeline assumes the diffusion model already produced a complete, multi-view-consistent set of initial Gaussian splats; if that first stage is wrong, later fitting and binding cannot recover a good avatar.

What would settle it

Take the same single input image, deliberately degrade or replace the initial splat stage (for example by removing the text prompts that suppress the Janus effect), run the identical fitting and binding steps, and check whether novel side/back views and expression transfers still match ground-truth NeRSemble frames at the reported LPIPS/PSNR levels.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Personal photorealistic head avatars become creatable from one ordinary photo rather than multi-view video or studio capture.
  • Real-time VR/AR head rendering from unobserved angles (side, back) is feasible without per-subject neural optimization at inference.
  • FLAME expression and pose parameters alone are sufficient to drive the avatar, so any FLAME tracker or HMD sensor pipeline can animate it.
  • The modular split lets stronger future single-image Gaussian generators be swapped in without redesigning the animation stage.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If initial-splat quality is the bottleneck, the same binding idea could be stress-tested on non-head objects once a reliable single-image Gaussian generator exists for those categories.
  • Density-aware splat scaling may transfer to other mesh-rigged Gaussian avatars (hands, bodies) where surface stretch under articulation is the main visual failure mode.
  • Because generation is feed-forward and binding is sparse, on-device or edge reconstruction of personal avatars becomes a realistic systems target once the diffusion backbone is distilled.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. S-Avatar proposes a three-stage pipeline for one-shot animatable 3D head avatars: (1) feed-forward generation of a canonical 3D Gaussian splat set S from a single image via ImageDream multi-view diffusion and an LGM-style asymmetric U-Net; (2) gradient-based fitting of FLAME (shape, expression, pose, and rigid transform/scale) to S using a weighted Chamfer + landmark objective (Eqs. 7–9); and (3) construction of a sparse inverse-distance binding template B (COO) that deforms splat positions by nearest-FLAME-vertex displacements and adapts splat scales via relative mesh density ratios (Eqs. 13–17). The avatar is then driven in real time by novel FLAME parameters. On 45 NeRSemble subjects across 16 views, the method reports better LPIPS/PSNR/SSIM than ROME, Portrait4D-v1/v2, VOODOO, and GPAvatar under a per-method face-crop protocol, with ablations on N_closest, scale adaptation, loss weights, and COO sparsity, plus qualitative extreme-view, cross-reenactment, in-the-wild, and VR demo results.

Significance. One-shot photorealistic, expression-controllable head avatars with real-time 3DGS rendering remain practically important for VR/AR. The decoupled design—build a diffusion-guided canonical splat field first, then attach FLAME control via a fixed geometric binding—is a clear and reusable systems contribution relative to methods that optimize radiance fields or Gaussians directly from sparse views. Strengths include a working real-time path (COO binding, ~35 FPS with scale adaptation), honest limitations, ablations that isolate N_closest and density-based scale updates, and demonstrated VR/interactive tooling. If the extreme-view consistency and identity-preserving animation claims hold under stricter full-head evaluation, the work would be a useful advance in accessible avatar creation. Credit is due for modularity (LGM can be swapped) and for releasing a project page with demos.

major comments (3)
  1. [§3.2, Limitations, Appendix Fig. 11] Limitations and §3.2 state that final quality is permanently capped by the frozen initial splat set S from LGM/ImageDream; fitting (§3.3) only aligns FLAME and binding (§3.4) only applies fixed IDW displacements plus density-ratio scales—neither can invent missing side/back geometry or undo Janus/inconsistency. The central claims of extreme-view 3D consistency and superior novel-view realism therefore inherit unquantified upstream error. The manuscript shows prompt-related failures in Appendix Fig. 11 but reports no failure rate, multi-view consistency metric, or oracle upper bound of S on the 45-subject NeRSemble split. Without that characterization (or a quantitative full-head side/back metric), the extreme-view narrative rests mainly on qualitative Figs. 5–6 and cannot be independently verified.
  2. [§4.1–4.2, Table 1] Table 1 gains are measured after per-method face cropping “so that the comparison is performed within the same facial region” (§4.1–4.2). This protocol can mask full-head failures (ears, scalp, back) that the paper’s extreme-view claim specifically advertises. Absolute PSNR remains modest (~16.1). Please add (i) a fixed full-head or head-silhouette evaluation protocol shared across methods, and/or (ii) region-specific metrics for side/back views, plus error bars or paired significance tests over the 45 subjects so that the SOTA ranking is statistically supported rather than point-estimate only.
  3. [§3.4, Eqs. 10–17, Table 2] The binding model (Eqs. 10–17) assumes that local inverse-distance weighting of N_closest FLAME vertices plus a density-ratio scale update adequately approximates non-rigid facial deformation of Gaussians. Table 2 only sweeps N_closest for aggregate PSNR; there is no stress test on large expression magnitudes, jaw/neck extremes, or topological failure cases (lip contact, cheek stretch). Given that animation quality is a primary claim, please quantify expression-range fidelity (e.g., error vs. held-out extreme expressions, or comparison against a learned residual deformer) and discuss when IDW binding breaks.
minor comments (5)
  1. [§3.1] Eq. (2) writes Σ = R S S^T R^T but the surrounding text uses s for scale; keep notation consistent with Kerbl et al. (scale vector vs. matrix).
  2. [Fig. 2] Several figure glyphs and FLAME parameter symbols in Fig. 2 are corrupted or use non-ASCII lookalikes (e.g., β, ψ, θ render inconsistently), which hurts readability in the PDF.
  3. [§4.1] Implementation details give Adam β1/β2 and 60 epochs but not wall-clock breakdown for the diffusion multi-view stage vs. fitting on the same hardware as the baselines; a single end-to-end timing table would help.
  4. [§2.3] Related work cites many concurrent one-shot Gaussian head methods; a short explicit comparison table (input type, canonical vs. direct optimize, real-time, full-head) would clarify positioning versus GPAvatar, SEGA, LAM, etc.
  5. [Abstract, Acknowledgments, References] Typos/style: “S-A vatar” spacing artifacts in the abstract block; “we was supported” in Acknowledgments; arXiv IDs in the 2602.* range for self-citations look premature—double-check bibliographic entries.

Circularity Check

0 steps flagged

No circularity: pipeline is modular engineering; metrics come from external NeRSemble GT and independent baselines, not from quantities defined by the fit.

full rationale

S-Avatar’s derivation is a three-stage engineering pipeline (LGM/ImageDream feed-forward initial splats S → Chamfer+landmark FLAME fit → fixed sparse IDW binding template B with density-based scale). None of the load-bearing equations reduce a claimed prediction to its own inputs. Binding weights w_ij are computed from Euclidean distances between fixed S and fitted FLAME vertices (Eqs. 10–17); they are not fitted to LPIPS/PSNR/SSIM. Novel-view/expression scores in Table 1 are measured against held-out multi-view NeRSemble ground truth and compared to external baselines (ROME, Portrait4D, VOODOO, GPAvatar). Ablations (N_closest, scale adaptation, λ_c/λ_l, COO) vary internal hyperparameters and report the same external metrics; they do not redefine the target. Self-citations (OFERA, lab VR demos) support only application demos in the appendix and do not justify the central quantitative claims or uniqueness of the method. The paper’s own Limitations explicitly treat initial splat quality as an upstream dependency rather than claiming to derive it from the evaluation. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The method rests on pretrained multi-view diffusion and LGM as black-box generators, on FLAME as the correct deformation space for heads, and on hand-chosen binding and loss hyperparameters. No new physical entity is postulated; the binding template is an engineered data structure whose justification is empirical quality/speed tradeoffs.

free parameters (5)
  • λ_c, λ_l (fitting loss weights) = 0.01 / 0.99
    Balance Chamfer vs landmark terms; chosen by ablation grid and fixed at 0.01/0.99 for reported results.
  • N_closest = 10
    Number of FLAME vertices influencing each splat; selected for PSNR/time tradeoff on their tests.
  • FLAME optimization learning rates and 60 epochs = lr shape 0.005; expr/pose 0.0005; transform 0.01; 60 epochs
    Adam schedules for shape/expression/pose/transform; hand-set implementation detail that affects fit quality.
  • Positive/negative text prompts for ImageDream/LGM = Fixed English positive/negative strings in appendix
    Prompt text materially changes initial splat quality (Janus/artifacts without prompts); engineered, not learned end-to-end here.
  • ε in inverse-distance weights = unspecified small constant
    Stabilizer to avoid division by zero in w_ij; small constant choice.
axioms (5)
  • domain assumption Pretrained ImageDream+LGM produce view-consistent initial 3DGS adequate for heads from one RGB crop.
    Invoked in §3.2 as the generation module; entire pipeline quality is gated on this (§5 Limitations).
  • domain assumption FLAME shape/pose/expression space plus rigid transform can align to and drive head geometry including regions FLAME does not model well (hair, ears detail).
    Fitting and binding (§3.3–3.4) assume FLAME vertices are the right control handles for splat motion.
  • ad hoc to paper Local inverse-distance weighting of nearest FLAME vertices plus density-ratio scale updates approximates non-rigid facial deformation of Gaussians.
    Binding template construction and Eqs. 13–17; justified by ablation, not derived from continuum mechanics.
  • ad hoc to paper Per-method face cropping yields a fair comparison across heterogeneous avatar outputs.
    Stated in §4.1 Datasets; affects all reported LPIPS/PSNR/SSIM.
  • standard math Standard 3DGS alpha-blending and FLAME LBS definitions hold as in Kerbl et al. and Li et al.
    Preliminaries §3.1; background, not re-proved.
invented entities (2)
  • Binding template B (sparse NFLAME×Nsplat weight matrix in COO form) no independent evidence
    purpose: Map FLAME vertex displacements and density changes onto Gaussian positions and scales for real-time expression control.
    Defined in §3.4; central engineered object of the paper. Independent evidence is only the empirical render metrics and speed, not an external physical measurement.
  • Relative-density Gaussian scale adaptation (σ from P and P') no independent evidence
    purpose: Grow/shrink splats when local mesh stretches/compresses to preserve apparent surface density under expression.
    Eqs. 15–17; ablation Table 3 shows PSNR gain. No external validation beyond this paper's renders.

pith-pipeline@v1.2.0-daily-grok45 · 24672 in / 3460 out tokens · 72557 ms · 2026-07-31T15:52:01.064852+00:00 · methodology

0 comments
read the original abstract

We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation module and strategies for animating 3D Gaussian Splatting (3DGS). While single-image head avatar reconstruction is crucial for lifelike Virtual Reality (VR) applications, existing approaches often struggle to preserve 3D consistency under unseen viewpoints. S-Avatar addresses this limitation through a three-stage pipeline. First, a high-resolution 3DGS is synthesized directly from a single image using a diffusion-based Gaussian splat generation module. Next, the parametric head model FLAME is aligned with the generated 3DGS by optimizing its parameters and spatial transformations. Finally, to adapt the 3DGS to FLAME variations, we construct a binding template that encodes the spatial relationship between the initial splats and FLAME. The dynamic 3D head avatar can then be rendered in real time by deforming the 3DGS with the binding template. By combining diffusion-guided canonical 3DGS generation with FLAME-based control, our method achieves efficient and accurate reconstruction with enhanced 3D consistency. Evaluations on public datasets demonstrate that S-Avatar outperforms state-of-the-art methods in novel-view and expression generation, achieving superior realism and consistency. Consequently, our approach represents a significant advance in accessible avatar creation, applicable to a wide range of VR/AR applications. The project page is available at https://github.com/hailsong/savatar.

Figures

Figures reproduced from arXiv: 2607.28164 by Hail Song, Jiwon Yang, Seokhwan Yang, Woojin Cho, Woontack Woo.

Figure 1
Figure 1. Figure 1: We introduce S-Avatar, a novel method for generating 3D human head avatars from a single image input. Utilizing a diffusion-based Gaussian splat generation module and a fitting/binding strategy for dynamic 3D Gaussian Splatting with expression control, our approach enables one-shot head avatar generation. This method produces photorealistic avatars that can be rendered from novel viewpoints and controlled … view at source ↗
Figure 2
Figure 2. Figure 2: System diagram of the proposed method, S-Avatar. Our method reconstructs a photorealistic head avatar from a single image input through a three-step modeling process: (1) Generating initial splats using Gaussian splats generation module, (2) Fitting the FLAME model to the generated Gaussian splats, and (3) Creating a binding template. During the rendering phase, the generated avatar can display various fac… view at source ↗
Figure 3
Figure 3. Figure 3: The fitting losses of S-Avatar. (A) shows the chamfer dis [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The binding process of S-Avatar. (A) shows the process of [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Multi-view self-reenactment rendering results for S-Avatar and baseline methods. Both our method and the baseline generate avatars [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative comparison of cross-reenactment results. Each avatar is reconstructed from a source image and driven by the expression [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Head avatar generation results from a one-shot, in-the [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: Interactive viewer for visualizing the output of S-Avatar. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: VR demo results using OFERA. The proposed method [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Rendering results of the initial splats S based on whether the text prompt input is provided for Gaussian generation module. It can be observed that providing a text prompt results in improved generation outcomes. Impact of Prompt on initial splats Generation. To assess the impact of the optional text prompt input for Gaussian generation module, we observed the generation results of the initial splats wit… view at source ↗
Figure 12
Figure 12. Figure 12: Additional rendering results from our method, S-Avatar [PITH_FULL_IMAGE:figures/full_fig_p015_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 1 canonical work pages

  1. [1]

    Aliev, A

    K.-A. Aliev, A. Sevastopolsky, M. Kolos, D. Ulyanov, and V . Lempit- sky. Neural point-based graphics. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Pro- ceedings, Part XXII 16, pp. 696–712. Springer, 2020. 2

  2. [2]

    Alldieck, M

    T. Alldieck, M. Magnor, W. Xu, C. Theobalt, and G. Pons-Moll. Video based reconstruction of 3d people models. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8387– 8397, 2018. 2

  3. [3]

    Athar, Z

    S. Athar, Z. Shu, and D. Samaras. Flame-in-nerf: Neural control of radiance fields for free view face animation. In2023 IEEE 17th In- ternational Conference on Automatic Face and Gesture Recognition (FG), pp. 1–8. IEEE, 2023. 2

  4. [4]

    Blanz and T

    V . Blanz and T. Vetter. A morphable model for the synthesis of 3d faces. InSeminal Graphics Papers: Pushing the Boundaries, V olume 2, pp. 157–164. 2023. 1, 2

  5. [5]

    M. C. B ¨uhler, K. Sarkar, T. Shah, G. Li, D. Wang, L. Helminger, S. Orts-Escolano, D. Lagun, O. Hilliges, T. Beeler, et al. Preface: A data-driven volumetric prior for few-shot ultra high-resolution face synthesis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3402–3413, 2023. 2

  6. [6]

    C. Cao, T. Simon, J. K. Kim, G. Schwartz, M. Zollhoefer, S.-S. Saito, S. Lombardi, S.-E. Wei, D. Belko, S.-I. Yu, et al. Authentic volumetric avatars from a phone scan.ACM Transactions on Graphics (TOG), 41(4):1–19, 2022. 2

  7. [7]

    Chen and W

    G. Chen and W. Wang. A survey on 3d gaussian splatting.arXiv preprint arXiv:2401.03890, 2024. 3

  8. [8]

    X. Chen, M. Mihajlovic, S. Wang, S. Prokudin, and S. Tang. Mor- phable diffusion: 3d-consistent diffusion for single-image avatar cre- ation. InProceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pp. 10359–10370, 2024. 2

  9. [9]

    Y . Chen, L. Wang, Q. Li, H. Xiao, S. Zhang, H. Yao, and Y . Liu. Monogaussianavatar: Monocular gaussian point-based head avatar. arXiv preprint arXiv:2312.04558, 2023. 1, 2

  10. [10]

    Z. Chen, F. Wang, Y . Wang, and H. Liu. Text-to-3d using gaussian splatting. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 21401–21412, 2024. 2

  11. [11]

    X. Chu, Y . Li, A. Zeng, T. Yang, L. Lin, Y . Liu, and T. Harada. Gpa- vatar: Generalizable and precise head avatar from image (s).arXiv preprint arXiv:2401.10215, 2024. 2, 6

  12. [12]

    Dan ˇeˇcek, M

    R. Dan ˇeˇcek, M. J. Black, and T. Bolkart. Emoca: Emotion driven monocular face capture and animation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20311–20322, 2022. 1, 2

  13. [13]

    Y . Deng, D. Wang, X. Ren, X. Chen, and B. Wang. Portrait4d: Learn- ing one-shot 4d head avatar synthesis using synthetic data. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 7119–7130, 2024. 2, 6

  14. [14]

    Y . Deng, D. Wang, and B. Wang. Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer. InEuropean Conference on Computer Vision, pp. 316–333. Springer, 2024. 2, 6

  15. [15]

    Y . Deng, J. Yang, S. Xu, D. Chen, Y . Jia, and X. Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pp. 0–0, 2019. 1, 2

  16. [16]

    H. Fan, H. Su, and L. J. Guibas. A point set generation network for 3d object reconstruction from a single image. InProceedings of the IEEE conference on computer vision and pattern recognition, pp. 605–613,

  17. [17]

    Y . Feng, H. Feng, M. J. Black, and T. Bolkart. Learning an animatable detailed 3d face model from in-the-wild images.ACM Transactions on Graphics (ToG), 40(4):1–13, 2021. 1, 2

  18. [18]

    Gafni, J

    G. Gafni, J. Thies, M. Zollhofer, and M. Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 8649–8658, 2021. 2

  19. [19]

    Garrido, M

    P. Garrido, M. Zollh ¨ofer, D. Casas, L. Valgaerts, K. Varanasi, P. P´erez, and C. Theobalt. Reconstruction of personalized 3d face rigs from monocular video.ACM Transactions on Graphics (TOG), 35(3):1– 15, 2016. 1, 2

  20. [20]

    Giebenhain, T

    S. Giebenhain, T. Kirschstein, M. R ¨unz, L. Agapito, and M. Nießner. Npga: Neural parametric gaussian avatars. InSIGGRAPH Asia 2024 Conference Papers, pp. 1–11, 2024. 2

  21. [21]

    C. Guo, Z. Su, J. Wang, S. Li, X. Chang, Z. Li, Y . Zhao, G. Wang, and R. Huang. Sega: Drivable 3d gaussian head avatar from a single image.arXiv preprint arXiv:2504.14373, 2025. 2

  22. [22]

    X. He, J. Chen, S. Peng, D. Huang, Y . Li, X. Huang, C. Yuan, W. Ouyang, and T. He. Gvgen: Text-to-3d generation with volumet- ric representation. InEuropean Conference on Computer Vision, pp. 463–479. Springer, 2024. 3

  23. [23]

    Y . He, X. Gu, X. Ye, C. Xu, Z. Zhao, Y . Dong, W. Yuan, Z. Dong, and L. Bo. Lam: Large avatar model for one-shot animatable gaussian head.arXiv preprint arXiv:2502.17796, 2025. 2

  24. [24]

    C. A. R. Hoare. Algorithm 64: quicksort.Communications of the ACM, 4(7):321, 1961. 5

  25. [25]

    M. R. Hugues and S. G. Petiton. Sparse matrix formats evaluation and optimization on a gpu. In2010 IEEE 12th International Conference on High Performance Computing and Communications (HPCC), pp. 122–129. IEEE, 2010. 6

  26. [26]

    Huynh-Thu and M

    Q. Huynh-Thu and M. Ghanbari. Scope of validity of psnr in im- age/video quality assessment.Electronics letters, 44(13):800–801,

  27. [27]

    Z. Ke, J. Sun, K. Li, Q. Yan, and R. W. Lau. Modnet: Real-time trimap-free portrait matting via objective decomposition. InAAAI,

  28. [28]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4):1–14, 2023. 1, 2, 3

  29. [29]

    Khakhulin, V

    T. Khakhulin, V . Sklyarova, V . Lempitsky, and E. Zakharov. Real- istic one-shot mesh-based head avatars. InEuropean Conference on Computer Vision, pp. 345–362. Springer, 2022. 2, 6

  30. [30]

    D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 4

  31. [31]

    Kirschstein, S

    T. Kirschstein, S. Qian, S. Giebenhain, T. Walter, and M. Nießner. Nersemble: Multi-view radiance field reconstruction of human heads. ACM Trans. Graph., 42(4), jul 2023. doi:10.1145/35924556, 7

  32. [32]

    Lee, Y .-T

    Y .-C. Lee, Y .-T. Chen, A. Wang, T.-H. Liao, B. Y . Feng, and J.-B. Huang. Vividdream: Generating 3d scene with ambient dynamics. arXiv preprint arXiv:2405.20334, 2024. 3

  33. [33]

    J. Li, C. Cao, G. Schwartz, R. Khirodkar, C. Richardt, T. Simon, Y . Sheikh, and S. Saito. Uravatar: Universal relightable gaussian codec avatars. InSIGGRAPH Asia 2024 Conference Papers, pp. 1–11,

  34. [34]

    T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero. Learning a model of facial shape and expression from 4d scans.ACM Trans. Graph., 36(6):194–1, 2017. 1, 2

  35. [35]

    W. Li, L. Zhang, D. Wang, B. Zhao, Z. Wang, M. Chen, B. Zhang, Z. Wang, L. Bo, and X. Li. One-shot high-fidelity talking-head syn- thesis with deformable neural radiance field. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17969–17978, 2023. 2

  36. [36]

    X. Li, S. De Mello, S. Liu, K. Nagano, U. Iqbal, and J. Kautz. General- izable one-shot 3d neural head avatar.Advances in Neural Information Processing Systems, 36:47239–47250, 2023. 2

  37. [37]

    Loper, N

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. Smpl: A skinned multi-person linear model.ACM transactions on graphics (TOG), 34(6):1–16, 2015. 2

  38. [38]

    W. Lyu, Y . Zhou, M.-H. Yang, and Z. Shu. Facelift: Single image to 3d head with view generation and gs-lrm, 2024. 9

  39. [39]

    S. Ma, T. Simon, J. Saragih, D. Wang, Y . Li, F. De La Torre, and Y . Sheikh. Pixel codec avatars. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 64–73,

  40. [40]

    Melas-Kyriazi, I

    L. Melas-Kyriazi, I. Laina, C. Rupprecht, N. Neverova, A. Vedaldi, O. Gafni, and F. Kokkinos. Im-3d: Iterative multiview diffusion and reconstruction for high-quality 3d generation.arXiv preprint arXiv:2402.08682, 2024. 3

  41. [41]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis.Communications of the ACM, 65(1):99–106, 2021. 1, 2

  42. [42]

    S. Moon, H. M. Lew, S. Lee, J.-S. Kang, and G.-M. Park. Geoavatar: Adaptive geometrical gaussian splatting for 3d head avatar.arXiv preprint arXiv:2507.18155, 2025. 2

  43. [43]

    Y . Mu, X. Zuo, C. Guo, Y . Wang, J. Lu, X. Wu, S. Xu, P. Dai, Y . Yan, and L. Cheng. Gsd: View-guided gaussian splatting diffusion for 3d reconstruction. InEuropean Conference on Computer Vision, pp. 55–

  44. [44]

    Pavlakos, V

    G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black. Expressive body capture: 3d hands, face, and body from a single image. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10975– 10985, 2019. 2

  45. [45]

    S. Qian, T. Kirschstein, L. Schoneveld, D. Davoli, S. Giebenhain, and M. Nießner. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians.arXiv preprint arXiv:2312.02069, 2023. 1, 2

  46. [46]

    Romero, D

    J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Mod- eling and capturing hands and bodies together.arXiv preprint arXiv:2201.02610, 2022. 2

  47. [47]

    Saito, J

    S. Saito, J. Yang, Q. Ma, and M. J. Black. Scanimate: Weakly super- vised learning of skinned clothed avatar networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pp. 2886–2897, 2021. 2

  48. [48]

    Sanyal, T

    S. Sanyal, T. Bolkart, H. Feng, and M. J. Black. Learning to regress 3d face shape and expression from an image without 3d supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7763–7772, 2019. 1, 2

  49. [49]

    Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y . Zhang, M. Fan, and Z. Wang. Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1606– 1616, 2024. 2

  50. [50]

    Y . Shen, K. Zhou, H. Wang, Y . Yang, and T. Shao. High-fidelity 3d object generation from single image with rgbn-volume gaussian re- construction model.arXiv preprint arXiv:2504.01512, 2025. 9

  51. [51]

    Smailbegovic, G

    F. Smailbegovic, G. N. Gaydadjiev, and S. Vassiliadis. Sparse ma- trix storage format. InProceedings of the 16th Annual Workshop on Circuits, Systems and Signal Processing, pp. 445–448, 2005. 6

  52. [52]

    H. Song. Toward realistic 3d avatar generation with dynamic 3d gaus- sian splatting for ar/vr communication. In2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), pp. 869–870. IEEE, 2024. 2

  53. [53]

    H. Song, B. Yoon, W. Cho, and W. Woo. Rc-smpl: Real-time cumu- lative smpl-based avatar body generation. In2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 89–98. IEEE, 2023. 2

  54. [54]

    H. Song, B. Yoon, S. Yang, S. Kang, H. Kim, H. Metzmacher, and W. Woo. Vrgaussianavatar: Integrating 3d gaussian avatars into vr. arXiv preprint arXiv:2602.01674, 2026. 2

  55. [55]

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision, pp. 1–18. Springer, 2024. 3, 4

  56. [56]

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation.arXiv preprint arXiv:2309.16653, 2023. 2

  57. [57]

    Tewari, F

    A. Tewari, F. Bernard, P. Garrido, G. Bharaj, M. Elgharib, H.-P. Sei- del, P. P´erez, M. Zollhofer, and C. Theobalt. Fml: Face model learning from videos. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pp. 10812–10822, 2019. 1, 2

  58. [58]

    Tewari, H.-P

    A. Tewari, H.-P. Seidel, M. Elgharib, C. Theobalt, et al. Learning complete 3d morphable face models from images and videos. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp. 3361–3371, 2021. 1

  59. [59]

    P. Tran, E. Zakharov, L.-N. Ho, A. T. Tran, L. Hu, and H. Li. V oodoo 3d: volumetric portrait disentanglement for one-shot 3d head reen- actment. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10336–10348, 2024. 2, 6

  60. [60]

    Wang and Y

    P. Wang and Y . Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation.arXiv preprint arXiv:2312.02201, 2023. 4

  61. [61]

    S. Wang, M. Mihajlovic, Q. Ma, A. Geiger, and S. Tang. Metaavatar: Learning animatable clothed human models from few depth images. Advances in Neural Information Processing Systems, 34:2810–2822,

  62. [62]

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004. 6

  63. [63]

    Xiang, X

    J. Xiang, X. Gao, Y . Guo, and J. Zhang. Flashavatar: High-fidelity head avatar with efficient gaussian embedding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1802–1812, 2024. 2

  64. [64]

    Y . Xu, B. Chen, Z. Li, H. Zhang, L. Wang, Z. Zheng, and Y . Liu. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 1931–1941, 2024. 1, 2

  65. [65]

    Y . Xu, H. Zhang, L. Wang, X. Zhao, H. Huang, G. Qi, and Y . Liu. Latentavatar: Learning latent expression code for expressive neural head avatar. InACM SIGGRAPH 2023 Conference Proceedings, pp. 1–10, 2023. 1

  66. [66]

    H. Yang, M. Zheng, C. Ma, Y .-K. Lai, P. Wan, and H. Huang. Vrmm: A volumetric relightable morphable head model. InACM SIGGRAPH 2024 Conference Papers, pp. 1–11, 2024. 2

  67. [67]

    S. Yang, B. Yoon, S. Kang, H. Song, and W. Woo. Ofera: Blendshape- driven 3d gaussian control for occluded facial expression to realistic avatars in vr.arXiv preprint arXiv:2602.01748, 2026. 2, 8

  68. [68]

    T. Yi, J. Fang, J. Wang, G. Wu, L. Xie, X. Zhang, W. Liu, Q. Tian, and X. Wang. Gaussiandreamer: Fast generation from text to 3d gaus- sians by bridging 2d and 3d diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6796–6807, 2024. 2

  69. [69]

    W. Yu, Y . Fan, Y . Zhang, X. Wang, F. Yin, Y . Bai, Y .-P. Cao, Y . Shan, Y . Wu, Z. Sun, et al. Nofa: Nerf-based one-shot facial avatar re- construction. InACM SIGGRAPH 2023 Conference Proceedings, pp. 1–12, 2023. 2

  70. [70]

    Z. Yu, Z. Bai, A. Meka, F. Tan, Q. Xu, R. Pandey, S. Fanello, H. S. Park, and Y . Zhang. One2avatar: Generative implicit head avatar for few-shot user adaptation.arXiv preprint arXiv:2402.11909, 2024. 2

  71. [71]

    Zhang, Y

    B. Zhang, Y . Cheng, J. Yang, C. Wang, F. Zhao, Y . Tang, D. Chen, and B. Guo. Gaussiancube: Structuring gaussian splatting using optimal transport for 3d generative modeling.arXiv e-prints, pp. arXiv–2403,

  72. [72]

    Zhang, Y

    D. Zhang, Y . Liu, L. Lin, Y . Zhu, K. Chen, M. Qin, Y . Li, and H. Wang. Hravatar: High-quality and relightable gaussian head avatar. InPro- ceedings of the Computer Vision and Pattern Recognition Conference, pp. 26285–26296, 2025. 2

  73. [73]

    Zhang, P

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018. 6

  74. [74]

    Z. Zhao, Z. Bao, Q. Li, G. Qiu, and K. Liu. Psavatar: A point-based morphable shape model for real-time head avatar creation with 3d gaussian splatting.arXiv preprint arXiv:2401.12900, 2024. 1, 2

  75. [75]

    Zheng, C

    X. Zheng, C. Wen, Z. Li, W. Zhang, Z. Su, X. Chang, Y . Zhao, Z. Lv, X. Zhang, Y . Zhang, et al. Headgap: Few-shot 3d head avatar via generalizable gaussian priors.arXiv preprint arXiv:2408.06019, 2024. 2

  76. [76]

    Zheng, V

    Y . Zheng, V . F. Abrevaya, M. C. B ¨uhler, X. Chen, M. J. Black, and O. Hilliges. Im avatar: Implicit morphable head avatars from videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13545–13555, 2022. 1, 2

  77. [77]

    Z. Zhou, F. Ma, H. Fan, and T.-S. Chua. Zero-1-to-a: Zero-shot one image to animatable head avatars using video diffusion. InProceed- ings of the Computer Vision and Pattern Recognition Conference, pp. 15941–15952, 2025. 2

  78. [78]

    Zielonka, T

    W. Zielonka, T. Bolkart, and J. Thies. Towards metrical reconstruction of human faces. InEuropean Conference on Computer Vision, pp. 250–269. Springer, 2022. 1, 2, 6 S-A vatar : Diffusion-Guided Gaussian Head A vatars from a Single Image Appendix. Figure 8: Reconstructed avatars rendered on Meta Quest 3. Our method supports photo-realistic and stereoscopi...