Pith. sign in

REVIEW 4 major objections 5 minor 116 references

CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read CAP4D proposes a single morphable multi-view diffusion pipeline that reconstructs photoreal, animatable 4D portrait avatars from anywhere from one to one hundred reference images and renders them in real time.

desk verdict Strong engineering and a genuinely useful stochastic conditioning idea, but the headline self-reenactment claims currently hinge on an unstated train/test split that a referee must resolve. read the letter →

arxiv 2412.12093 v1 pith:HRCYUZ5G submitted 2024-12-16 cs.CV

classification cs.CV
keywords 4Davatarreconstructionmulti-viewdiffusionmodel3DGaussiansplattingmorphableportraitreenactmentnovelviewsynthesisstochasticconditioningreal-timerendering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that one pipeline can turn any number of reference photos of a face—one, ten, or a hundred—into a single animatable 3D avatar that moves and renders in real time. The key idea is to condition a diffusion model not only on the input photos but on the estimated 3D shape, expression, and camera direction of every photo, so it can generate hundreds of new views with new expressions that stay consistent with the subject's identity. Those generated views are then distilled into a 3D Gaussian avatar whose expression-driven deformations are learned by a small network, giving a representation that can be driven by another person's video or by speech. If the central claim holds, content creators and visual effects pipelines could use the same method whether they have a single selfie or a studio capture rig.

What carries the argument

The morphable multi-view diffusion model (MMDM) is the central object: a latent diffusion model initialized from Stable Diffusion 2.1, with 3D attention across images and cross-attention removed, conditioned on six channels of per-image geometry: 3D pose maps, expression deformation maps, view-direction maps, and masks. It is what turns an arbitrary set of reference images into a large, expression-diverse set of self-consistent views. The second load-bearing mechanism is stochastic I/O conditioning: at each DDIM timestep, reference and generated latents are randomly shuffled and processed in batches, so all images participate jointly in denoising even though the network can only see a few at a time. Finally, the 4D avatar represents the subject as 3D Gaussians attached to a FLAME head mesh with a U-Net predicting expression-dependent UV deformations, which makes the distilled avatar animatable and real-time renderable.

What would settle it

Check the nine Nersemble evaluation sequence IDs against the training corpus and rerun the self-reenactment comparison with the model retrained on a strict subject-level holdout; if the single- and few-image advantages over the baselines disappear, the claimed generalization is not real.

Watch

Extended reading notes

Core claim

CAP4D's central discovery is that the gap between single-image and multi-view avatar reconstruction can be bridged by treating novel-view generation as a multi-view diffusion problem with 3D morphable-model conditioning, then distilling the generated images into an animatable Gaussian representation. The diffusion model takes up to four reference images at a time and generates images for target viewpoints, poses, and expressions specified by a FLAME model; a stochastic input-output conditioning loop shuffles and resamples reference and generated images at every diffusion timestep, so the same joint denoising process can absorb one or many reference images and emit hundreds of consistent novel views. These views are then used to optimize a 3D Gaussian splat avatar attached to a remeshed FLAME head, with expression-dependent corrective deformations predicted by a U-Net. The result is claimed to outperform prior methods on self- and cross-reenactment, to sharpen as more reference images are added, and to remain animatable and renderable in real time.

Load-bearing premise

The load-bearing premise is that the nine Nersemble subjects used for testing were not part of the 6,317 subjects the diffusion model trained on; the paper evaluates on Nersemble after training on Nersemble and reports only a camera-view holdout, not a subject-level split, so its single- and few-image gains could in part reflect memorized identities.

Editorial extensions

If this is right

  • A single captured image becomes a fully animatable 4D avatar, closing much of the fidelity gap with multi-view studio methods.
  • Adding more reference images improves the reconstructed avatar, and the same pipeline scales from one to hundreds without reconfiguration.
  • The avatar can be driven by another person's video or by speech, and the rendered output remains temporally consistent.
  • Because the MMDM conditions on 3D morphable-model parameters, avatars can be edited in 2D, such as with makeup or relighting, and then reanimated as 4D.
  • Single- and few-image self-reenactment is claimed to outperform prior single-view and multi-view baselines on photometric fidelity, identity preservation, and temporal consistency.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The stochastic I/O conditioning loop is not tied to heads: any multi-view diffusion model with a small context window could use the same shuffle-and-resample trick to generate hundreds of mutually consistent views for static scenes, objects, or full bodies.
  • Because the final avatar is only as expressive as the FLAME parameters that drive it, the method's animation range for hair, glasses, tongue, and jaw details is a likely ceiling; moving beyond mesh-bound Gaussians would be the natural next step.
  • The observation that the diffusion model plateaus while the final avatar keeps improving with hundreds of references suggests the generate-and-reconstruct decomposition is the part that scales, so future work should focus on cheaper generation rather than bigger diffusion models.
  • A decisive experiment not reported in the paper is a strict subject-level train/test split on the Nersemble data; the reported single- and few-image gains would be much more convincing if they survive that split.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes CAP4D, a two-stage pipeline for reconstructing animatable 4D portrait avatars from one to hundreds of reference images. In the first stage, a morphable multi-view diffusion model (MMDM), initialized from Stable Diffusion 2.1 and conditioned on FLAME-based pose, expression, view-direction, and mask maps, generates many novel views with controlled expressions using a stochastic input/output conditioning procedure that alternates reference and generated image subsets across diffusion timesteps. In the second stage, the generated and reference images are used to fit a real-time 4D avatar based on GaussianAvatars, augmented with a UV-space deformation U-Net and an LPIPS loss. The method is evaluated on self-reenactment (Nersemble, with 1/10/100 reference images) and cross-reenactment (FFHQ references driven by VFHQ/Nersemble videos), reporting improved PSNR, LPIPS, CSIM, JOD, and human-preference results over several baselines, along with ablations of the main components.

Significance. If the empirical claims hold, CAP4D would be a valuable unified solution bridging single-image and multi-view avatar reconstruction, with a practical real-time rendering avatar and a generation stage that scales from one to hundreds of references. The stochastic I/O conditioning idea is interesting and the ablations are reasonably broad, covering conditioning signals, sampling strategy, reconstruction losses, and the number of generated views. The paper also presents failure cases and a brief ethics statement, which is commendable. However, the central generalization claim depends on the self-reenactment evaluation, and that evaluation has a potential training/test overlap with the MMDM's training corpus, as well as internal table inconsistencies that currently prevent a confident assessment of the method's actual performance.

major comments (4)
  1. [§4, §5.1, Supp. D] The paper never states that the nine Nersemble sequences used for self-reenactment evaluation in §5.1 were excluded from the MMDM training corpus described in §4. Supp. D says that during training, 'we randomly select R reference images and G target images from all views and frames within a sequence with equal probability,' and the training pool explicitly includes Nersemble. If the evaluation sequences were among those training sequences, the single- and few-image results in Table 1 would reflect subject memorization rather than generalization, and the comparison would be unfair because the single-view baselines (Voodoo3D, GAGAvatar, Real3D, Portrait4D-v2) are not trained on Nersemble subjects. Please state explicitly whether a subject-level train/test split was enforced; if it was not, the evaluation needs to be rerun on held-out subjects before the central claim can be assessed.
  2. [§5.3, Tables 1 and 3] The ablation study in §5.3 claims that all ablations in Table 3 use 10 reference images, but the '4D rep.' rows report PSNR 21.69, LPIPS 0.311, CSIM 0.633, JOD 5.67, which exactly match the single-reference CAP4D row in Table 1 (21.69, 0.311, 0.633, 5.672). In addition, the MMDM-only row in Table 1 at 10 references reports CSIM 0.804, while the corresponding 'sampling/Ours' row in Table 3 reports 0.779. These inconsistencies make it impossible to know which setting the 4D-representation ablations actually used; please correct the tables or state the setting used for each ablation.
  3. [Algorithm 1] Algorithm 1 is internally inconsistent as printed. The first instruction in the t-loop shuffles generated latents into (Z'gen,t, C'gen), but the batch selection immediately overwrites these variables using the original, unshuffled (Zgen,t, Cgen)[iG'+1 : (i+1)G']. Moreover, the DDIM update inside the loop writes to Z'gen,t-1 for the current batch only, and there is no statement that reassembles Zgen,t-1 before the next outer iteration. Since the stochastic I/O conditioning is a central contribution and no code is provided, this pseudocode needs to be corrected so that the sampling procedure is reproducible.
  4. [§5.1, Table 1] Table 1 reports only point estimates over the nine Nersemble sequences, with no error bars, per-sequence results, or significance tests. Some reported margins are modest (e.g., single-reference PSNR 21.69 vs. GAGAvatar's 20.78; 100-reference PSNR 23.30 vs. FlashAvatar's 22.87), so the statement that CAP4D 'significantly outperforms every baseline' is not yet supported. Please provide variance estimates or per-subject results, particularly for the single- and few-image settings where the train/test overlap concern is most acute.
minor comments (5)
  1. [Table 2 caption, §5.2, Supp. E.3] Table 2 caption states the user study had 23 participants, while §5.2 and Supp. E.3 state 24; please correct the inconsistency.
  2. [Table 1 caption] The caption of Table 1 describes 'single-image (left) and multi-image (right)' results, but the table has three blocks (single, 10, 100 reference images); please update the caption.
  3. [§6 vs §4] The Discussion (§6) says generation takes 'up to 8 hours,' while §4 reports 840-image generation taking ~4 hours on 4xRTX6000 GPUs; please clarify whether the 8-hour figure includes 4D avatar reconstruction.
  4. [Fig. S9 caption] The caption of Fig. S9 contains the typo 'Quanitity' for 'Quantity'.
  5. [Supp. D] The expression database used for sampling is built from the Nersemble dataset; please clarify that only expression parameters are used, not identity or appearance information, to avoid any appearance leakage through the sampling step.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the generate-and-distill pipeline is self-contained; the only self-citation (face tracker) is a tool dependency, and the missing Nersemble subject-level split is a correctness risk, not a circular step.

full rationale

CAP4D's claimed derivation is a generate-and-distill pipeline: a multi-view diffusion model is trained on a collection of datasets, then at inference it synthesizes novel views from reference images, and these views are used to fit a Gaussian-splatting avatar. None of the evaluation targets (held-out Nersemble viewpoints, FFHQ identities) is used as a fitting input in the same step that reports the metric. The only self-citation of the authors' own work is the off-the-shelf 3D face tracker [83], used to produce FLAME parameters for conditioning; this is a tool dependency, not a load-bearing derivation, because the tracker's output does not encode the target result (novel-view image quality or identity preservation), and the avatar metrics are computed against ground-truth images for which separate FLAME fits are made. There is no equation in which a fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work. The main methodological risk—that the nine Nersemble self-reenactment subjects may not have been excluded from MMDM training, since Section 4 lists Nersemble in the training pool and Section 5.1 holds out only 4 of 16 camera viewpoints—is a data-leakage/correctness concern, not a circularity of the derivation chain; the paper's own text does not exhibit the equation X = Y pattern required for a circularity finding. Accordingly, the circularity score is 1 (minor self-citation that is not load-bearing).

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim loads on the learned diffusion prior and the evaluation protocol rather than on a mathematical derivation. The free parameters are standard training and architecture hyperparameters chosen by hand; none are fitted to the evaluation benchmarks. The main unverified assumptions are the coverage of the training prior and the fairness of the train/evaluation split.

free parameters (6)
  • lambda_LPIPS schedule = 0 to 0.9
    Weight of LPIPS loss in the avatar fitting loss (Eq. 6, Supplementary C); chosen by hand and annealed linearly.
  • lambda_deform = 0.4
    Weight for deformation map regularization loss (Eq. 7, Supplementary C).
  • lambda_rot = 0.005
    Weight for Gaussian rotation regularization loss (Eq. 7, Supplementary C).
  • G (generated views) = 840
    Number of novel views generated for avatar fitting; ablations show G=840 improves over 420 but 1260 adds little (Supplementary E.5).
  • CFG guidance weight = 2
    Classifier-free guidance weight used during MMDM sampling (Supplementary A.1).
  • View sampling bounds (psi_max, theta_max) = 55 degrees, 20 degrees
    Limits on the azimuth/elevation ellipse for sampled novel views (Eq. 4, Supplementary A.3). Evaluation views reach up to about 49 degrees azimuth, inside this bound.
assumptions (5)
  • domain assumption FLAME 3DMM provides a sufficient parametric description of head shape, expression, and pose for both conditioning and avatar animation.
    The MMDM and the avatar are conditioned on FLAME geometry. If FLAME cannot represent a subject (or geometry like hair and glasses), generation and fitting degrade, as shown in the failure cases (Fig. S10).
  • domain assumption The head tracker [83] yields accurate 3DMM parameters, camera intrinsics, and extrinsics for training and inference.
    All conditioning maps and camera poses come from this off-the-shelf tracker (Section 3.1, Supplementary A.2). Tracking errors directly corrupt the conditioning signals.
  • ad hoc to paper Stochastic I/O conditioning maintains a globally consistent appearance across all generated views.
    The method assumes that referencing random subsets of images at each DDIM step yields a consistent joint distribution. This is validated only through ablations (Section 5.3, Fig. S7), not by a formal guarantee.
  • domain assumption Multi-view attention over the reference and generated latents enforces cross-view consistency.
    Architectural assumption inherited from CAT3D (Section 3.1). The model is trained to align views via 3D attention, but the consistency is not proven formally.
  • domain assumption Background matting [55] reliably removes backgrounds in reference images without discarding identity-relevant content.
    Used as an automated preprocessing step (Supplementary A.2). Failure of matting can introduce artifacts, as acknowledged in the failure cases.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models." pith.science (2026). https://pith.science/paper/HRCYUZ5G

@misc{pith2026241212093,
  author       = {Pith},
  title        = {Pith review of: CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HRCYUZ5G}},
  note         = {Machine review of arXiv:2412.12093}
}
abstract

Reconstructing photorealistic and dynamic portrait avatars from images is essential to many applications including advertising, visual effects, and virtual reality. Depending on the application, avatar reconstruction involves different capture setups and constraints $-$ for example, visual effects studios use camera arrays to capture hundreds of reference images, while content creators may seek to animate a single portrait image downloaded from the internet. As such, there is a large and heterogeneous ecosystem of methods for avatar reconstruction. Techniques based on multi-view stereo or neural rendering achieve the highest quality results, but require hundreds of reference images. Recent generative models produce convincing avatars from a single reference image, but visual fidelity yet lags behind multi-view techniques. Here, we present CAP4D: an approach that uses a morphable multi-view diffusion model to reconstruct photoreal 4D (dynamic 3D) portrait avatars from any number of reference images (i.e., one to 100) and animate and render them in real time. Our approach demonstrates state-of-the-art performance for single-, few-, and multi-image 4D portrait avatar reconstruction, and takes steps to bridge the gap in visual fidelity between single-image and multi-view reconstruction techniques.

Figures

Figures reproduced from arXiv: 2412.12093 by the authors.

Figure 1
Figure 1. We present CAP4D: a method that generates 4D portrait avatars based on an arbitrary number of reference images (e.g., from one [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of CAP4D. (a) The method takes as input an arbitrary number of reference images Iref that are encoded into the latent space of a variational autoencoder [72]. An off-the-shelf face tracker estimates a 3DMM, Mref, for each reference image, from which we derive conditioning signals that describe camera view direction, Vref, head pose Pref, and expression Eref. We associate additional conditioning signals with… view at source ↗
Figure 3
Figure 3. Self-reenactment. Our approach is more realistic than baseline methods for self-reenactment from a single reference image (row 1), 10 reference images (row 2) and 100 reference images (row 3). The MMDM output (MMDM only) produces the most realistic output at the cost of temporal consistency compared to our reconstructed 4D Avatar (CAP4D). single reference image Method PSNR↑ LPIPS↓ CSIM↑ JOD↑ Voodoo3D [86] 19.05 0.38… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Cross-reenactment. Avatars are reconstructed from a single reference image (col. 1), and their expressions are driven by frames of a driving video (col. 2). The camera moves according to the indicated horizontal (H) and vertical (V) view angle. CAP4D faithfully recover…
Figure 5
Figure 5. Figure 5: Extensions. We demonstrate 4D appearance editing and relighting by applying CAP4D to images edited using off-the-shelf models [66, 109]. We also animate CAP4D avatars with a method that predicts 3DMM expressions from speech [94] (see supplement). human preference Metho…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

116 extracted references · 67 canonical work pages

  1. [1]

    Abdelrahman, Thorsten Hempel, Aly Khalifa, Ayoub Al-Hamadi, and Laslo Dinges

    Ahmed A. Abdelrahman, Thorsten Hempel, Aly Khalifa, Ayoub Al-Hamadi, and Laslo Dinges. L2CS-Net : Fine- grained gaze estimation in unconstrained environments. In Proc. ICFSP, pages 98–102, 2023. 6, 5

  2. [2]

    The digital Emily project: Achieving a photorealis- tic digital actor

    Oleg Alexander, Mike Rogers, William Lambeth, Jen-Yuan Chiang, Wan-Chun Ma, Chuan-Chang Wang, and Paul De- bevec. The digital Emily project: Achieving a photorealis- tic digital actor. IEEE Comput. Graph. Appl., 30(4):20–31,

  3. [3]

    RigNeRF: Fully controllable neural 3D portraits

    ShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman, and Zhixin Shu. RigNeRF: Fully controllable neural 3D portraits. In Proc. CVPR, 2022. 3

  4. [4]

    Bridging the gap: Studio-like avatar creation from a monocular phone capture

    ShahRukh Athar, Shunsuke Saito, Zhengyu Yang, Stanislav Pidhorsky, and Chen Cao. Bridging the gap: Studio-like avatar creation from a monocular phone capture. In Proc. ECCV, 2024. 3

  5. [5]

    TC4D: Trajectory-conditioned text-to-4D generation

    Sherwin Bahmani, Xian Liu, Wang Yifan, Ivan Sko- rokhodov, Victor Rong, Ziwei Liu, Xihui Liu, Jeong Joon Park, Sergey Tulyakov, Gordon Wetzstein, et al. TC4D: Trajectory-conditioned text-to-4D generation. In Proc. ECCV, 2024. 3

  6. [6]

    4D-fy: Text-to-4D generation using hy- brid score distillation sampling

    Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gor- don Wetzstein, Leonidas Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, and David B Lindell. 4D-fy: Text-to-4D generation using hy- brid score distillation sampling. In Proc. CVPR, 2024. 3

  7. [7]

    Learning personal- ized high quality volumetric head avatars from monocular RGB videos

    Ziqian Bai, Feitong Tan, Zeng Huang, Kripasindhu Sarkar, Danhang Tang, Di Qiu, Abhimitra Meka, Ruofei Du, Ming- song Dou, Sergio Orts-Escolano, et al. Learning personal- ized high quality volumetric head avatars from monocular RGB videos. In Proc. CVPR, 2023. 3

  8. [8]

    Im- agen 3

    Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brichtova, Andrew Bunner, Kelvin Chan, Yichang Chen, Sander Dieleman, Yuqing Du, Zach Eaton-Rosen, et al. Im- agen 3. arXiv preprint arXiv:2408.07009, 2024. 3

Show all 116 references
  1. [9]

    A morphable model for the synthesis of 3D faces

    V olker Blanz and Thomas Vetter. A morphable model for the synthesis of 3D faces. In Proc. SIGGRAPH, 1999. 1, 3

  2. [10]

    Stable video diffusion: Scaling latent video diffusion models to large datasets

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 1, 2

  3. [11]

    Align your latents: High-resolution video synthesis with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proc. CVPR, 2023. 3

  4. [12]

    Video generation models as world simulators, 2024

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, et al. Video generation models as world simulators, 2024. 3

  5. [13]

    How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks)

    Adrian Bulat and Georgios Tzimiropoulos. How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks). In Proc. ICCV,

  6. [14]

    Neural head reenactment with latent pose de- scriptors

    Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. Neural head reenactment with latent pose de- scriptors. In Proc. CVPR, 2020. 3

  7. [15]

    Authentic volumetric avatars from a phone scan

    Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, et al. Authentic volumetric avatars from a phone scan. ACM Trans. Graph., 41(4):1–19, 2022. 3

  8. [16]

    VideoCrafter2: Overcoming data limitations for high-quality video diffu- sion models

    Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter2: Overcoming data limitations for high-quality video diffu- sion models. In Proc. CVPR, 2024. 3

  9. [17]

    AniFaceDiff: High-fidelity face reenactment via facial parametric conditioned diffusion models

    Ken Chen, Sachith Seneviratne, Wei Wang, Dongting Hu, Sanjay Saha, Md Tarek Hasan, Sanka Rasnayaka, Tamasha Malepathirana, Mingming Gong, and Saman Halgamuge. AniFaceDiff: High-fidelity face reenactment via facial parametric conditioned diffusion models. arXiv preprint arXiv:2...

  10. [18]

    Morphable Diffusion: 3D- consistent diffusion for single-image avatar creation

    Xiyi Chen, Marko Mihajlovic, Shaofei Wang, Sergey Prokudin, and Siyu Tang. Morphable Diffusion: 3D- consistent diffusion for single-image avatar creation. In Proc. CVPR, 2024. 1, 3

  11. [19]

    Generalizable and an- imatable Gaussian head avatar

    Xuangeng Chu and Tatsuya Harada. Generalizable and an- imatable Gaussian head avatar. In Proc. NeurIPS, 2024. 3, 6, 7, 8, 5

  12. [20]

    Statistical modeling of craniofacial shape and texture

    Hang Dai, Nick Pears, William Smith, and Christian Dun- can. Statistical modeling of craniofacial shape and texture. Int. J. Comput. Vis., 128(2):547–571, 2020. 1

  13. [21]

    ArcFace: Additive angu- lar margin loss for deep face recognition

    Jiankang Deng, Jia Guo, Jing Yang, Niannan Xue, Irene Kotsia, and Stefanos Zafeiriou. ArcFace: Additive angu- lar margin loss for deep face recognition. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 44(10): 5962–5979, 2022. 6

  14. [22]

    Portrait4D: Learning one-shot 4D head avatar synthesis using synthetic data

    Yu Deng, Duomin Wang, Xiaohang Ren, Xingyu Chen, and Baoyuan Wang. Portrait4D: Learning one-shot 4D head avatar synthesis using synthetic data. In Proc. CVPR, 2024. 3

  15. [23]

    Portrait4D- v2: Pseudo multi-view data creates better 4D head synthe- sizer

    Yu Deng, Duomin Wang, and baoyuan Wang. Portrait4D- v2: Pseudo multi-view data creates better 4D head synthe- sizer. arXiv, 2024. 6, 7, 8

  16. [24]

    DiffusionRig: Learning personalized priors for facial appearance editing

    Zheng Ding, Xuaner Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, and Xiuming Zhang. DiffusionRig: Learning personalized priors for facial appearance editing. In Proc. CVPR, 2023. 1, 6, 7

  17. [25]

    HeadGAN: One-shot neural head synthesis and editing

    Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. HeadGAN: One-shot neural head synthesis and editing. In Proc. ICCV, 2021. 3

  18. [26]

    9 MegaPortraits: One-shot megapixel neural head avatars

    Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Alek- sei Ivakhnenko, Victor Lempitsky, and Egor Zakharov. 9 MegaPortraits: One-shot megapixel neural head avatars. In Proc. ACM-MM, 2022. 3

  19. [27]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Proc. ICML, 2024. 3

  20. [28]

    Black, and Timo Bolkart

    Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics (ToG), Proc. SIGGRAPH, 40(8), 2021. 3, 5

  21. [29]

    Multi-view stereo: A tutorial

    Yasutaka Furukawa, Carlos Hern ´andez, et al. Multi-view stereo: A tutorial. Found. Trends Comput. Graph. Vis., 9 (1-2):1–148, 2015. 1

  22. [30]

    Dynamic neural radiance fields for monocular 4D facial avatar reconstruction

    Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4D facial avatar reconstruction. In Proc. CVPR, 2021. 1, 3

  23. [31]

    Srini- vasan, Jonathan T

    Ruiqi Gao*, Aleksander Holynski*, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srini- vasan, Jonathan T. Barron, and Ben Poole*. Cat3d: Cre- ate anything in 3d with multi-view diffusion models. Proc. NeurIPS, 2024. 1, 3, 4, 6, 2

  24. [32]

    Reconstructing detailed dynamic face geometry from monocular video

    Pablo Garrido, Levi Valgaerts, Chenglei Wu, and Christian Theobalt. Reconstructing detailed dynamic face geometry from monocular video. ACM Trans. Graph., 32(6):1–10,

  25. [33]

    Learning neural parametric head models

    Simon Giebenhain, Tobias Kirschstein, Markos Geor- gopoulos, Martin R ¨unz, Lourdes Agapito, and Matthias Nießner. Learning neural parametric head models. In Proc. CVPR, 2023. 1

  26. [34]

    MonoNPHM: Dynamic head reconstruction from monocular videos

    Simon Giebenhain, Tobias Kirschstein, Markos Geor- gopoulos, Martin R ¨unz, Lourdes Agapito, and Matthias Nießner. MonoNPHM: Dynamic head reconstruction from monocular videos. In Proc. CVPR, 2024. 3

  27. [35]

    Neural head avatars from monocular rgb videos

    Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. Neural head avatars from monocular rgb videos. In Proc. CVPR,

  28. [36]

    DiffPortrait3D: Controllable diffusion for zero-shot portrait view synthesis

    Yuming Gu, Hongyi Xu, You Xie, Guoxian Song, Yichun Shi, Di Chang, Jing Yang, and Linjie Luo. DiffPortrait3D: Controllable diffusion for zero-shot portrait view synthesis. In Proc. CVPR, 2024. 1

  29. [37]

    Live- Portrait: Efficient portrait animation with stitching and re- targeting control

    Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. Live- Portrait: Efficient portrait animation with stitching and re- targeting control. arXiv preprint arXiv:2407.03168, 2024. 3

  30. [38]

    AnimateDiff: Animate your personalized text- to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. AnimateDiff: Animate your personalized text- to-image diffusion models without specific tuning. In Proc. ICLR, 2024. 3

  31. [39]

    Hancock and Jeremy N

    Jeffrey T. Hancock and Jeremy N. Bailenson. The social impact of deepfakes. Cyberpsychology, Behavior, and So- cial Networking, 24(3):149–152, 2021. PMID: 33760669. 8

  32. [40]

    Classifier-free diffusion guidance

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 6

  33. [41]

    LRM: Large reconstruction model for single im- age to 3D

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. LRM: Large reconstruction model for single im- age to 3D. arXiv preprint arXiv:2311.04400, 2023. 3

  34. [42]

    simple diffusion: End-to-end diffusion for high resolution images

    Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. In Proc. ICML, 2023. 1

  35. [43]

    GaussianAvatar: Towards realistic human avatar modeling from a single video via animatable 3D Gaussians

    Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Shengping Zhang, and Liqiang Nie. GaussianAvatar: Towards realistic human avatar modeling from a single video via animatable 3D Gaussians. In Proc. CVPR, 2024. 3

  36. [44]

    Dynamic 3D avatar creation from hand-held video input

    Alexandru Eugen Ichim, Sofien Bouaziz, and Mark Pauly. Dynamic 3D avatar creation from hand-held video input. ACM Trans. Graph., 34(4):1–14, 2015. 1, 3

  37. [45]

    A style-based generator architecture for generative adversarial networks

    Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proc. CVPR, 2019. 1

  38. [46]

    A style-based generator architecture for generative adversarial networks

    Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. IEEE Transactions on Pattern Analysis & Machine Intelli- gence, 43(12):4217–4228, 2021. 6

  39. [47]

    3D Gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):1–14,

  40. [48]

    Realistic one-shot mesh-based head avatars

    Taras Khakhulin, Vanessa Sklyarova, Victor Lempitsky, and Egor Zakharov. Realistic one-shot mesh-based head avatars. In Proc. ECCV, 2022. 3

  41. [49]

    Learn- ing to generate conditional tri-plane for 3D-aware expres- sion controllable portrait animation

    Taekyung Ki, Dongchan Min, and Gyeongsu Chae. Learn- ing to generate conditional tri-plane for 3D-aware expres- sion controllable portrait animation. In Proc. ECCV, 2024. 3

  42. [50]

    Nersemble: Multi-view ra- diance field reconstruction of human heads

    Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view ra- diance field reconstruction of human heads. ACM Trans. Graph., 42(4):1–14, 2023. 1, 6, 3, 4, 5

  43. [51]

    DiffusionAvatars: Deferred diffusion for high- fidelity 3D head avatars

    Tobias Kirschstein, Simon Giebenhain, and Matthias Nießner. DiffusionAvatars: Deferred diffusion for high- fidelity 3D head avatars. In Proc. CVPR, 2024. 5

  44. [52]

    Collaborative video diffusion: Consistent multi- video generation with camera control

    Zhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu, Hongsheng Li, Leonidas Guibas, and Gordon Wet- zstein. Collaborative video diffusion: Consistent multi- video generation with camera control. arXiv preprint arXiv:2405.17414, 2024. 3

  45. [53]

    Learning a model of facial shape and ex- pression from 4D scans

    Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and ex- pression from 4D scans. ACM Trans. Graph., 36(6):194–1,

  46. [54]

    Generalizable one-shot 3D neu- ral head avatar

    Xueting Li, Shalini De Mello, Sifei Liu, Koki Nagano, Umar Iqbal, and Jan Kautz. Generalizable one-shot 3D neu- ral head avatar. Proc. NeurIPS, 2024. 3

  47. [55]

    Robust high-resolution video mat- ting with temporal guidance, 2021

    Shanchuan Lin, Linjie Yang, Imran Saleemi, and Soumyadip Sengupta. Robust high-resolution video mat- ting with temporal guidance, 2021. 1 10

  48. [56]

    Common diffusion noise schedules and sample steps are flawed

    Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common diffusion noise schedules and sample steps are flawed. In Proc. WACV, pages 5392–5399, Los Alamitos, CA, USA, 2024. IEEE Computer Society. 1

  49. [57]

    Zero-1-to-3: Zero-shot one image to 3D object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. In Proc. CVPR, 2023. 1

  50. [58]

    Mix- ture of volumetric primitives for efficient neural rendering

    Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mix- ture of volumetric primitives for efficient neural rendering. ACM Trans. Graph., 40(4):1–13, 2021. 1

  51. [59]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proc. ICLR, 2019. 6

  52. [60]

    Pixel codec avatars

    Shugao Ma, Tomas Simon, Jason Saragih, Dawei Wang, Yuecheng Li, Fernando De La Torre, and Yaser Sheikh. Pixel codec avatars. In Proc. CVPR, 2021. 3

  53. [61]

    OTAvatar: One-shot talking face avatar with con- trollable tri-plane rendering

    Zhiyuan Ma, Xiangyu Zhu, Guo-Jun Qi, Zhen Lei, and Lei Zhang. OTAvatar: One-shot talking face avatar with con- trollable tri-plane rendering. In Proc. CVPR, 2023. 3

  54. [62]

    Mantiuk, Gyorgy Denes, Alexandre Chapiro, An- ton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, and Anjul Patney

    Rafał K. Mantiuk, Gyorgy Denes, Alexandre Chapiro, An- ton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, and Anjul Patney. FovVideoVDP: a visible difference pre- dictor for wide field-of-view video. ACM Trans. Graph., 40 (4), 2021. 6

  55. [63]

    Jewett, Si- mon Venshtain, Christopher Heilman, Yueh-Tung Chen, Sidi Fu, Mohamed Ezzeldin A

    Julieta Martinez, Emily Kim, Javier Romero, Timur Bagautdinov, Shunsuke Saito, Shoou-I Yu, Stuart Ander- son, Michael Zollh¨ofer, Te-Li Wang, Shaojie Bai, Chenghui Li, Shih-En Wei, Rohan Joshi, Wyatt Borsos, Tomas Si- mon, Jason Saragih, Paul Theodosis, Alexander Greene, Anjan...

  56. [64]

    NeRF: Representing scenes as neural radiance fields for view syn- thesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view syn- thesis. Commun. ACM., 65(1):99–106, 2021. 3, 5

  57. [65]

    V oxCeleb: A large-scale speaker identification dataset

    Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. V oxCeleb: A large-scale speaker identification dataset. In Proc. Interspeech, 2017. 1

  58. [66]

    DiFaReli: Diffusion face relighting

    Puntawat Ponglertnapakorn, Nontawat Tritrong, and Supa- sorn Suwajanakorn. DiFaReli: Diffusion face relighting. In Proc. CVPR, 2023. 8

  59. [67]

    DreamFusion: Text-to-3D using 2D diffusion

    Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. DreamFusion: Text-to-3D using 2D diffusion. InProc. ICLR, 2023. 3

  60. [68]

    Joker: Conditional 3D head synthesis with extreme facial expressions

    Malte Prinzler, Egor Zakharov, Vanessa Sklyarova, Berna Kabadayi, and Justus Thies. Joker: Conditional 3D head synthesis with extreme facial expressions. arXiv preprint arXiv:2410.16395, 2024. 1, 3

  61. [69]

    GaussianAvatars: Photorealistic head avatars with rigged 3D Gaussians

    Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Da- vide Davoli, Simon Giebenhain, and Matthias Nießner. GaussianAvatars: Photorealistic head avatars with rigged 3D Gaussians. In Proc. CVPR, 2024. 3, 4, 5, 6, 7, 10

  62. [70]

    LOLNeRF: Learn from one look

    Daniel Rebain, Mark Matthews, Kwang Moo Yi, Dmitry Lagun, and Andrea Tagliasacchi. LOLNeRF: Learn from one look. In Proc. CVPR, 2022. 3

  63. [71]

    Pirenderer: Controllable portrait image generation via semantic neural rendering

    Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu. Pirenderer: Controllable portrait image generation via semantic neural rendering. In Proc. ICCV, 2021. 3

  64. [72]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proc. CVPR ,

  65. [73]

    U- Net: Convolutional networks for biomedical image seg- mentation, 2015

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- Net: Convolutional networks for biomedical image seg- mentation, 2015. 6, 4

  66. [74]

    Relightable gaussian codec avatars

    Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. Relightable gaussian codec avatars. In Proc. CVPR, 2024. 8

  67. [75]

    PSA – a new scalable space partition based se- lection algorithm for MOEAs

    Shaul Salomon, Gideon Avigad, Alex Goldvard, and Oliver Sch¨utze. PSA – a new scalable space partition based se- lection algorithm for MOEAs. In Proc. EVOLVE, 2013. 3, 4

  68. [76]

    SplattingAvatar: Realistic real-time human avatars with mesh-embedded Gaussian splatting

    Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang. SplattingAvatar: Realistic real-time human avatars with mesh-embedded Gaussian splatting. In Proc. CVPR, 2024. 3

  69. [77]

    Zero123++: a single image to consis- tent multi-view diffusion base model

    Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consis- tent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023. 3

  70. [78]

    MVDream: Multi-view diffusion for 3D generation

    Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. MVDream: Multi-view diffusion for 3D generation. In Proc. ICLR, 2023. 1, 3

  71. [79]

    First order motion model for image animation

    Aliaksandr Siarohin, St ´ephane Lathuili `ere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation. Proc. NeurIPS, 2019. 1, 3

  72. [80]

    Denois- ing diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In Proc. ICLR, 2020. 3, 5

  73. [81]

    DreamCraft3D: Hierar- chical 3D generation with bootstrapped diffusion prior

    Jingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang, Wen Liu, Zhenda Xie, and Yebin Liu. DreamCraft3D: Hierar- chical 3D generation with bootstrapped diffusion prior. In Proc. ICLR, 2024. 3

  74. [82]

    Fourier fea- tures let networks learn high frequency functions in low di- mensional domains

    Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. Fourier fea- tures let networks learn high frequency functions in low di- mensional domains. Proc. NeurIPS, 2020. 5, 1 11

  75. [83]

    3D face tracking from 2D video through iterative dense uv to image flow

    Felix Taubner, Prashant Raina, Mathieu Tuli, Eu Wern Teh, Chul Lee, and Jinmiao Huang. 3D face tracking from 2D video through iterative dense uv to image flow. In Proc. CVPR, 2024. 1, 4, 6, 5

  76. [84]

    Advances in neural rendering

    Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srini- vasan, Edgar Tretschk, Wang Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lom- bardi, et al. Advances in neural rendering. Comput. Graph. Forum, 41(2):703–735, 2022. 1

  77. [85]

    EMO: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions

    Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. EMO: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions. arXiv preprint arXiv:2402.17485, 2024. 1

  78. [86]

    VOODOO 3D: V olumetric portrait disentanglement for one-shot 3D head reenactment

    Phong Tran, Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran, Liwen Hu, and Hao Li. VOODOO 3D: V olumetric portrait disentanglement for one-shot 3D head reenactment. Proc. CVPR, 2024. 3, 6, 7, 8

  79. [87]

    Real- time radiance fields for single-image portrait view synthe- sis

    Alex Trevithick, Matthew Chan, Michael Stengel, Eric Chan, Chao Liu, Zhiding Yu, Sameh Khamis, Manmohan Chandraker, Ravi Ramamoorthi, and Koki Nagano. Real- time radiance fields for single-image portrait view synthe- sis. ACM Trans. Graph., 42(4):1–15, 2023. 3

  80. [88]

    MEAD: A large-scale audio-visual dataset for emo- tional talking-face generation

    Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. MEAD: A large-scale audio-visual dataset for emo- tional talking-face generation. In Proc .ECCV, 2020. 6, 5

  81. [89]

    One- shot free-view neural talking-head synthesis for video con- ferencing

    Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One- shot free-view neural talking-head synthesis for video con- ferencing. In Proc. CVPR, 2021. 3

  82. [90]

    Novel view synthesis with diffusion models

    Daniel Watson, William Chan, Ricardo Martin Bru- alla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. In Proc. ICLR, 2023. 3

  83. [91]

    FlashAvatar: High-fidelity head avatar with efficient gaus- sian embedding

    Jun Xiang, Xuan Gao, Yudong Guo, and Juyong Zhang. FlashAvatar: High-fidelity head avatar with efficient gaus- sian embedding. In Proc. CVPR, 2024. 1, 3, 6, 7, 5

  84. [92]

    VFHQ: A high-quality dataset and bench- mark for video face super-resolution

    Liangbin Xie, Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan. VFHQ: A high-quality dataset and bench- mark for video face super-resolution. In Proc. CVPRW,

  85. [93]

    X-portrait: Expressive portrait ani- mation with hierarchical motion attention

    You Xie, Hongyi Xu, Guoxian Song, Chao Wang, Yichun Shi, and Linjie Luo. X-portrait: Expressive portrait ani- mation with hierarchical motion attention. In Proc. SIG- GRAPH, 2024. 1

  86. [94]

    CodeTalker: Speech- driven 3D facial animation with discrete motion prior

    Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. CodeTalker: Speech- driven 3D facial animation with discrete motion prior. In Proc. CVPR, pages 12780–12790, 2023. 8

  87. [95]

    Deep 3D portrait from a single image

    Sicheng Xu, Jiaolong Yang, Dong Chen, Fang Wen, Yu Deng, Yunde Jia, and Xin Tong. Deep 3D portrait from a single image. In Proc. CVPR, 2020. 3

  88. [96]

    Vasa-1: Lifelike audio-driven talking faces generated in real time

    Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. Vasa-1: Lifelike audio-driven talking faces generated in real time. arXiv preprint arXiv:2404.10667 ,

  89. [97]

    Gaussian head avatar: Ultra high-fidelity head avatar via dynamic Gaus- sians

    Yuelang Xu, Benwang Chen, Zhe Li, Hongwen Zhang, Lizhen Wang, Zerong Zheng, and Yebin Liu. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic Gaus- sians. In Proc. CVPR, 2024. 3

  90. [98]

    PV3D: A 3D generative model for portrait video genera- tion

    Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Wenqing Zhang, Song Bai, Jiashi Feng, and Mike Zheng Shou. PV3D: A 3D generative model for portrait video genera- tion. In Proc. ICLR, 2023. 3

  91. [99]

    FaceScape: a large- scale high quality 3D face dataset and detailed riggable 3D face prediction

    Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. FaceScape: a large- scale high quality 3D face dataset and detailed riggable 3D face prediction. In Proc. CVPR, 2020. 1, 3

  92. [100]

    Real3D-Portrait: One- shot realistic 3D talking portrait synthesis

    Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang, We- ichuang Li, Jiawei Huang, Ziyue Jiang, Jinzheng He, Rongjie Huang, Jinglin Liu, et al. Real3D-Portrait: One- shot realistic 3D talking portrait synthesis. In Proc. ICLR,

  93. [101]

    StyleHEAT: One-shot high- resolution editable talking face generation via pre-trained stylegan

    Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang. StyleHEAT: One-shot high- resolution editable talking face generation via pre-trained stylegan. In Proc. ECCV. Springer, 2022. 3

  94. [102]

    NOFA: NeRF-based one-shot facial avatar reconstruction

    Wangbo Yu, Yanbo Fan, Yong Zhang, Xuan Wang, Fei Yin, Yunpeng Bai, Yan-Pei Cao, Ying Shan, Yang Wu, Zhongqian Sun, et al. NOFA: NeRF-based one-shot facial avatar reconstruction. In Proc. SIGGRAPH, 2023. 3

  95. [103]

    Few-shot adversarial learning of realistic neural talking head models

    Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. Few-shot adversarial learning of realistic neural talking head models. In Proc. ICCV, 2019. 3

  96. [104]

    Fast bi-layer neural synthesis of one-shot realistic head avatars

    Egor Zakharov, Aleksei Ivakhnenko, Aliaksandra Shysheya, and Victor Lempitsky. Fast bi-layer neural synthesis of one-shot realistic head avatars. In Proc. ECCV, 2020

  97. [105]

    MetaPortrait: Identity-preserving talking head generation with fast personalized adaptation

    Bowen Zhang, Chenyang Qi, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, and Fang Wen. MetaPortrait: Identity-preserving talking head generation with fast personalized adaptation. In Proc. CVPR, 2023. 3

  98. [106]

    RodinHD: High-fidelity 3D avatar generation with diffusion models

    Bowen Zhang, Yiji Cheng, Chunyu Wang, Ting Zhang, Jiaolong Yang, Yansong Tang, Feng Zhao, Dong Chen, and Baining Guo. RodinHD: High-fidelity 3D avatar generation with diffusion models. arXiv preprint arXiv:2407.06938 ,

  99. [107]

    MediaPipe hands: On-device real-time hand tracking, 2020

    Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka, George Sung, Chuo-Ling Chang, and Matthias Grundmann. MediaPipe hands: On-device real-time hand tracking, 2020. 5

  100. [108]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. CVPR, 2018. 6, 4

  101. [109]

    Stable- Makeup: When real-world makeup transfer meets diffusion model, 2024

    Yuxuan Zhang, Lifu Wei, Qing Zhang, Yiren Song, Jiaming Liu, Huaxia Li, Xu Tang, Yao Hu, and Haibo Zhao. Stable- Makeup: When real-world makeup transfer meets diffusion model, 2024. 8

  102. [110]

    HAvatar: High-fidelity 12 head avatar via facial model conditioned neural radiance field

    Xiaochen Zhao, Lizhen Wang, Jingxiang Sun, Hongwen Zhang, Jinli Suo, and Yebin Liu. HAvatar: High-fidelity 12 head avatar via facial model conditioned neural radiance field. ACM Trans. Graph., 43(1):1–16, 2023. 3

  103. [111]

    I M avatar: Implicit morphable head avatars from videos

    Yufeng Zheng, Victoria Fern ´andez Abrevaya, Marcel C B¨uhler, Xu Chen, Michael J Black, and Otmar Hilliges. I M avatar: Implicit morphable head avatars from videos. In Proc. CVPR, 2022. 3

  104. [112]

    PointAvatar: Deformable point- based head avatars from videos

    Yufeng Zheng, Wang Yifan, Gordon Wetzstein, Michael J Black, and Otmar Hilliges. PointAvatar: Deformable point- based head avatars from videos. In Proc. CVPR, 2023. 3

  105. [113]

    Head- studio: Text to animatable head avatars with 3D Gaussian splatting

    Zhenglin Zhou, Fan Ma, Hehe Fan, and Yi Yang. Head- studio: Text to animatable head avatars with 3D Gaussian splatting. In Proc. ECCV, 2024. 3

  106. [114]

    Mo- FaNeRF: Morphable facial neural radiance field

    Yiyu Zhuang, Hao Zhu, Xusen Sun, and Xun Cao. Mo- FaNeRF: Morphable facial neural radiance field. In Proc. ECCV, 2022. 3

  107. [115]

    Instant volumetric head avatars

    Wojciech Zielonka, Timo Bolkart, and Justus Thies. Instant volumetric head avatars. Proc. CVPR, 2022. 1, 3

  108. [116]

    To- wards metrical reconstruction of human faces

    Wojciech Zielonka, Timo Bolkart, and Justus Thies. To- wards metrical reconstruction of human faces. In Proc. ECCV, 2022. 1, 3 13 CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models Supplementary Material This document includes supplementa...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.