REVIEW 4 major objections 5 minor 51 references
ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read ConsistentAvatar claims that aligning a high-frequency Fourier detail map from a coarse 3D proxy to real video frames makes diffusion-based talking-head generation consistent in time, pose, and expression, and reports state-of-the-art…
desk verdict A plausible incremental diffusion-based talking-head method with a genuinely new TSD detail-alignment idea, but the temporal-consistency claim is under-supported and the evaluation has leakage and overclaim issues. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Temporally-Sensitive Detail (TSD) map: a per-frame image formed by Fourier transforming the coarse RGB output, suppressing all frequencies below $w=10$, and inverse-transforming, so that it retains contours, expression edges, and high-frequency details that shift between adjacent frames. The argument is carried by two diffusion stages: the temporal consistency diffusion module (TSDM), which uses extra IP-Adapter-style cross-attention layers to align the coarse TSD to the ground-truth TSD and thereby learn the temporal pattern; and the fully consistent diffusion module, which uses ControlNet to condition the final generation on the aligned TSD, the coarse head normal, and an emotion embedding chosen by nearest-neighbor matching of DECA expression vectors against MEAD. The aligned TSD is the load-bearing guidance; the normal and emotion conditions add 3D and expression control once temporal stability is in place.
What would settle it
Run the full pipeline with several Fourier cutoffs (for instance, $w=5$, $10$, $20$, and a learned high-pass) on the same videos and measure adjacent-frame optical-flow error and DECA pose/expression error; if $w=10$ is not clearly the best or if the improvement over using unaligned TSD disappears for some cutoff, the claim that TSD defined by this cutoff carries the temporal pattern is not supported. A second decisive check is to evaluate on an unseen identity with large pose and expression changes: if temporal consistency degrades there, the alignment does not generalize beyond training conditions.
Extended reading notes
Core claim
On its own terms, the paper establishes the following: a temporally-sensitive detail map, obtained by taking the high-frequency Fourier component (cutoff $w=10$) of a coarse RGB render from INSTA and of the ground-truth frame, can be aligned to real-frame detail through a small diffusion module, and that aligned map, when fed as a condition alongside head normals and a CLIP emotion embedding into a ControlNet-based diffusion renderer, yields avatars that are closer to ground truth and substantially more stable across time, pose, and expression than prior methods. The authors' core claim is that the aligned TSD 'represents the temporal patterns' and constrains the diffusion process to generate temporally stable talking heads, with this reliable guidance compensating for inaccuracies in the other conditions. Quantitative comparisons on three datasets report lower L2, higher PSNR/SSIM, lower LPIPS, lower pose error, and lower expression error than the compared baselines.
Load-bearing premise
The load-bearing premise is that the high-frequency Fourier band at cutoff $w=10$ captures exactly the details that change between frames; if that band omits or distorts the relevant contours and expression edges, the alignment mechanism cannot stabilize generation.
Editorial extensions
If this is right
- Generated talking-head videos should show adjacent-frame optical flow comparable to real video, rather than the high jitter seen in DiffusionRig and the baseline without aligned TSD.
- Pose error (from DECA-estimated coefficients) and expression error should remain low across different viewpoints and expressions, not just on the training identities.
- The method retains high image quality while running in about 10 denoising steps thanks to LCM, making it practical for interactive avatar generation.
- Each condition plays a distinct role: removing aligned TSD hurts temporal and expression consistency, removing the normal condition hurts 3D consistency, and removing the emotion embedding leaves expressions less accurate.
Reading between the lines
- A natural extension is to test how sensitive the result is to the Fourier cutoff $w=10$; the paper reports no ablation on $w$, so a sweep over cutoffs would clarify whether the benefit comes from the specific band or from any high-frequency alignment.
- The same two-stage recipe—align a cheap proxy's high-frequency residual to the target, then condition the generator on it—could transfer to other conditional generation tasks where the cheap condition is accurate in structure but noisy in detail, such as pose-conditioned human video or depth-conditioned scene rendering.
- Because emotion labels are assigned by cosine similarity of DECA expression vectors, the method inherits DECA's expression ambiguities; an end-to-end learned emotion encoder or multiple hypothesis labels might sharpen expression control further.
- The reported temporal metric is optical-flow magnitude between adjacent frames; complementing it with point-tracking or learned video-quality metrics would test whether the stability holds beyond flow-specific artifacts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces ConsistentAvatar, a diffusion-based framework for talking-head avatar generation from monocular RGB video. The method first obtains coarse RGB and normal outputs from INSTA, then defines a Temporally-Sensitive Detail (TSD) map by high-pass Fourier filtering. A temporal consistency diffusion module (TSDM) is trained to align the coarse TSD to the TSD of the ground-truth frame. A fully consistent diffusion module (FCSD) then generates the final avatar conditioned on the aligned TSD, INSTA normal, and an emotion text embedding obtained by matching the target expression to MEAD via DECA. Experiments on the INSTA, PointAvatar, and NeRFace datasets show improved L2, PSNR, SSIM, LPIPS, pose error, and expression error over several baselines, and a qualitative optical-flow comparison indicates improved temporal stability. The paper also uses LCM and an SDXL-style refiner to reduce inference to about 2.4 seconds. The central claim is that aligning TSD, which 'represents the temporal patterns,' constrains the diffusion process to produce temporally stable and fully consistent avatars.
Significance. If the temporal-alignment mechanism is real, the TSD representation is a useful idea for improving consistency in diffusion-based talking-head generation. The paper provides a staged design that is easy to follow and reports consistent gains over multiple metrics and datasets, along with ablations of the TSD, normal, and emotion conditions. The computational efficiency gain from LCM is a positive aspect. However, the central temporal-consistency claim is not supported by a quantitative temporal metric or an ablation that isolates temporal ordering, and the emotion-labeling protocol has a potential leakage that may inflate the reported expression-consistency improvement. These issues need to be resolved before the claims can be accepted.
major comments (4)
- [Sec. 4.1, Eq. (2)-(3)] The definition of the Temporally-Sensitive Detail (TSD) map is not reproducible from the text. Equation (2) writes a one-dimensional Fourier integral with respect to a continuous variable t for an image I_rgb_i, defines W only as 'the frequency set of the image,' and then fixes w=10 without specifying units or the two-dimensional masking procedure. Since TSD is the central contribution and the input to both diffusion modules, please provide the exact 2D filtering algorithm (e.g., FFT radius, band-pass mask, or high-pass threshold) and a sensitivity analysis for the cutoff w=10.
- [Sec. 4.1, Sec. 5.2 (Fig. 7)] The temporal-consistency mechanism is asserted but not isolated. The TSDM (Eqs. 4-6) aligns each frame independently using only per-frame TSD, pose, and expression; there is no recurrence, no adjacent-frame input, and no video-level loss. Therefore any temporal stability improvement could stem from better per-frame reconstruction rather than from modeling temporal patterns. The only temporal evidence is the qualitative optical-flow plot in Fig. 7, with no quantitative numbers. Please add a quantitative temporal-consistency metric and an ablation that feeds temporally mismatched TSD (e.g., TSD from a different frame) to show that the temporal alignment itself, not just the aligned TSD's per-frame accuracy, is responsible for the improvement.
- [Sec. 4.2 (Eq. 7) and Sec. 5.2 (Expression Error)] The emotion-labeling protocol introduces a circularity in the expression-consistency evaluation. Equation (7) assigns an emotion label by matching the DECA expression vector of the target frame I_i against MEAD, and the Expression Error (EE) metric in Sec. 5.2 also uses DECA to compare generated expression coefficients to the same ground-truth target. Thus, at test time, the conditioning contains information directly derived from the target expression, which can inflate EE results independently of the generative model. Please evaluate expression consistency with emotion labels that are not computed from the target frame (e.g., from audio, manual annotation, or a separate emotion-conditioning experiment), and report whether the EE gain persists.
- [Sec. 5.1 (Evaluation Protocol) and Sec. 5.2 (Temporal consistency)] The evaluation protocol uses a per-identity split: the last 350 frames of each video are held out for testing, and the rest are used for training. Consequently, the method is only tested on identities seen during training, and the claim of 'fully consistent talking head avatar' is not evaluated for generalization to unseen identities. If the method is intended as a per-identity personalization system, please state that explicitly; if the title's generality is intended, add a cross-identity evaluation. Also, the optical-flow quantitative comparison promised in the Fig. 1 caption does not appear in the paper; the numbers should be reported.
minor comments (5)
- [Sec. 6] The abstract and Sec. 1 claim 'fully consistent' avatars, but Sec. 6 lists teeth and eyeball inaccuracies; consider using 'improved consistency' or 'consistent' with caveats.
- [Sec. 4.2] The text uses 'STD' once near the emotion-condition description; this should read 'TSD.'
- [Sec. 5.1] The dataset description says 'a resolution of 5122'; this should be '512×512.'
- [Tab. 1 and Sec. 5.2] The sentence about reducing inference time is vague; it should explicitly compare the 'w/o LCM' row (8.20s) with the 'Ours' row (2.40s) to attribute the gain to LCM.
- [Tab. 1 and Tab. 2] Given the small number of training videos (10 for INSTA), please report per-sequence results or error bars across repeated runs to support the quantitative claims.
Circularity Check
No circularity: the pipeline is a supervised conditional-diffusion system, and the temporal-consistency claim is an emergent downstream effect rather than a quantity fitted and then re-measured.
full rationale
ConsistentAvatar is an empirical, supervised pipeline rather than a derivation, and I find no load-bearing step that reduces to its own inputs by construction. Stage 2 (TSDM) is trained with the per-frame denoising loss L_TSD = || eps_hat_t - eps_t ||^2_2 (Eq. 6) to map the coarse TSD of INSTA's RGB output to the ground-truth TSD of the corresponding video frame; Stage 3 (FCSD) is trained with the analogous per-frame loss L_portrait (Eq. 9). Neither loss uses the optical-flow temporal metric, adjacent-frame inputs, or a video-level objective, so the reported temporal consistency is an emergent downstream property, not a quantity that is optimized and then re-measured. The TSD defined in Eqs. 2-3 is a per-frame Fourier high-frequency/contour map; calling it "temporal patterns" is a semantic overclaim and a correctness risk (the cutoff w=10 is not justified and no recurrence is used), but a per-frame condition is not circular merely because temporal smoothness is inherited from the driving pose and expression. The emotion condition in Eq. 7 is derived from the target frame's DECA expression, and the EE metric in Sec. 5.2 also uses DECA, which is a mild target-information leak for the expression-consistency ablation; however, the 8-class emotion label is a coarse categorical projection of the 100-dim expression vector, so the EE result is not forced by construction. The paper cites only external systems (INSTA, DECA, MEAD, SD, LCM, SDXL, IP-Adapter, ControlNet) with no load-bearing self-citation chain, and its image-quality and 3D-consistency evaluations are standard reconstruction benchmarks. The acknowledged limitation (missing teeth and eyeball geometry, Sec. 6) is independent of the claimed mechanism. Overall, there is no circularity; the weaknesses are experimental-control and interpretability issues, not self-referential derivation.
Assumptions & free parameters
free parameters (1)
- Fourier cutoff frequency w =
10
assumptions (5)
- domain assumption INSTA provides a sufficiently accurate coarse RGB and normal proxy for the target head pose and expression.
- domain assumption A Fourier high-pass filter at w=10 isolates the temporally sensitive details (contours, high-frequency features) needed for consistency.
- domain assumption DECA expression vectors from the target frame can be matched by cosine similarity to MEAD clips to recover the correct emotion label.
- domain assumption Pretrained Stable Diffusion, ControlNet, IP-Adapter, and LCM can be combined with the learned conditions without interfering.
- standard math Fourier transform and inverse transform behave as standard linear operations for image filtering.
invented entities (1)
-
Temporally-Sensitive Detail (TSD) map
Cite this review
Pith. "Pith review of ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance." pith.science (2026). https://pith.science/paper/CVKO6UJG
@misc{pith2026241115436,
author = {Pith},
title = {Pith review of: ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance},
year = {2026},
howpublished = {\url{https://pith.science/paper/CVKO6UJG}},
note = {Machine review of arXiv:2411.15436}
}
read the original abstract
Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framework for fully consistent and high-fidelity talking avatar generation. Instead of directly employing multi-modal conditions to the diffusion process, our method learns to first model the temporal representation for stability between adjacent frames. Specifically, we propose a Temporally-Sensitive Detail (TSD) map containing high-frequency feature and contours that vary significantly along the time axis. Using a temporal consistent diffusion module, we learn to align TSD of the initial result to that of the video frame ground truth. The final avatar is generated by a fully consistent diffusion module, conditioned on the aligned TSD, rough head normal, and emotion prompt embedding. We find that the aligned TSD, which represents the temporal patterns, constrains the diffusion process to generate temporally stable talking head. Further, its reliable guidance complements the inaccuracy of other conditions, suppressing the accumulated error while improving the consistency on various aspects. Extensive experiments demonstrate that ConsistentAvatar outperforms the state-of-the-art methods on the generated appearance, 3D, expression and temporal consistency. Project page: https://njust-yang.github.io/ConsistentAvatar.github.io/
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
ShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman, and Zhixin Shu
-
[2]
Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques. 187–194
work page 1999
-
[3]
EMOCA: Emotion Driven Monocular Face Capture and Animation
Radek Danecek, Michael J. Black, and Timo Bolkart. 2022. EMOCA: Emotion Driven Monocular Face Capture and Animation. arXiv:2204.11312 [cs.CV]
work page Pith review arXiv 2022
-
[4]
Yu Deng, Jiaolong Yang, Dong Chen, Fang Wen, and Xin Tong. 2020. Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning. arXiv:2004.11660 [cs.CV]
work page Pith review arXiv 2020
-
[5]
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. 2020. Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set. arXiv:1903.08527 [cs.CV]
arXiv 2020
-
[6]
Zheng Ding, Xuaner Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, and Xiuming Zhang. 2023. DiffusionRig: Learning Personalized Priors for Facial Appearance Editing. arXiv:2304.06711 [cs.CV]
work page Pith review arXiv 2023
-
[7]
Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. 2021. Learn- ing an Animatable Detailed 3D Face Model from In-The-Wild Images. arXiv:2012.04012 [cs.CV]
work page Pith review arXiv 2021
-
[8]
Guy Gafni, Justus Thies, Michael Zollhöfer, and Matthias Nießner. 2021. Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. 8649–8658
work page 2021
Show all 51 references
-
[9]
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz, Lour- des Agapito, and Matthias Nießner. 2023. Learning Neural Parametric Head Models. arXiv:2212.02761 [cs.CV]
2023 arXiv
-
[10]
Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. 2022. Neural Head Avatars from Monocular RGB Videos. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022...
2022
-
[11]
Kuangxiao Gu, Yuqian Zhou, and Thomas Huang. 2019. FLNet: Landmark Driven Fetching and Learning Network for Faithful Talking Facial Animation Synthesis. arXiv:1911.09224 [cs.CV]
2019 arXiv
-
[12]
Yue Han, Jiangning Zhang, Junwei Zhu, Xiangtai Li, Yanhao Ge, Wei Li, Chengjie Wang, Yong Liu, Xiaoming Liu, and Ying Tai. 2023. A Generalist FaceX via Learning Unified Facial Representation. arXiv:2401.00551 [cs.CV]
2023 arXiv
-
[13]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. arXiv:2006.11239 [cs.LG]
2020 arXiv
-
[14]
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. 2021. Audio-Driven Emotional Video Portraits. arXiv:2104.07452 [cs.CV]
2021 arXiv
-
[15]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A Style-Based Generator Archi- tecture for Generative Adversarial Networks. arXiv:1812.04948 [cs.NE]
2019 arXiv
-
[16]
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and Improving the Image Quality of StyleGAN. arXiv:1912.04958 [cs.CV]
2020 arXiv
-
[17]
Taras Khakhulin, Vanessa Sklyarova, Victor Lempitsky, and Egor Zakharov. 2022. Realistic One-shot Mesh-based Head Avatars. arXiv:2206.08343 [cs.CV]
2022 arXiv
-
[18]
Gyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi, Junho Kim, and Eunho Yang. 2023. Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video Encoding. arXiv:2212.02802 [cs.CV]
2023 arXiv
-
[19]
Minchul Kim, Feng Liu, Anil Jain, and Xiaoming Liu. 2023. DCFace: Synthetic Face Generation with Dual Condition Diffusion Model. arXiv:2304.07060 [cs.CV]
2023 arXiv
-
[20]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , Yoshua Bengio and Yann LeCun (Eds.)
2015
-
[21]
Tobias Kirschstein, Simon Giebenhain, and Matthias Nießner. 2023. Dif- fusionAvatars: Deferred Diffusion for High-fidelity 3D Head Avatars. arXiv:2311.18635 [cs.CV]
2023 arXiv
-
[22]
Black, Hao Li, and Javier Romero
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero. 2017. Learning a model of facial shape and expression from 4D scans. ACM Trans. Graph. 36, 6 (2017), 194:1–194:17
2017
-
[23]
Simian Luo, Yiqin Tan, Longbo Huang, Jian Li, and Hang Zhao. 2023. Latent Con- sistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv:2310.04378 [cs.CV]
2023 arXiv
-
[24]
Thu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang, and Niloy J. Mitra. 2020. BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processi...
2020
-
[25]
Alex Nichol and Prafulla Dhariwal. 2021. Improved Denoising Diffusion Proba- bilistic Models. arXiv:2102.09672 [cs.LG]
2021 arXiv
-
[26]
Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. 2022. Unsuper- vised Learning of Efficient Geometry-Aware Neural Articulated Representations. arXiv:2204.08839 [cs.CV]
2022 arXiv
-
[27]
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. 2023. EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation. arXiv:2303.11089 [cs.CV]
2023 arXiv
-
[28]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV]
2023 arXiv
-
[29]
Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. 2023. GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. arXiv:2312.02069 [cs.CV]
2023 arXiv
-
[30]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...
2021 arXiv
-
[31]
Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.)
Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 2021. Advances in Neural Information Process- ing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual
2021
-
[32]
Li, and Shan Liu
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li, and Shan Liu. 2021. PIRen- derer: Controllable Portrait Image Generation via Semantic Neural Rendering. arXiv:2109.08379 [cs.CV]
2021 arXiv
-
[33]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In CVPR. 10684–10695
2022
-
[34]
Soubhik Sanyal, Timo Bolkart, Haiwen Feng, and Michael J. Black. 2019. Learning to Regress 3D Face Shape and Expression from an Image without 3D Supervision. arXiv:1905.06817 [cs.CV]
2019 arXiv
-
[35]
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023. DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation. arXiv:2301.03786 [cs.CV]
2023 arXiv
-
[36]
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2020. First Order Motion Model for Image Animation. arXiv:2003.00196 [cs.CV]
2020 arXiv
-
[37]
Sanjana Sinha, Sandika Biswas, Ravindra Yadav, and Brojeshwar Bhowmick
-
[38]
Shuai Tan, Bin Ji, and Ye Pan. 2023. EMMN: Emotional Motion Memory Net- work for Audio-driven Emotional Talking Face Generation. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . 22089–22099
2023
-
[39]
arXiv:2205.01155 [cs.CV]
Emotion-Controllable Generalized Talking Face Generation. arXiv:2205.01155 [cs.CV]
-
[40]
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020. MEAD: A Large-Scale Audio- Visual Dataset for Emotional Talking-Face Generation. InComputer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, Augus...
2020
-
[41]
Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj, Florian Bernard, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, and Christian Theobalt. 2020. StyleRig: Rigging StyleGAN for 3D Control Over Portrait Images. In 2020 IEEE/CVF Con- ference on Computer Vision and Pattern Recog...
2020
-
[42]
Yuelang Xu, Lizhen Wang, Xiaochen Zhao, Hongwen Zhang, and Yebin Liu. 2023. AvatarMAV: Fast 3D Head Avatar Reconstruction Using Motion-Aware Neural Voxels. In ACM SIGGRAPH 2023 Conference Proceedings, SIGGRAPH 2023, Los Angeles, CA, USA, August 6-10, 2023 , Erik Brunvand, Alla...
2023
-
[43]
Yibo Xia, Lizhen Wang, Xiang Deng, Xiaoyan Luo, and Yebin Liu. 2023. GMTalker: Gaussian Mixture based Emotional talking video Portraits. arXiv:2312.07669 [cs.CV]
2023 arXiv
-
[44]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. arXiv:2302.05543 [cs.CV]
2023 arXiv
-
[45]
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023. IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models. arXiv:2308.06721 [cs.CV]
2023 arXiv
-
[46]
Ruiqi Zhao, Tianyi Wu, and Guodong Guo. 2021. Sparse to Dense Motion Transfer for Face Image Animation. arXiv:2109.00471 [cs.CV]
2021 arXiv
-
[47]
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2023. SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation. arXiv:2211.12194 [cs.CV]
2023 arXiv
-
[48]
Black, and Otmar Hilliges
Yufeng Zheng, Wang Yifan, Gordon Wetzstein, Michael J. Black, and Otmar Hilliges. 2023. PointAvatar: Deformable Point-based Head Avatars from Videos. arXiv:2212.08377 [cs.CV]
2023 arXiv
-
[49]
Bühler, Xu Chen, Michael J
Yufeng Zheng, Victoria Fernández Abrevaya, Marcel C. Bühler, Xu Chen, Michael J. Black, and Otmar Hilliges. 2022. I M Avatar: Implicit Morphable Head Avatars from Videos. In IEEE/CVF Conference on Computer Vision and MM’24, October 28 - November 1, 2024, Melbourne, Australia. ...
2022
-
[51]
Wojciech Zielonka, Timo Bolkart, and Justus Thies. 2023. Instant Volumetric Head Avatars. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4574–4584
2023
-
[2022]
arXiv:2206.06481 [cs.CV]
RigNeRF: Fully Controllable Neural 3D Portraits. arXiv:2206.06481 [cs.CV]
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.