Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Conditioning a multimodal diffusion transformer on a joint identity-expression representation, plus consistent attention at inference, can generate fine-grained facial expressions while keeping the target identity consistent across the expr

desk verdict Plausible abstract, unreadable body: don't verify anything from this copy; revisit if a clean version appears. read the letter →

arxiv 2508.09461 v1 pith:JYI26TZJ submitted 2025-08-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords avatargenerationfine-grainedfacialexpressionsidentitypreservationdiffusiontransformerconsistentattentionpersonalizedavatarsexpressionsynthesismultimodal
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

GEN-AFFECT is a framework for turning a person into a customized 2D avatar that can show many fine-grained facial expressions while still looking like the same person. The paper's central claim is that a single learned representation can carry both who the person is and which expression they are making, and that feeding that representation into a multimodal diffusion transformer avoids the usual trade-off between expression accuracy and identity preservation. To keep the whole set of generated expressions consistent with one identity, the paper adds consistent attention at inference, letting the generated images share identity information across the expression array. If correct, this would give game, virtual-communication, education, and content-creation pipelines a direct way to generate a coherent expressive avatar set from one reference image. The paper reports that this approach outperforms previous methods on expression accuracy, identity preservation, and identity consistency.

What carries the argument

The central object is the identity-expression representation, a single learned embedding that fuses who the avatar is with the expression it should show, used to condition a multimodal diffusion transformer. The second load-bearing mechanism is consistent attention at inference, which shares identity-related information across the set of generated expressions so the whole array stays anchored to one identity.

What would settle it

Use a held-out identity, generate the full set of fine-grained expressions, and compare face-recognition embeddings of the generated images. If the average distance between expressions of the same identity approaches the distance between different identities, or if an expression classifier cannot reliably tell the generated expressions apart, the identity-expression representation is not doing its job.

Watch

Extended reading notes

Core claim

The paper proposes that identity and expression should not be treated as separate conditions fighting for control of the generated face. Instead, GEN-AFFECT extracts an identity-expression representation that fuses both factors and conditions a multimodal diffusion transformer on it. At inference, the framework applies consistent attention across the set of generated expressions so that information about the target identity is shared among all generated images. The claimed outcome is a set of avatars that each render a specific fine-grained expression accurately while remaining recognizably the same person. The paper asserts this design beats previous state-of-the-art methods on expression a

Load-bearing premise

The load-bearing premise is that identity and expression can be encoded together in one learned representation without the two factors bleeding into each other, so changing the expression does not change the identity and fixing the identity does not erase the expression.

Editorial extensions

If this is right

  • A single reference image could produce an entire avatar set in which each image shows a distinct fine-grained expression but still reads as the same person.
  • Expression fidelity and identity preservation do not have to be traded off: the same conditioning mechanism handles both goals.
  • Consistent attention offers an inference-time route to identity consistency across a generated expression set, without requiring additional training.
  • Custom avatar pipelines for gaming, virtual communication, education, and content creation could use one framework to generate a user's full expressive avatar set.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The joint identity-expression representation could be probed directly to test whether it truly disentangles identity from expression; interpolating between expressions should leave the identity direction of the embedding unchanged, an analysis the paper does not report.
  • Consistent attention at inference suggests a general recipe for any batch of related generations, such as video frames or multi-view images, where consistency across the set matters.
  • Because the claim rests on a single shared representation, generalization to unseen or rare identities may require additional identity anchors; this is a testable extension beyond the paper's reported results.
  • Reporting expression accuracy per fine-grained category, rather than only aggregate accuracy, would reveal whether identity preservation subtly dulls the most delicate expressions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes GEN-AFFECT, a framework for personalized 2D avatar generation that conditions a multimodal diffusion transformer on an extracted identity-expression representation and employs consistent attention at inference to maintain identity across a set of fine-grained facial expressions. The abstract claims superior performance over previous state-of-the-art methods in expression accuracy, identity preservation, and identity consistency across expressions. However, the supplied full text is almost entirely unreadable mojibake, so the method architecture, training objective, inference procedure, and experimental results cannot be inspected or verified.

Significance. If the claims hold, this work would address a genuine and difficult problem in avatar generation: preserving identity while generating fine-grained facial expressions. The proposed ideas—joint identity-expression conditioning and inference-time information sharing—are plausible and potentially useful to the community. I credit the authors for a clear abstract and a well-defined application. However, as delivered, the manuscript does not allow evaluation: no reproducible code, no accessible tables or metrics, no parameter-free derivations, and no way to confirm the stated superior performance. The significance therefore remains prospective rather than demonstrated.

major comments (4)
  1. [Full text] The supplied full text is unreadable due to widespread encoding corruption: equations, figures, tables, baselines, and metrics are inaccessible. The central claim of 'superior performance compared to previous state-of-the-art methods' cannot be checked. This is load-bearing for the paper's contribution; without readable experiments, the claim is unsupported.
  2. [Full text page header] The embedded header reads 'arXiv:2508.09462v1 [cs.LG]', while the stated record is arXiv:2508.09461 (cs.CV). This version/provenance mismatch makes it unclear which manuscript is being reviewed and undermines reproducibility. The authors should supply a clean text with matching identifiers.
  3. [Abstract / Method] The method conditions on an 'extracted identity-expression representation', but no disentanglement loss, orthogonality constraint, or representation analysis is described in the readable portion. If identity and expression are confounded in this representation, the two headline goals—expression accuracy and identity preservation—will trade off. The authors should provide an ablation that fixes identity while varying expressions and measures identity-embedding distance, and vice versa.
  4. [Abstract / Inference] The 'consistent attention at inference' is underspecified. If it is implemented as unconstrained cross-sample attention, it could inflate identity-consistency metrics by averaging away fine-grained expression-specific details. The authors should specify the attention restriction and show empirically that it does not suppress expression diversity, e.g., by comparing generated-expression variance with and without consistent attention.
minor comments (4)
  1. [Title] The capitalization 'identiTy' is inconsistent; use 'Identity'.
  2. [Abstract] The abstract repeats 'fine-grained facial expressions' and 'array of generated expressions' multiple times; tightening would improve readability.
  3. [Related work] The corrupted text prevents access to references. Once restored, the paper should explicitly compare with recent diffusion-based avatar generation and expression-transfer methods, including both identity and expression metrics.
  4. [Abstract / Scope] The phrase 'different forms of customized 2D avatars' is vague; specifying the target domain (e.g., cartoon avatars, portrait stylization, or photorealistic avatars) would sharpen the contribution.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable from the readable abstract; corrupted body prevents deeper inspection.

full rationale

Only the abstract is legible; the supplied full text is mojibake, and the embedded page header carries arXiv:2508.09462 while the stated record is 2508.09461. Under the applicable standard, circularity must be demonstrated by quoting the paper and exhibiting a specific reduction (e.g., Eq. X equals Eq. Y by construction, or a fitted parameter renamed as a prediction). The abstract describes conditioning a multimodal diffusion transformer on an extracted identity-expression representation and employing consistent attention at inference for information sharing across generated expressions. It does not define that representation as the identity metric, does not state that any evaluation quantity is fitted or derived from the conditioning input by construction, and contains no visible self-citations or imported uniqueness theorems. The strongest claim is comparative superiority over prior methods, which is an empirical claim rather than a derivation from its own inputs. The reader's hypothesized risk—that the conditioning embedding and the identity evaluation metric may share the same embedding family—is not established by any quoted text and would, at most, be a representational-entanglement concern, not a demonstrated circular step. Because no load-bearing circular reduction can be exhibited from the readable material, the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

The supplied full text is encoding-corrupted, so this ledger is abstract-level. No free parameters could be extracted (the abstract enumerates none). The five axioms are the load-bearing assumptions: representational disentanglement, diffusion conditioning fidelity, identity-selective sharing behavior of consistent attention, external evaluation fairness, and standard diffusion-transformer background. The two invented entities are internal architectural constructs without independent evidence. The ledger is provisional and would need the readable method and experiments sections to be completed.

assumptions (5)
  • domain assumption Identity and expression information can be jointly encoded in one representation extracted from a reference image, with sufficient disentanglement that either factor can vary without distorting the other.
    The framework 'conditions a multimodal diffusion transformer on an extracted identity-expression representation' (abstract). If the embedding entangles identity and expression, both identity preservation and expression accuracy fail simultaneously.
  • domain assumption A diffusion transformer conditioned on the identity-expression representation can render fine-grained expression differences in output pixels.
    The generative claim assumes the conditioned denoiser has the capacity and conditioning fidelity to realize subtle expression variations; load-bearing but unverifiable from the corrupted text.
  • domain assumption Inference-time consistent attention across the expression set improves identity consistency without suppressing expression-specific detail.
    The abstract's consistency mechanism assumes cross-sample information sharing is identity-selective; if shared features homogenize expressions, the fine-grained expression accuracy claim fails.
  • domain assumption The evaluation uses externally grounded metrics and fairly tuned baselines for expression accuracy, identity preservation, and identity consistency.
    The claimed SOTA result depends on benchmark and metric choices that cannot be inspected in the unreadable experiments section.
  • standard math Diffusion transformers (DiT-style denoisers) provide the established background generative machinery used by the method.
    Background assumption from prior literature; the paper extends rather than re-derives the diffusion machinery.
invented entities (2)
  • identity-expression representation
    purpose: joint conditioning signal injected into the diffusion transformer to control identity and expression jointly
    Introduced in the abstract as the extracted conditioning representation. It is an internal latent; its utility is evidenced only by the paper's own reported metrics, which are unreadable in this copy. No external falsifiable handle is provided.
  • consistent attention (inference-time)
    purpose: cross-sample information sharing across the set of generated expressions to maintain identity consistency
    Described in the abstract as 'consistent attention at inference for information sharing across the set of generated expressions.' Its effectiveness is evidenced only by the paper's consistency evaluation, unreadable here; no independent external evidence exists.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy." pith.science (2026). https://pith.science/paper/JYI26TZJ

@misc{pith2026250809461,
  author       = {Pith},
  title        = {Pith review of: Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JYI26TZJ}},
  note         = {Machine review of arXiv:2508.09461}
}
read the original abstract

Different forms of customized 2D avatars are widely used in gaming applications, virtual communication, education, and content creation. However, existing approaches often fail to capture fine-grained facial expressions and struggle to preserve identity across different expressions. We propose GEN-AFFECT, a novel framework for personalized avatar generation that generates expressive and identity-consistent avatars with a diverse set of facial expressions. Our framework proposes conditioning a multimodal diffusion transformer on an extracted identity-expression representation. This enables identity preservation and representation of a wide range of facial expressions. GEN-AFFECT additionally employs consistent attention at inference for information sharing across the set of generated expressions, enabling the generation process to maintain identity consistency over the array of generated fine-grained expressions. GEN-AFFECT demonstrates superior performance compared to previous state-of-the-art methods on the basis of the accuracy of the generated expressions, the preservation of the identity and the consistency of the target identity across an array of fine-grained facial expressions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 34 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Emo S tyle: ne-shot facial expression editing using continuous emotion parameters

    Bita Azari and Angelica Lim. Emo S tyle: ne-shot facial expression editing using continuous emotion parameters. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 6385--6394, 2024

  3. [3]

    Layer normalization

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

  4. [4]

    Improving image generation with better captions

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2 0 (3): 0 8, 2023

  5. [5]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv\'e J\'egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the International Conference on Computer Vision (ICCV), 2021

  6. [6]

    Semantic-rich facial emotional expression recognition

    Keyu Chen, Xu Yang, Changjie Fan, Wei Zhang, and Yu Ding. Semantic-rich facial emotional expression recognition. IEEE Transactions on Affective Computing, 13 0 (4): 0 1906--1916, 2022

  7. [7]

    Star GAN: U nified generative adversarial networks for multi-domain image-to-image translation

    Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. Star GAN: U nified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8789--8797, 2018

  8. [8]

    GAN mut: L earning interpretable conditional space for gamut of emotions

    Stefano d'Apolito, Danda Pani Paudel, Zhiwu Huang, Andres Romero, and Luc Van Gool. GAN mut: L earning interpretable conditional space for gamut of emotions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 568--577, 2021

Show all 52 references
  1. [9]

    Arcface: Additive angular margin loss for deep face recognition

    Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4690--4699, 2019

  2. [10]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M \"u ller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on mach...

  3. [11]

    AI -based avatars are changing the way we learn and teach: benefits and challenges

    Maximilian C Fink, Seth A Robinson, and Bernhard Ertl. AI -based avatars are changing the way we learn and teach: benefits and challenges. In Frontiers in Education, page 1416307. Frontiers Media SA, 2024

  4. [12]

    An image is worth one word: Personalizing text-to-image generation using textual inversion

    Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. An image is worth one word: Personalizing text-to-image generation using textual inversion. In The Eleventh International Conference on Learning Representations, 2023

  5. [13]

    Generative adversarial networks

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63 0 (11): 0 139--144, 2020

  6. [14]

    Pulid: Pure and lightning id customization via contrastive alignment

    Zinan Guo, Yanze Wu, Chen Zhuowei, Peng Zhang, Qian He, et al. Pulid: Pure and lightning id customization via contrastive alignment. Advances in neural information processing systems, 37: 0 36777--36804, 2024

  7. [15]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33: 0 6840--6851, 2020

  8. [16]

    Progressive growing of gans for improved quality, stability, and variation

    Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017

  9. [17]

    Photomaker: Customizing realistic human photos via stacked id embedding

    Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, and Ying Shan. Photomaker: Customizing realistic human photos via stacked id embedding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8640--8650, 2024

  10. [18]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023

  11. [19]

    Towards a simultaneous and granular identity-expression control in personalized face generation

    Renshuai Liu, Bowen Ma, Wei Zhang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, and Xuan Cheng. Towards a simultaneous and granular identity-expression control in personalized face generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  12. [20]

    Flow straight and fast: Learning to generate and transfer data with rectified flow

    Xingchao Liu, Chengyue Gong, and qiang liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2023

  13. [21]

    Deep learning face attributes in the wild

    Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), 2015

  14. [22]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019

  15. [23]

    Affectnet: A database for facial expression, valence, and arousal computing in the wild

    Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor. Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing, 10 0 (1): 0 18--31, 2017

  16. [24]

    Avatar design types and user engagement in digital educational games during evaluation phase

    Dinna N Mohd Nizam, Dylia Nursakinaz Rudiyansah, Nooralisa Mohd Tuah, Zaidatol Haslinda Abdullah Sani, and Kornchulee Sungkaew. Avatar design types and user engagement in digital educational games during evaluation phase. International Journal of Electrical and Computer Engine...

  17. [25]

    DINOv2: L earning robust visual features without supervision

    Maxime Oquab, Timoth \'e e Darcet, Th \'e o Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: L earning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023

  18. [26]

    All T ogether: E ffect of avatars in mixed-modality conferencing environments

    Payod Panda, Molly Jane Nicholas, Mar Gonzalez-Franco, Kori Inkpen, Eyal Ofek, Ross Cutler, Ken Hinckley, and Jaron Lanier. All T ogether: E ffect of avatars in mixed-modality conferencing environments. In Proceedings of the 1st Annual Meeting of the Symposium on Human-Compute...

  19. [27]

    in-the-wild

    Foivos Paraperas Papantoniou, Panagiotis P. Filntisis, Petros Maragos, and Anastasios Roussos. Neural emotion director: Speech-preserving semantic control of facial expressions in "in-the-wild" videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  20. [28]

    Portraitbooth: A versatile portrait model for fast identity-preserved personalization

    Xu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai, Donghao Luo, Jiangning Zhang, Wei Lin, Taisong Jin, Chengjie Wang, and Rongrong Ji. Portraitbooth: A versatile portrait model for fast identity-preserved personalization. In Proceedings of the IEEE/CVF Conference on Computer Vision ...

  21. [29]

    Photorealistic and identity-preserving image-based emotion manipulation with latent diffusion models

    Ioannis Pikoulis, Panagiotis P Filntisis, and Petros Maragos. Photorealistic and identity-preserving image-based emotion manipulation with latent diffusion models. arXiv preprint arXiv:2308.03183, 2023

  22. [30]

    Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer

    Albert Pumarola, Antonio Agudo, Aleix M. Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. Ganimation: Anatomically-aware facial animation from a single image. In Computer Vision – ECCV 2018: 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, P...

  23. [31]

    GAN imation: A natomically-aware facial animation from a single image

    Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. GAN imation: A natomically-aware facial animation from a single image. In Proceedings of the European conference on Computer Vision (ECCV), pages 818--833, 2018 b

  24. [32]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pa...

  25. [33]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684--10695, 2022

  26. [34]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22500...

  27. [35]

    Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information...

  28. [36]

    Does avatar design in educational games promote a positive emotional experience among learners? E-learning and Digital Media, 18 0 (5): 0 422--440, 2021

    Kogilathah Segaran, Ahmad Zamzuri Mohamad Ali, and Tan Wee Hoe. Does avatar design in educational games promote a positive emotional experience among learners? E-learning and Digital Media, 18 0 (5): 0 422--440, 2021

  29. [37]

    Emotion knowledge: further exploration of a prototype approach

    Phillip Shaver, Judith Schwartz, Donald Kirson, and Cary O'connor. Emotion knowledge: further exploration of a prototype approach. Journal of personality and social psychology, 52 0 (6): 0 1061, 1987

  30. [38]

    Face2diffusion for fast and editable face personalization

    Kaede Shiohara and Toshihiko Yamasaki. Face2diffusion for fast and editable face personalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6850--6859, 2024

  31. [39]

    IC face: I nterpretable and controllable face reenactment using GAN s

    Soumya Tripathy, Juho Kannala, and Esa Rahtu. IC face: I nterpretable and controllable face reenactment using GAN s. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3385--3394, 2020

  32. [40]

    Facegan: Facial attribute controllable reenactment gan

    Soumya Tripathy, Juho Kannala, and Esa Rahtu. Facegan: Facial attribute controllable reenactment gan. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021

  33. [41]

    Face0: Instantaneously conditioning a text-to-image model on a face

    Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. Face0: Instantaneously conditioning a text-to-image model on a face. In SIGGRAPH Asia 2023 Conference Papers, pages 1--10, 2023

  34. [42]

    Instantid: Zero-shot identity-preserving generation in seconds

    Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and Anthony Chen. Instantid: Zero-shot identity-preserving generation in seconds. arXiv preprint arXiv:2401.07519, 2024

  35. [43]

    Freeman, Fr\' e do Durand, and Song Han

    Guangxuan Xiao, Tianwei Yin, William T. Freeman, Fr\' e do Durand, and Song Han. Fastcomposer: Tuning-free multi-subject image generation with localized attention. Int. J. Comput. Vision, 133 0 (3): 0 1175–1194, 2024

  36. [44]

    Facestudio: Put your face everywhere in seconds, 2023

    Yuxuan Yan, Chi Zhang, Rui Wang, Pei Cheng, Gang Yu, and Bin Fu. Facestudio: Put your face everywhere in seconds, 2023

  37. [45]

    Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023 a

  38. [46]

    Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. 2023 b

  39. [47]

    Metaportrait: Identity-preserving talking head generation with fast personalized adaptation

    Bowen Zhang, Chenyang Qi, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, and Fang Wen. Metaportrait: Identity-preserving talking head generation with fast personalized adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  40. [48]

    Adding conditional control to text-to-image diffusion models, 2023 b

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023 b

  41. [49]

    Action unit driven facial expression synthesis from a single image with patch attentive GAN

    Yong Zhao, Le Yang, Ercheng Pei, Meshia Cedric Oveneke, Mitchel Alioscha-Perez, Longfei Li, Dongmei Jiang, and Hichem Sahli. Action unit driven facial expression synthesis from a single image with patch attentive GAN . In Computer Graphics Forum, pages 47--61. Wiley Online Lib...

  42. [50]

    POSTER: A pyramid cross-fusion transformer network for facial expression recognition

    Ce Zheng, Matias Mendieta, and Chen Chen. POSTER: A pyramid cross-fusion transformer network for facial expression recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (CVPR), pages 3146--3155, 2023

  43. [51]

    Storydiffusion: Consistent self-attention for long-range image and video generation

    Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou. Storydiffusion: Consistent self-attention for long-range image and video generation. Advances in Neural Information Processing Systems, 37: 0 110315--110340, 2024

  44. [52]

    4d facial expression diffusion model

    Kaifeng Zou, Sylvain Faisan, Boyang Yu, Sebastien Valette, and Hyewon Seo. 4d facial expression diffusion model. ACM Trans. Multimedia Comput. Commun. Appl., 21 0 (1), 2024

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.