Pith. sign in

REVIEW 4 major objections 5 minor 39 references

Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that perceived believability of AI child avatars is governed by audiovisual congruence: mismatched adult voices reduce realism even when facial expressions are well crafted, and anger becomes much harder to recognize…

desk verdict The anger-recognition result is solid, but the claimed realism benefit of silence is real only for sadness and is confounded by the unmatched adult-voice condition; still worth peer review. read the letter →

arxiv 2506.13477 v2 pith:YDHIA3HC submitted 2025-06-16 cs.HC cs.CV

classification cs.HCcs.CV
keywords Human-ComputerInteractionInteractiveAvatarsEmotionallyExpressivePerceivedRealismUnrealEngineMetaHumanNVIDIAAudio2Faceaudiovisualcongruencechildinterviewtraining
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that emotionally believable child avatars, of the kind needed for training interviewers who talk with children about abuse, are limited less by rendering quality than by how well voice, face, and expression cohere. The authors built a real-time pipeline that couples Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face, which derives facial animation from vocal prosody, and tested two child avatars in a between-subjects study of 70 participants who rated emotional clarity, facial realism, and empathy for joy, sadness, and anger. Their central result is an audiovisual congruence effect: when a young adult female text-to-speech voice was paired with a child-like face, muted visual-only clips were rated as more realistic than clips with audio, while anger became markedly harder to recognize without sound. A sympathetic reading of the paper is that matching a voice to a face matters more than polishing either channel alone, and that high-arousal emotions like anger need audio while low-arousal emotions like sadness can stand on facial cues.

What carries the argument

The load-bearing mechanism is the prosody-to-expression mapping performed by NVIDIA Omniverse Audio2Face, which converts spectral and prosodic features of synthesized speech into facial action unit activations and blendshape weights that drive a Unreal Engine 5 MetaHuman facial rig through the Live Link plugin; a two-computer split separates speech generation from rendering to keep the loop interactive. The argument's interpretive machinery is the valence-arousal asymmetry: audio cues are treated as the main carrier of emotional arousal, visual cues as the main carrier of valence, which is why anger (high arousal, negative valence) is the emotion most damaged by silence, while sadness and joy survive on visual cues. The paper also uses FACS-coded markers such as brow depression (AU1+4), lip-corner droop (AU15), brow tension (AU4), and Duchenne markers (AU6+12) to explain which facial signals remain readable without audio, and it uses an anchoring stimulus to calibrate participant ratings before the main clips.

What would settle it

Run the same 70-person between-subjects study again with age-matched child text-to-speech voices, using one TTS system or matching prosody more tightly; if audio+visual clips are then rated at least as realistic as visual-only clips and anger recognition in the audio condition stays high, the congruence interpretation is supported, whereas if visual-only still wins realism even with well-matched child voices, the silence advantage is caused by the audio channel itself rather than voice-age mismatch.

Watch

Extended reading notes

Core claim

The paper demonstrates a working real-time architecture for prosody-driven facial emotion on photorealistic child avatars and reports that its user study found perceived realism fell when audio was added: the presence of audio reduced the odds of higher realism ratings by about 74% ($OR=0.26$, $p=.038$). The recognition results were emotion-dependent: sadness and joy were recognized from faces alone, with sadness stable for both avatars, whereas anger recognition dropped sharply without audio (for Emory from $M=3.83$ to $M=1.43$, for Amelia from $M=3.11$ to $M=2.17$ on the five-point scale). A three-way interaction between condition, avatar, and emotion congruency ($OR=0.212$, $p<.001$) and avatar-specific effects (Emory overall harder to read, $OR=0.202$, $p=.002$) show that the value of audio depends on which avatar and which emotion is shown. The paper also finds that facial morphology shifts interpretation: Amelia's softer features supported sadness and joy, while Emory's angular features conveyed anger better when voice was present. The authors conclude that perceived believability depends on audiovisual congruence and facial geometry together, not on fidelity in any single channel.

Load-bearing premise

The study assumes that two young adult female text-to-speech voices, chosen by the researchers' judgment and supported by their earlier finding that female voices can represent child characters, are an acceptable stand-in for the child voices these avatars would need; if properly age-matched child voices had been used, the realism advantage of silence could shrink or disappear.

Editorial extensions

If this is right

  • For low-arousal emotions such as sadness and joy, a visual-only presentation can be rated more realistic than an audio-visual one whenever the voice is not fully congruent with the avatar's face, because silencing the clip removes a source of dissonance.
  • High-arousal emotions such as anger cannot be reliably recognized from facial animation alone; training applications that need anger detection must supply synchronized vocal cues or accept a large recognition drop.
  • Avatar facial morphology is not neutral: softer, rounder features aid low-arousal believability, while angular features make anger more salient, so designers should select geometry to match the emotional demands of the training scenario.
  • Voice-age congruence is a first-order design constraint for child avatars; the adult-voice compromise used here is identified as a confound that likely reduced the audio condition's realism.
  • The distributed two-PC pipeline shows real-time prosody-driven emotional animation is technically feasible, but the remaining realism bottleneck is low-level synchronization and prosody-to-expression alignment rather than raw rendering quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the congruence interpretation extends, an adaptive design would use audio only for high-arousal moments or switch to a well-matched child voice, making silence no longer preferable; this is a natural next test the paper does not run.
  • A direct follow-up experiment varying voice age, TTS system, and emotion (including fear and disgust) in a factorial design would separate voice-age mismatch from general TTS prosody mismatch; if silence still beats audio with age-matched child voices, the effect is about audio itself, not age.
  • The non-significant empathy difference, combined with higher visual-only realism, suggests interview training could deliberately include a silent replay mode so trainees practice reading facial cues without vocal interference, a design option the paper leaves implicit.
  • The qualitative comments about faces looking 'scary' or 'rigid' point to a possible connection between audiovisual mismatch and the uncanny valley, implying that congruence failures, not just fidelity failures, may trigger discomfort in sensitive contexts.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a real-time architecture that combines Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to produce prosody-driven facial expressions on photorealistic child avatars, and reports a between-subjects user study (N=70) comparing audio+visual with visual-only presentation of joy, sadness, and anger. The central empirical claim is that perceived believability depends on audiovisual congruence: a mismatched adult female voice undermines even well-crafted facial expressions, silencing audio improved perceived realism, and anger was markedly harder to recognize without audio. The paper also reports avatar-morphology effects and qualitative feedback about lip-sync and uncanny-valley issues.

Significance. If the claims held as stated, the paper would offer useful design guidance for emotionally expressive avatars in sensitive training contexts, and the public code release is a strength. The anger-recognition drop in the visual-only condition is visible in the descriptive statistics and partly corroborated by t-tests, and the authors are transparent about the voice-age mismatch in the abstract, Section 3.2, and Section 5.1. However, the manuscript's headline generalization—that visual-only presentations are more realistic because audio mismatches undermine believability—is not supported by the models as reported: the main effect of condition on facial-expression realism is nonsignificant in Table 6, and the significant realism result in Table 7 compares conditions on different questionnaire item sets. The paper is therefore a promising empirical contribution whose interpretive claims need substantial reframing or additional analysis.

major comments (4)
  1. [Section 3.2 and Section 5.1] The only audio condition in this study uses young adult female TTS voices for both child avatars, a mismatch the authors acknowledge in the abstract and in Section 5.1. Because no age-matched child-voice condition is included, the observed realism advantage of visual-only clips cannot distinguish among three explanations: (a) a mismatched voice specifically reduces realism, (b) any audio reduces realism because of lip-sync or animation artifacts, or (c) the adult-voice-on-child-avatar mismatch drives the effect. The abstract and conclusion state that 'perceived believability hinges on audiovisual congruence,' but the design can only support a narrower claim about the tested adult-voice conditions. This is a load-bearing limitation for the paper's central generalization.
  2. [Section 4.1.4, Table 7] The realism model in Table 7 compares audio+visual and visual-only conditions using different item sets: voice tone and dialogue content were rated only in the audio+visual condition, while the visual-only condition rated only facial expressions and visual appearance. The significant OR of 0.26 (p=.038) could therefore reflect the composition of the aggregated realism score rather than the effect of audio per se. The analysis should be rerun on the common items (facial expressions and visual appearance) only, or modeled at the item level with item type as a factor, to support the claim that audio reduces perceived realism.
  3. [Section 4.1.3, Tables 6 and 8] The claim that 'silencing the clips improved perceived realism' overstates the evidence. In Table 6, the main effect of visual-only on facial-expression realism is not significant (p=.526); the only significant realism-related effect is the interaction with sadness (p=.023, OR=4.554). The supplementary t-tests in Table 8 show significant FDR-corrected facial-expression realism differences only for the two sad scenarios (Emory-Sad d=0.68, p=.021; Amelia-Sad d=0.67, p=.021), while the angry, joy, and other scenarios are not significant after FDR correction. The conclusion should be reframed as an emotion-specific effect, not a general realism benefit of removing audio.
  4. [Section 4.2.1 and Section 5] Qualitative comments about lip-sync and desynchronization make hypothesis (b) above—that any audio, regardless of voice age—is plausible. The manuscript itself notes that '[i]ssues with lip synchronization were also noted as detrimental to realism' and that 'the inclusion of voice amplified scrutiny of temporal and spatial alignment.' This strengthens the need for a matched-voice condition or at least for conclusions that explicitly acknowledge that the observed effects may be driven by synchronization artifacts rather than by voice-age congruence specifically.
minor comments (5)
  1. [Section 4.1.1, Table 3] The model-comparison table for facial realism contains internally inconsistent values: for example, the row with logLik=-380.78 and AIC=939.57 implies a parameter count that is not reported and is implausibly large, and the listed LRT values do not match the differences in logLik between adjacent rows. Please reconcile the reported fit statistics.
  2. [Section 4.1.1, Table 2] The emotion-recognition model comparison reports 'no.par' and a single LRT value, but it is unclear which two models are being compared. Please specify the null and alternative models and report the degrees of freedom of the chi-square statistic.
  3. [Figure 5] The figure contains a text-encoding artifact in the legend label ('Visual/uni00ADOnly'); please fix this. Also, the figure caption says 'color-coded bars distinguishing the experimental conditions,' but the printed version relies on gray shading; please ensure the distinction is accessible in grayscale.
  4. [Section 4.1.4, paragraph 2] The phrase 'a significant 74% reduction in the likelihood of higher realism ratings' is imprecise: OR=0.26 means the odds of a higher rating are reduced by approximately 74%, not the likelihood. Please use odds-based wording consistently.
  5. [Section 3.3] The ethics statement says the study 'did not require formal ethical approval,' while also noting that some participants reported discomfort and that no content warnings were provided. For a journal submission, please clarify the institutional ethics framework under which this determination was made, and consider reporting whether any debriefing was provided.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical user study whose outcomes are measured, not derived from its inputs; the one self-citation is design justification, not load-bearing.

full rationale

This paper is an empirical user study rather than a derivation chain, so the equation-level circularity patterns do not apply. The central claims (anger recognition drops without audio, visual-only clips are rated more realistic, audiovisual congruence matters) are supported by newly collected participant ratings and mixed-effects models (e.g., Table 7 OR=0.26, p=.038; Table 8 anger drops), not by re-reporting fitted parameters as predictions. The only self-citation, reference [34], is used in Section 3.2 to justify selecting young adult female TTS voices for child avatars; it informs stimulus design but does not define the outcome measures or the reported statistical contrasts, and the paper explicitly labels the voice-age mismatch as a confound in the abstract and Section 5.1. The absence of a matched-age child-voice condition is a genuine threat to external validity and weakens the generality of the audiovisual-congruence conclusion, but that is a confound, not a circular reduction. No self-definitional step, no fitted-input-called-prediction, and no load-bearing self-citation chain are present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

This is an empirical systems paper, so there are no fitted derivation constants. The main hand-chosen design settings are the exaggerated Audio2Face mode and the TTS voice selection. The central claim rests on standard statistical modeling assumptions and on the validity of the emotional stimuli. No new theoretical entities are introduced.

free parameters (1)
  • Audio2Face emotion intensity mode = exaggerated
    The 'exaggerated' mode was chosen by hand after preliminary observations; it affects all stimulus clips and could influence perceived realism and recognition, but it was not fitted to the participant data.
assumptions (4)
  • domain assumption Cumulative link mixed-effects models with random intercepts are appropriate for the five-point Likert ratings.
    Statistical conclusions rely on the validity of the ordinal model and the proportional-odds relaxation described in Section 4.1.1.
  • domain assumption The low-fidelity anchoring stimulus calibrates participants' realism ratings as intended.
    The study inserts a low-fidelity anchor before main stimuli, citing anchoring literature [35]; if the anchor failed, the magnitude of realism ratings could shift.
  • domain assumption Audio2Face's FACS-derived blendshape expressions in exaggerated mode validly represent joy, sadness, and anger for child avatars.
    The stimuli are assumed to be recognizable instances of the target emotions; the paper does not independently validate the animation against FACS labeling.
  • domain assumption The use of young adult female voices for child avatars is an acceptable stand-in, based on the authors' prior study [34].
    This justifies the voice-age mismatch design; if the prior result is not transferable, the mismatch may be larger than intended and the results may not generalize.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications." pith.science (2026). https://pith.science/paper/YDHIA3HC

@misc{pith2026250613477,
  author       = {Pith},
  title        = {Pith review of: Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YDHIA3HC}},
  note         = {Machine review of arXiv:2506.13477}
}
read the original abstract

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a real-time architecture combining Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to generate facial expressions from vocal prosody in photorealistic child avatars. Due to limited TTS options, both avatars were voiced using young adult female models from two systems to better fit character profiles, introducing a voice-age mismatch. This confound may affect audiovisual alignment. We used a two-PC setup to decouple speech generation from GPU-intensive rendering, enabling low-latency interaction in desktop and VR. A between-subjects study (N=70) compared audio+visual vs. visual-only conditions as participants rated emotional clarity, facial realism, and empathy for avatars expressing joy, sadness, and anger. While emotions were generally recognized - especially sadness and joy - anger was harder to detect without audio, highlighting the role of voice in high-arousal expressions. Interestingly, silencing clips improved perceived realism by removing mismatches between voice and animation, especially when tone or age felt incongruent. These results emphasize the importance of audiovisual congruence: mismatched voice undermines expression, while a good match can enhance weaker visuals - posing challenges for emotionally coherent avatars in sensitive contexts.

Figures

Figures reproduced from arXiv: 2506.13477 by the authors.

Figure 1
Figure 1. Example frames of the two avatars used in the user study: "Emory" (male-presenting) and "Amelia" (female [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Architecture for real-time emotionally expressive child avatars. PC1 processes speech input (Google Cloud [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Scripted utterances used for emotional avatar expressions in the user study. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Low-fidelity anchoring stimulus used at the beginning of the user study to calibrate participants’ expectations [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Mean ratings (M ± SD) of emotional expressiveness across five emotions (Angry, Sad, Joy, Fear, Disgust) for the avatars Emory and Amelia, under both audio+visual and visual-only conditions. Bars represent group means; values next to each bar report the mean and standar…
Figure 6
Figure 6. Figure 6: Forest plot showing OR and 95% CI for predictors of emotion recognition accuracy. The dashed vertical line [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Mean ratings (M ± SD) perceived Facial Expressions realism for each video clips under both audio+visual and visual-only conditions [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Forest plot showing OR with 95% CI for the model predictors of perceived realism of facial expressions. The [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Mean ratings (M ± SD) of perceived realism across different aspects (facial expressions, visual appearance, voice tone, dialogue content) and empathy, compared between audio+visual and visual-only conditions [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

39 extracted references · 37 canonical work pages

  1. [1]

    Supervising 3d talking head avatars with analysis-by- audio-synthesis,

    R. Danˇeˇcek, C. Schmitt, S. Polikovsky, and M. J. Black, “Supervising 3d talking head avatars with analysis-by- audio-synthesis,” arXiv preprint arXiv:2504.13386, 2025

  2. [2]

    Affective computing: Recent advances, challenges, and future trends,

    G. Pei, H. Li, Y . Lu, Y . Wang, S. Hua, and T. Li, “Affective computing: Recent advances, challenges, and future trends,” Intelligent Computing, vol. 3, p. 0076, 2024

  3. [3]

    Emotionally_expressive_child_avatars,

    P. Salehi, “Emotionally_expressive_child_avatars,” https://github.com/pegahsalehi/Emotionally_Expressive_ Child_Avatars, 2025, gitHub repository, accessed July 7, 2025

  4. [4]

    A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills,

    P. Salehi, S. Z. Hassan, G. A. Baugerud, M. Powell, M. S. Johnson, D. Johansen, S. S. Sabet, M. A. Riegler, and P. Halvorsen, “A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills,” IEEE Access, 2024

  5. [5]

    Research on the application of digital human production based on photoscan realistic head 3d scanning and unreal engine metahuman technology in the metaverse,

    Y . Pan, K. Kim, J. Lee, Y . Sang, and J. Cheon, “Research on the application of digital human production based on photoscan realistic head 3d scanning and unreal engine metahuman technology in the metaverse,” International journal of advanced smart convergence, vol. 11, no. 3, pp. 102–118, 2022

  6. [6]

    NVIDIA Audio2Face 3D,

    NVIDIA, “NVIDIA Audio2Face 3D,” https://build.nvidia.com/nvidia/audio2face-3d, 2025, accessed: 2025-04-04

  7. [7]

    Adapting software with affective computing: a systematic review,

    R. V . Aranha, C. G. Corrêa, and F. L. Nunes, “Adapting software with affective computing: a systematic review,” IEEE Transactions on Affective Computing, vol. 12, no. 4, pp. 883–899, 2019

  8. [8]

    Learning word vectors for sentiment analysis,

    A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y . Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, 2011, pp. 142–150

Show all 39 references
  1. [9]

    Recognizing action units for facial expression analysis,

    Y .-I. Tian, T. Kanade, and J. F. Cohn, “Recognizing action units for facial expression analysis,”IEEE Transactions on pattern analysis and machine intelligence, vol. 23, no. 2, pp. 97–115, 2001. 17 A PREPRINT - S EPTEMBER 22, 2025

  2. [10]

    Empirical evidence relating eeg signal duration to emotion classification performance,

    E. T. Pereira, H. M. Gomes, L. R. Veloso, and M. R. A. Mota, “Empirical evidence relating eeg signal duration to emotion classification performance,” IEEE Transactions on Affective Computing, vol. 12, no. 1, pp. 154–164, 2018

  3. [11]

    Designing emotionally sentient agents,

    D. McDuff and M. Czerwinski, “Designing emotionally sentient agents,” Communications of the ACM, vol. 61, no. 12, pp. 74–83, 2018

  4. [12]

    Implementation of a chatbot system using ai and nlp,

    T. Lalwani, S. Bhalotia, A. Pal, V . Rathod, and S. Bisen, “Implementation of a chatbot system using ai and nlp,” International Journal of Innovative Research in Computer Science & Technology (IJIRCST) Volume-6, Issue-3, 2018

  5. [13]

    Machines and mindlessness: Social responses to computers,

    C. Nass and Y . Moon, “Machines and mindlessness: Social responses to computers,”Journal of social issues, vol. 56, no. 1, pp. 81–103, 2000

  6. [14]

    The human face as a dynamic tool for social communication,

    R. E. Jack and P. G. Schyns, “The human face as a dynamic tool for social communication,” Current Biology, vol. 25, no. 14, pp. R621–R634, 2015

  7. [15]

    What is emotion?

    M. Cabanac, “What is emotion?” Behavioural processes, vol. 60, no. 2, pp. 69–83, 2002

  8. [16]

    Cognitive, social, and physiological determinants of emotional state

    S. Schachter and J. Singer, “Cognitive, social, and physiological determinants of emotional state.”Psychological review, vol. 69, no. 5, p. 379, 1962

  9. [17]

    Evaluating the modeling and use of emotion in virtual humans,

    J. Gratch and S. Marsella, “Evaluating the modeling and use of emotion in virtual humans,” in Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems, 2004. AAMAS 2004., vol. 1. IEEE Computer Society, 2004, pp. 320–327

  10. [18]

    Lessons from emotion psychology for the design of lifelike characters,

    ——, “Lessons from emotion psychology for the design of lifelike characters,” Applied Artificial Intelligence, vol. 19, no. 3-4, pp. 215–233, 2005

  11. [19]

    Audio-driven emotional video portraits,

    X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14 080–14 089

  12. [20]

    Neural emotion director: Speech-preserving semantic control of facial expressions in

    F. P. Papantoniou, P. P. Filntisis, P. Maragos, and A. Roussos, “Neural emotion director: Speech-preserving semantic control of facial expressions in" in-the-wild" videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18 781–18 790

  13. [21]

    Neural voice puppetry: Audio-driven facial reenactment,

    J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16. Springer, 2020, pp. 716–731

  14. [22]

    Learning a model of facial shape and expression from 4d scans

    T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans.” ACM Trans. Graph., vol. 36, no. 6, pp. 194–1, 2017

  15. [23]

    Emotional speech-driven animation with content-emotion disentanglement,

    R. Danˇeˇcek, K. Chhatre, S. Tripathi, Y . Wen, M. Black, and T. Bolkart, “Emotional speech-driven animation with content-emotion disentanglement,” in SIGGRAPH Asia 2023 Conference Papers, 2023, pp. 1–13

  16. [24]

    Emotalk: Speech-driven emotional disentanglement for 3d face animation,

    Z. Peng, H. Wu, Z. Song, H. Xu, X. Zhu, J. He, H. Liu, and Z. Fan, “Emotalk: Speech-driven emotional disentanglement for 3d face animation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 20 687–20 697

  17. [25]

    Metahuman,

    Epic Games, “Metahuman,” https://www.unrealengine.com/en-US/metahuman, 2025, accessed 30 April 2025

  18. [26]

    Rig logic white paper v2,

    ——, “Rig logic white paper v2,” https://cdn2.unrealengine.com/rig-logic-whitepaper-v2-5c9f23f7e210.pdf, 2021, accessed 30 April 2025

  19. [27]

    Using subsurface scattering in unreal engine materials,

    Epic Games Documentation Team, “Using subsurface scattering in unreal engine materials,” https://dev.epicgames. com/documentation/en-us/unreal-engine/using-subsurface-scattering-in-unreal-engine-materials, 2025, accessed 30 April 2025

  20. [28]

    Google cloud speech-to-text,

    G. Cloud, “Google cloud speech-to-text,” https://cloud.google.com/speech-to-text, 2023, accessed 5 May 2025

  21. [29]

    Chatgpt,

    OpenAI, “Chatgpt,” https://chat.openai.com, 2024, accessed 5 May 2025

  22. [30]

    Amazon polly neural voices,

    A. W. Services, “Amazon polly neural voices,” https://docs.aws.amazon.com/polly/latest/dg/neural-voices.html, 2025, accessed 5 May 2025

  23. [31]

    Nvidia audio2face: Revolutionizing facial animation in real-time,

    S. Daly, “Nvidia audio2face: Revolutionizing facial animation in real-time,” https://9meters.com/technology/ai/ nvidia-audio2face-revolutionizing-facial-animation-in-real-time, 2024, accessed 5 May 2025

  24. [32]

    Audio2face live link ue plugin — user manual,

    N. Corporation, “Audio2face live link ue plugin — user manual,” https://docs.omniverse.nvidia.com/audio2face/ latest/user-manual/livelink-ue-plugin.html, 2025, accessed 5 May 2025

  25. [33]

    Multimodal affect models: An investigation of relative salience of audio and visual cues for emotion prediction,

    J. Wu, T. Dang, V . Sethu, and E. Ambikairajah, “Multimodal affect models: An investigation of relative salience of audio and visual cues for emotion prediction,” Frontiers in Computer Science, vol. 3, p. 767767, 2021. 18 A PREPRINT - S EPTEMBER 22, 2025

  26. [34]

    Is more realistic better? a comparison of game engine and gan-based avatars for investigative interviews of children,

    P. Salehi, S. Z. Hassan, S. Shafiee Sabet, G. Astrid Baugerud, M. Sinkerud Johnson, P. Halvorsen, and M. A. Riegler, “Is more realistic better? a comparison of game engine and gan-based avatars for investigative interviews of children,” in Proceedings of the 3rd ACM Workshop o...

  27. [35]

    What stimuli are necessary for anchoring effects to occur?

    Y . Onuki, H. Honda, and K. Ueda, “What stimuli are necessary for anchoring effects to occur?” Frontiers in Psychology, vol. 12, p. 602372, 2021

  28. [36]

    Facial action coding system,

    P. Ekman and W. V . Friesen, “Facial action coding system,”Environmental Psychology & Nonverbal Behavior, 1978

  29. [37]

    An fmri study of affective congruence across visual and auditory modalities,

    C. Gao, C. E. Weber, D. H. Wedell, and S. V . Shinkareva, “An fmri study of affective congruence across visual and auditory modalities,” Journal of cognitive neuroscience, vol. 32, no. 7, pp. 1251–1262, 2020

  30. [38]

    Cross-modal interaction between auditory and visual input impacts memory retrieval,

    V . Marian, S. Hayakawa, and S. R. Schroeder, “Cross-modal interaction between auditory and visual input impacts memory retrieval,” Frontiers in Neuroscience, vol. 15, p. 661477, 2021

  31. [39]

    How is believability of a virtual agent related to warmth, competence, personification, and embodiment?

    V . Demeure, R. Niewiadomski, and C. Pelachaud, “How is believability of a virtual agent related to warmth, competence, personification, and embodiment?” Presence, vol. 20, no. 5, pp. 431–448, 2011. Appendix: Supplementary Material This appendix includes supplementary statisti...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.