Pith. sign in

REVIEW 3 major objections 3 minor 85 references

A multidimensional measurement of photorealistic avatar quality of experience

T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper argues that subjective human ratings are the only reliable way to measure photorealistic avatar quality of experience: standard objective metrics correlate weakly with nine of ten measured dimensions and only moderately with…

desk verdict Solid open-source avatar QoE framework with a credible negative result on objective metrics, but the no-uncanny-valley claim exceeds what the data can support. read the letter →

arxiv 2411.09066 v3 pith:XEYRNSWY submitted 2024-11-13 cs.HC cs.CVcs.GR

classification cs.HCcs.CVcs.GR
keywords photorealisticavatarsqualityofexperiencesubjectivetestingcrowdsourcinguncannyvalleyobjectivemetricstelecommunicationavatarrealism
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Photorealistic avatars are usually developed and compared using pixel- and feature-based metrics, but this paper argues those metrics cannot tell you whether an avatar can be trusted, feels comfortable to work with, or creeps people out. It contributes an open-source, crowdsourced test framework that measures avatar quality of experience along ten human-centered dimensions, and shows the framework is accurate against a panel of experts and reproducible across repeated runs. Applied to 19 avatar models spanning photorealistic to cartoon-like, the measurements reveal that PSNR, SSIM, LPIPS, FID, and FVD correlate weakly with nine of the ten dimensions and only moderately with emotion accuracy. The paper also reports that for avatars above a realism threshold, eight dimensions move together, so the survey can be shortened from ten items to three. Finally, it finds affinity grows linearly with realism and concludes that for photorealistic avatars in telecommunication there is no uncanny valley; lower-realism avatars are simply worse on trust, comfort, appropriateness, formality, and creepiness than real video.

What carries the argument

The load-bearing mechanism is the open-source crowdsourced subjective test framework, an implementation of the P.910 subjective video quality assessment standard extended to avatars, with rater qualification, display calibration, gold clips, trapping items, and repeated items to control data quality. It uses two survey templates: Template A asks for agreement with eight statements about the avatar alone, and Template B shows the avatar side-by-side with the real person to rate resemblance and emotion accuracy. The framework's output is a per-clip and per-condition mean opinion score for each dimension, and its validation against laboratory video studies and an expert panel is what gives the correlations their weight. The analytic load is carried by inter-dimension Pearson correlations, a principal component analysis that finds two components explaining 94% of the variance, and a linear fit of affinity against realism whose R²=0.966 is the basis for the no-uncanny-valley conclusion.

What would settle it

Have separate groups of raters each judge only one dimension of the same avatar clips, then compare those single-dimension scores with the all-items-at-once scores. If trust, comfort, and appropriateness no longer track realism when realism is not among the questions asked — or if the average inter-item correlation drops well below 0.996 — then the reported high correlations are a halo or order effect, and the paper's dimensionality reduction and no-uncanny-valley linearity would not survive independent measurement.

Watch

Extended reading notes

Core claim

The central discovery is that the quality of experience of a photorealistic avatar is a multidimensional subjective quantity that standard objective metrics do not capture. Using a validated crowdsourced survey, the paper measures ten dimensions — realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness (reported as not creepy), formality, affinity, resemblance to the person, and emotion accuracy — across 19 avatar models. The inter-item correlations among the eight Template A dimensions average PCC=0.996, and for avatars with realism above 2 on the 1–5 scale they average 0.999, which the paper interprets as a strong redundancy that permits reducing the ten dimensions to three via principal component analysis. Against PSNR, SSIM, LPIPS, FID, and FVD, the subjective dimensions show weak correlation, with only emotion accuracy reaching moderate correlation (e.g., PSNR=0.63, LPIPS=-0.67). In addition, affinity versus realism is well described by a straight line (R²=0.966), which the paper takes as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.

Load-bearing premise

The ten survey items are treated as independently meaningful dimensions, but the near-perfect inter-item correlations (average PCC=0.996) are exactly what a single global impression would look like, so if raters are not actually distinguishing the dimensions, the dimensionality reduction and the linear affinity–realism relation would be artifacts of rating co-movement rather than evidence about the underlying quality structure.

Editorial extensions

If this is right

  • Avatar developers who optimize only PSNR, SSIM, LPIPS, FID, or FVD will leave most of the quality-of-experience space unmeasured, so subjective testing becomes necessary for usability claims.
  • For avatars with realism above 2 on the 1–5 scale, one Template A item (e.g., realism) plus the two Template B items (emotion accuracy and resemblance) is enough, shrinking the survey from ten items to three.
  • Avatars that are less realistic than real video will score lower on trust, comfort using, comfort interacting, work appropriateness, formality, and affinity, and higher on creepiness; matching real video requires matching its realism.
  • Because affinity rises linearly with realism (R²=0.966), designers of telecommunication avatars should maximize realism rather than stop short of a presumed uncanny valley.
  • Some state-of-the-art photorealistic avatars approach real video on these subjective dimensions, but none reaches real-video realism, so the gap between avatars and real video remains measurable and large.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: if the near-perfect inter-item correlations are a halo effect, single-dimension ratings from independent groups will diverge from the all-items survey; running that comparison is a direct check on whether the ten dimensions are really separate constructs.
  • The no-uncanny-valley conclusion is tied to the stimulus set used here — 2D talking-head clips in telecommunication poses; the same survey applied to full-body avatars, VR/AR displays, or avatars with motion artifacts could still find a dip the paper does not test.
  • Because emotion accuracy and resemblance show moderate-to-strong correlation with PSNR, SSIM, and LPIPS when computed on the face region, a learned perceptual metric trained on subjective avatar ratings could plausibly close much of the objective-subjective gap, a direction the paper mentions as future work.
  • The framework is a passive-viewing test; an interactive variant in which a rater converses with a live avatar could change trust and comfort ratings, since the paper's own reproducibility and accuracy claims rest on passive observation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents an open-source crowdsourced test framework, based on the authors' P.910 implementation, for measuring ten subjective dimensions of photorealistic avatar quality of experience: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance, and emotion accuracy. The framework is validated against VQEG laboratory video-quality data and against a panel of experts, and its run-to-run reproducibility is reported. Using 19 avatar models and real-video baselines, the authors find that the subjective dimensions are weakly correlated with common objective metrics (PSNR, SSIM, LPIPS, FID, FVD), that eight Template A dimensions become nearly collinear for avatars with realism above a threshold, and that affinity is linearly related to realism, which they interpret as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.

Significance. If the claims hold, this is a useful contribution: an open, validated, reproducible subjective testing methodology for avatar quality of experience, a public test set, and evidence that objective metrics currently used in avatar development do not capture the usability dimensions that matter. The framework validation is genuinely strong: run-to-run clip-level correlations are high for Template A (Table 8, average PCC 0.962), model-level reproducibility is above 0.98 (Table 7), accuracy against an expert panel averages PCC 0.904 (Table 6), and the crowdsourcing implementation is validated against external VQEG lab data (Tables 2-4). The main risk is that the two most prominent perceptual conclusions, the dimensionality reduction and the absence of an uncanny valley, rest on near-perfect inter-item correlations that may reflect a general impression halo rather than distinct underlying constructs, and the no-uncanny-valley conclusion is drawn from a linear fit on clustered clip-level points without a direct test for a local dip.

major comments (3)
  1. [§5.5, Fig. 10] I would ask the authors to add an avatar-random-effect analysis at the model level, a direct test for a local dip in the near-photorealistic range, confidence intervals on the fit, or to substantially soften the conclusion and restrict it to the range actually sampled.
  2. [§5.3, Tables 9-11] Since the realism > 2 threshold is chosen post hoc and the claim directly feeds the developer guideline that 'only 3 dimensions need to be measured' (§6.2), the paper should provide stronger evidence of discriminant validity, for example a confirmatory factor analysis with a specified measurement model, or should reframe the claim as a practical shortcut rather than a statement about the underlying psychological dimensions.
  3. [§5.4, Tables 12-13] Please report confidence intervals or bootstrap intervals for Table 12, and adjust the abstract and conclusions to reflect the face-limited results.
minor comments (3)
  1. [§4.3] The text says the average clip-level run-run PCC is 0.968, but Table 8 gives Template A averages of 0.962 and a Template B emotion-accuracy average of 0.870; the text should match the table, and the lower emotion-accuracy reproducibility should be acknowledged as a caveat to the 'highly reproducible' claim.
  2. [§6] The first sentence lists 'PSNR, SSIM, LPIPS, FIV, FVD'; this should be FID, not FIV.
  3. [§3.1] The open-source tool is named P.910, identical to the ITU-T Recommendation P.910. This naming is likely to confuse readers; consider renaming the repository or adding a clarifying sentence in Section 3.1.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the framework's ratings are primary data and are validated against external VQEG lab data and an expert panel; the authors' self-citation of their own P.910 tool is not load-bearing.

full rationale

The central claims are built on new crowdsourced subjective ratings of avatar clips, not on fitted parameters or on prior results by the same authors. The P.910 crowdsourcing implementation is the authors' own tool, but its validity is established against external VQEG HDTV laboratory data (Section 3.1.9, PCC=0.952 per sequence) and the avatar extension is compared to an independent expert panel (Table 6, average PCC=0.904). Reproducibility is demonstrated across five mutually exclusive MTurk runs (Tables 7-8). The weak objective-metric correlations (Table 12) are computed from the collected MOS and standard objective metrics, so they are empirical findings rather than constructed equivalences. The most conceptually vulnerable claim is the no-uncanny-valley conclusion (Section 5.5), because the near-perfect inter-item correlations (Table 9, average PCC=0.996) suggest a strong halo effect; a linear affinity-realism fit (R2=0.966) may partly reflect rating co-movement rather than independent psychological dimensions. However, the paper openly reports these correlations and the limitations that the test is passive and 2D, and the linearity claim is an empirical inference, not a reduction of the conclusion to its inputs by definition. No step in the derivation chain equates a predicted quantity to a fitted parameter or imports a uniqueness theorem from the authors' prior work. Hence the circularity score is low.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No invented entities. The central empirical claims rest on three fitted quantities: the affinity-realism regression (slope and intercept, not reported) and the post hoc realism>2 split. The axioms are domain assumptions about passive 2D crowdsourced ratings, expert ground truth, construct independence of the ten scales, and the validity of PCA and objective metric preprocessing.

free parameters (3)
  • Affinity-realism linear fit slope = not reported
    Section 5.5 and Figure 10: the no-uncanny-valley conclusion is based on this regression (R2=0.966), but the slope is not reported.
  • Affinity-realism linear fit intercept = not reported
    Section 5.5 and Figure 10: the intercept of the same regression is fitted to the data and not reported.
  • Realism threshold for subgroup analysis = 2 on a 1-5 Likert scale
    Section 5.2: the split realism>2 is applied post hoc; it determines the claim that eight dimensions are strongly correlated and motivates the 10-to-3 dimensionality reduction.
assumptions (5)
  • domain assumption Passive viewing of short 2D video clips by crowdsourced raters is a valid proxy for avatar quality of experience in telecommunication scenarios.
    The survey shows clips and collects ratings without live interaction (Section 3.2); Section 6.3 lists 'passive test' and '2D displays' as limitations, so the framework's ecological validity for real telecom use is assumed.
  • domain assumption The 5-expert panel provides an accurate ground truth for the subjective dimensions.
    Section 4.4 validates crowdsourced ratings against a panel of five experts; expert consensus is itself a small-sample subjective judgment.
  • domain assumption The ten survey items measure distinct psychological constructs rather than a single overall impression.
    Section 3.2 presents ten separate items, but Table 9 shows average inter-item PCC=0.996, which is also what a halo effect would produce; the paper does not control for this.
  • standard math PCA with two retained components (94% variance) is a valid linear reduction of the ten items.
    Section 5.3 applies PCA and adopts a two-component model; this assumes the latent structure is linear and that 94% variance is sufficient.
  • domain assumption Objective metrics computed on background-removed, face-aligned original videos are a fair basis for comparing avatar output to the source.
    Section 5.4 transforms original videos using face landmarks and removes the background before computing PSNR/SSIM/LPIPS/FID/FVD; this preprocessing could change metric values and is assumed not to bias the comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A multidimensional measurement of photorealistic avatar quality of experience." pith.science (2026). https://pith.science/paper/XEYRNSWY

@misc{pith2026241109066,
  author       = {Pith},
  title        = {Pith review of: A multidimensional measurement of photorealistic avatar quality of experience},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XEYRNSWY}},
  note         = {Machine review of arXiv:2411.09066}
}
read the original abstract

Photorealistic avatars are human avatars that look, move, and talk like real people. The performance of photorealistic avatars has significantly improved recently based on objective metrics such as PSNR, SSIM, LPIPS, FID, and FVD. However, recent photorealistic avatar publications do not provide subjective tests of the avatars to measure human usability factors. We provide an open source test framework to subjectively measure photorealistic avatar performance in ten dimensions: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance to the person, and emotion accuracy. Using telecommunication scenarios, we show that the correlation of nine of these subjective metrics with PSNR, SSIM, LPIPS, FID, and FVD is weak, and moderate for emotion accuracy. The crowdsourced subjective test framework is highly reproducible and accurate when compared to a panel of experts. We analyze a wide range of avatars from photorealistic to cartoon-like and show that some photorealistic avatars are approaching real video performance based on these dimensions. We also find that for avatars above a certain level of realism, eight of these measured dimensions are strongly correlated. This means that avatars that are not as realistic as real video will have lower trust, comfortableness using, comfortableness interacting with, appropriateness for work, formality, and affinity, and higher creepiness compared to real video. In addition, because there is a strong linear relationship between avatar affinity and realism, there is no uncanny valley effect for photorealistic avatars in the telecommunication scenario. We suggest several extensions of this test framework for future work and discuss design implications for telecommunication systems. The test framework is available at https://github.com/microsoft/P.910.

Figures

Figures reproduced from arXiv: 2411.09066 by the authors.

Figure 1
Figure 1. Data Flow Diagram [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. The crowdsourcing test from the participant’s perspective. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Items in the survey. (a) Template A as represented in the survey including a trapping and repeated [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Avatars used for the survey each avatar, the number viewpoints, whether speech is included, and if there was a reference video driving the avatar. This test set should help answer RQ1, as we can compute objective and subjective metrics on it and compute their correlati…
Figure 5
Figure 5. Figure 5: (a) MOS scores per dimension across all models from Template A and (b) Template B. [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: Template A MOS scores per dimension across all models irrespective of the view angle. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: MOS scores per dimension across all models - No audio in clips [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Standard deviation across all dimensions in Template A versus realism. A best fit line is shown with a [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Modifying the original video to match the position of the avatar using face landmarks and removing the background before calculating objective metrics. (a) A frame from the original video, (b) the corresponding frame from the avatar, (c) an overlay of the modified orig…
Figure 10
Figure 10. Figure 10: Affinity versus realism for the video clips in Figure [PITH_FULL_IMAGE:figures/full_fig_p023_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 46 canonical work pages

  1. [1]

    Bailenson

    Jeremy N. Bailenson. 2021. Nonverbal overload: A theoretical argument for the causes of Zoom fatigue. Technology, Mind, and Behavior 2, 1 (Feb. 2021). https://doi.org/10.1037/tmb0000030

  2. [2]

    Biehl, Daniel Avrahami, and Anthony Dunnigan

    Jacob T. Biehl, Daniel Avrahami, and Anthony Dunnigan. 2015. Not Really There: Understanding Embodied Communication Affordances in Team Perception and Participation. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing . ACM, Vancouver BC Canada, 1567–1575. https: //doi.org/10.1145/2675133.2675220

  3. [3]

    Ryan Canales, Doug Roble, and Michael Neff. 2024. The Impact of Avatar Stylization on Trust. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR) . IEEE, Orlando, FL, USA, 418–428. https://doi.org/10.1109/VR58804.2024. 00063

  4. [4]

    Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, and Jason Saragih. 2022. Authentic volumetric avatars from a phone scan. ACM Transactions on Graphics 41, 4 (July 2022), 163:1–163:19. https://doi.org/10.1145/3528223.3530143

  5. [5]

    Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Bo Qiu, Stefano Zanetti, Ioannis Katsavounidis, and Alan C. Bovik. 2024. Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality. http://arxiv.org/abs/2408.07041 arXiv:2408.07041 [eess]

  6. [6]

    JH Clark. 1924. The Ishihara test for color blindness. American Journal of Physiological Optics (1924)

  7. [7]

    Christophe Croux and Catherine Dehon. 2010. Influence functions of the Spearman and Kendall correlation measures. Statistical methods & applications 19 (2010), 497–515

  8. [8]

    Yu Deng, Duomin Wang, Xiaohang Ren, Xingyu Chen, and Baoyuan Wang. 2024. Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data. In CVPR

Show all 85 references
  1. [9]

    Macdorman

    Alexander Diel, Sarah Weigelt, and Karl F. Macdorman. 2022. A Meta-analysis of the Uncanny Valley’s Independent and Dependent Variables. ACM Transactions on Human-Robot Interaction 11, 1 (March 2022), 1–33. https://doi.org/10. 1145/3470742

  2. [10]

    Geraldine Fauville, Mufan Luo, Anna C. M. Queiroz, Jeremy N. Bailenson, and Jeff Hancock. 2021. Nonverbal Mechanisms Predict Zoom Fatigue and Explain Why Women Experience Higher Levels than Men. SSRN Electronic Journal (2021). https://doi.org/10.2139/ssrn.3820035

  3. [11]

    Aisha Frampton-Clerk and Oyewole Oyekoya. 2022. Investigating the Perceived Realism of the Other User’s Look-Alike Avatars. In Proceedings of the 28th ACM Symposium on Virtual Reality Software and Technology . ACM, Tsukuba Japan, 1–5. https://doi.org/10.1145/3562939.3565636

  4. [12]

    Guo Freeman and Divine Maloney. 2021. Body, Avatar, and Me: The Presentation and Perception of Self in Social Virtual Reality. Proceedings of the ACM on Human-Computer Interaction 4, CSCW3 (Jan. 2021), 1–27. https://doi.org/ 10.1145/3432938

  5. [13]

    Chapman, and Adrian Smallbone

    Anthony Gale, Graham Spratt, Antony J. Chapman, and Adrian Smallbone. 1975. EEG correlates of eye contact and interpersonal distance. Biological Psychology 3, 4 (Dec. 1975), 237–245. https://doi.org/10.1016/0301-0511(75)90023-X

  6. [14]

    Maia Garau, Mel Slater, Vinoba Vinayagamoorthy, Andrea Brogni, Anthony Steed, and M Angela Sasse. 2003. The Impact of Avatar Realism and Eye Gaze Control on Perceived Quality of Communication in a Shared Immersive Virtual Environment. NEW HORIZONS (2003)

  7. [15]

    Cristina Gasch, Alireza Javanmardi, Azucena Garcia-Palacios, and Alain Pagani. 2024. Avatar quality: A study on presence and user preference. In 2024 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR). IEEE, Los Angeles, CA, USA, ...

  8. [16]

    Gonzalez and R

    R. Gonzalez and R. Woods. 2006. Digital image processing (3rd ed.). Prentice Hall

  9. [17]

    Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. 2024. Live- Portrait: Efficient Portrait Animation with Stitching and Retargeting Control. http://arxiv.org/abs/2407.03168 , Vol. 1, No. 1, Article . Publication date: April 2025. ...

  10. [18]

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural In- formation Processing Systems , Vol. 30. Curran Associates, Inc. https:/...

  11. [19]

    MacDorman

    Chin-Chang Ho and Karl F. MacDorman. 2017. Measuring the Uncanny Valley Effect: Refinements to Indices for Perceived Humanness, Attractiveness, and Eeriness. International Journal of Social Robotics 9, 1 (Jan. 2017), 129–139. https://doi.org/10.1007/s12369-016-0380-9

  12. [20]

    Tobias Hossfeld, Christian Keimel, Matthias Hirth, Bruno Gardlo, Julian Habigt, Klaus Diepold, and Phuoc Tran-Gia

  13. [21]

    Ziyao Huang, Fan Tang, Yong Zhang, Xiaodong Cun, Juan Cao, Jintao Li, and Tong-Yee Lee. 2024. Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework. In CVPR

  14. [22]

    Inkpen and Mara Sedlins

    Kori M. Inkpen and Mara Sedlins. 2011. Me and my avatar: exploring users’ comfort with avatars for workplace communication. In Proceedings of the ACM 2011 conference on Computer supported cooperative work . ACM, Hangzhou China, 383–386. https://doi.org/10.1145/1958824.1958883

  15. [23]

    ISO Ophthalmic optics and instruments. 2017. Ophthalmic optics — Visual acuity testing — Standard and clinical optotypes and their presentation . Standard ISO 8596:2017. International Organization for Standardization, Geneva, CH. https://www.iso.org/standard/69042.html

  16. [24]

    ITU-R Recommendation BT.500-14. 2019. Methodologies for the subjective assessment of the quality of television images . International Telecommunication Union, Geneva

  17. [25]

    ITU-T PSTR-CROWDS. 2018. Subjective evaluation of media quality using a crowdsourcing approach . International Telecommunication Union, Geneva

  18. [26]

    ITU-T Recommendation P.1320. 2022. Quality of experience assessment of extended reality meetings . International Telecommunication Union, Geneva

  19. [27]

    ITU-T Recommendation P.808. 2021. Subjective evaluation of speech quality with a crowdsourcing approach. International Telecommunication Union, Geneva

  20. [28]

    {ITU-T Recommendation P.910}. 2021. {Subjective video quality assessment methods for multimedia applications} . International Telecommunication Union, Geneva, Switzerland

  21. [29]

    ITU-T Recommendation P.910. 2021. Subjective video quality assessment methods for multimedia applications . Interna- tional Telecommunication Union, Geneva

  22. [30]

    ITU-T Recommendation P.913. 2021. Methods for the subjective assessment of video quality, audio quality and audiovisual quality of Internet video and distribution quality television in any environment . International Telecommunication Union, Geneva

  23. [31]

    ITU-T Recommendation P.918. 2021. Dimension-based subjective quality evaluation for video content . International Telecommunication Union, Geneva

  24. [32]

    J. Jung, M. Wien, and V. Baroncini. 2021. ISO/IEC JTC 1/SC 29/AG 5: Guidelines for remote experts viewing sessions

  25. [33]

    Sasa Junuzovic, Kori Inkpen, John Tang, Mara Sedlins, and Kristie Fisher. 2012. To see or not to see: a study comparing four-way avatar, video, and audio conferencing for work. In Proceedings of the 17th ACM international conference on Supporting group work. ACM, Sanibel Islan...

  26. [34]

    Christian Keimel, Julian Habigt, Clemens Horch, and Klaus Diepold. 2012. QualityCrowd-A Framework for Crowd-based Quality Evaluation. In Picture coding symposium. https://doi.org/10.1109/PCS.2012.6213338

  27. [35]

    M. G. Kendall. 1938. A New Measure of Rank Correlation. Biometrika 30, 1/2 (1938), 81–93. https://doi.org/10.2307/ 2332226 Publisher: [Oxford University Press, Biometrika Trust]

  28. [36]

    Tobias Kirschstein, Simon Giebenhain, and Matthias Nießner. 2024. DiffusionAvatars: Deferred Diffusion for High- fidelity 3D Head Avatars. In CVPR

  29. [37]

    Jari Kätsyri, Beatrice de Gelder, and Tapio Takala. 2019. Virtual Faces Evoke Only a Weak Uncanny Valley Effect: An Empirical Investigation With Controlled Virtual Face Images. Perception 48, 10 (Oct. 2019), 968–991. https: //doi.org/10.1177/0301006619869134 Publisher: SAGE Pu...

  30. [38]

    Jinghuai Lin, Johrine Cronjé, Carolin Wienrich, Paul Pauli, and Marc Erich Latoschik. 2023. Visual Indicators Represent- ing Avatars’ Authenticity in Social Virtual Reality and Their Impacts on Perceived Trustworthiness. IEEE Transactions on Visualization and Computer Graphics...

  31. [39]

    Jinlin Liu, Kai Yu, Mengyang Feng, Xiefan Guo, and Miaomiao Cui. 2024. Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation. http://arxiv.org/abs/2405.16393 arXiv:2405.16393 [cs]

  32. [40]

    Abrar Majeedi, Babak Naderi, Yasaman Hosseinkashi, Juhee Cho, Ruben Alvarez Martinez, and Ross Cutler. 2023. Full Reference Video Quality Assessment for Machine Learning-Based Video Codecs. http://arxiv.org/abs/2309.00769 , Vol. 1, No. 1, Article . Publication date: April 2025...

  33. [41]

    Maya Mathur and David Reichling. 2016. Navigating a social world with robot partners: A quantitative cartography of the Uncanny Valley. Cognition 146 (Jan. 2016), 22–32

  34. [42]

    Bülthoff

    Rachel McDonnell, Martin Breidt, and Heinrich H. Bülthoff. 2012. Render me real?: investigating the effect of render style on the perception of animated virtual humans. ACM Transactions on Graphics 31, 4 (Aug. 2012), 1–11. https://doi.org/10.1145/2185520.2185587

  35. [43]

    Masahiro Mori, Karl MacDorman, and Norri Kageki. 2012. The Uncanny Valley [From the Field]. IEEE Robotics & Automation Magazine 19, 2 (June 2012), 98–100. https://doi.org/10.1109/MRA.2012.2192811

  36. [44]

    Babak Naderi and Ross Cutler. 2020. An Open Source Implementation of ITU-T Recommendation P.808 with Validation. INTERSPEECH (Oct. 2020), 2862–2866. https://doi.org/10.21437/Interspeech.2020-2665

  37. [45]

    Babak Naderi and Ross Cutler. 2023. A crowdsourcing approach to video quality assessment. http://arxiv.org/abs/ 2204.06784 arXiv:2204.06784

  38. [46]

    Babak Naderi and Ross Cutler. 2024. A crowdsourcing approach to video quality assessment. In ICASSP

  39. [47]

    Babak Naderi and Sebastian Möller. 2020. Application of just-noticeable difference in quality as environment suitability test for crowdsourcing speech quality assessment task. In2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX). IEEE, 1–6

  40. [48]

    Nguyen and John Canny

    David T. Nguyen and John Canny. 2007. Multiview: improving trust in group video conferencing through spatial faithfulness. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI ’07 . ACM Press, San Jose, California, USA, 1465–1474. https://doi.org...

  41. [49]

    Nguyen and John Canny

    David T. Nguyen and John Canny. 2009. More than Face-to-Face: Empathy Effects of Video Framing

  42. [50]

    ITU-T Recommendation P.912. 2016. Subjective video quality assessment methods for recognition tasks

  43. [51]

    Vrushank Phadnis, Kristin Moore, and Mar Gonzalez Franco. 2023. Avatars in Work Meetings: Correlation Between Photorealism and Appeal. http://arxiv.org/abs/2304.01405 arXiv:2304.01405 [cs]

  44. [52]

    Margaret H Pinson and Lucjan Janowski. 2014. A new subjective audiovisual & video quality testing recommendation. An Era of Change (2014), 51

  45. [53]

    Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao. 2021. Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates. http://arxiv.org/abs/2108.08020 arXiv:2108.08020 [cs]

  46. [54]

    Benjamin Rainer, Markus Waltl, and Christian Timmerer. 2013. A web based subjective evaluation platform. In 2013 Fifth International Workshop on Quality of Multimedia Experience (QoMEX) . 24–25. https://doi.org/10.1109/QoMEX. 2013.6603196

  47. [55]

    Rakesh Rao Ramachandra Rao, Steve Göring, and Alexander Raake. 2021. Towards High Resolution Video Quality Assessment in the Crowd. In 2021 13th International Conference on Quality of Multimedia Experience (QoMEX) . 1–6. https://doi.org/10.1109/QoMEX51781.2021.9465425 ISSN: 2472-7814

  48. [56]

    Rabindra Ratan and Béatrice Hasler. 2014. Playing well with virtual classmates: Relating avatar design to group satisfaction. Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing (Feb. 2014), 564–573. https://doi.org/10.1145/2531602.2531732

  49. [57]

    Christos Sagonas, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic. 2013. A Semi-automatic Methodology for Facial Landmark Annotation. In 2013 IEEE Conference on Computer Vision and Pattern Recognition Workshops . IEEE, OR, USA, 896–903. https://doi.org/10.1109/CVPR...

  50. [58]

    Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. 2024. Relightable Gaussian Codec Avatars. In CVPR

  51. [59]

    Falk Schiffner and Sebastian Möller. 2017. Defining the relevant perceptual quality space for video and video-telephony. In 2017 Ninth International Conference on Quality of Multimedia Experience (QoMEX) . IEEE, 1–3

  52. [60]

    Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang

  53. [61]

    Haejung Suk and Teemu H. Laine. 2023. Influence of Avatar Facial Appearance on Users’ Perceived Embodiment and Presence in Immersive Virtual Reality. Electronics 12, 3 (Jan. 2023), 583. https://doi.org/10.3390/electronics12030583 Number: 3 Publisher: Multidisciplinary Digital ...

  54. [62]

    Michela Testolina and Touradj Ebrahimi. 2021. Review of subjective quality assessment methodologies and standards for compressed images evaluation. In Applications of Digital Image Processing XLIV , Vol. 11842. SPIE, 302–315. https: //doi.org/10.1117/12.2597813

  55. [63]

    Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. 2024. EMO: Emote Portrait Alive – Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions. http://arxiv.org/abs/2402.17485 arXiv:2402.17485 [cs]

  56. [64]

    Charlton

    Angela Tinwell, Deborah Abdel Nabi, and John P. Charlton. 2013. Perception of psychopathy and the Uncanny Valley in virtual characters. Computers in Human Behavior 29, 4 (July 2013), 1617–1625. https://doi.org/10.1016/j.chb.2013.01.008 , Vol. 1, No. 1, Article . Publication da...

  57. [65]

    Toshiko Tominaga, Takanori Hayashi, Jun Okamoto, and Akira Takahashi. 2010. Performance comparisons of subjective quality assessment methods for mobile video. In2010 Second International Workshop on Quality of Multimedia Experience (QoMEX). 82–87. https://doi.org/10.1109/QOMEX...

  58. [66]

    Nicholas Toothman and Michael Neff. 2019. The Impact of Avatar Tracking Errors on User Experience in VR. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR) . IEEE, Osaka, Japan, 756–766. https://doi.org/10.1109/ VR.2019.8798108

  59. [67]

    Bernhard Treutwein. 1995. Adaptive psychophysical procedures. Vision research 35, 17 (1995), 2503–2522

  60. [68]

    Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2019. FVD: A new metric for video generation. ICLR (2019)

  61. [69]

    Evgeniy Upenik, Michela Testolina, Joao Ascenso, Fernando Pereira, and Touradj Ebrahimi (Eds.). 2021. Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image Coding. IEEE Visual Communications and Image Processing (VCIP 2021) (2021)

  62. [70]

    Video Quality Experts Group. 2010. Report on the validation of video quality models for high definition video content. (2010). https://www.its.bldrdoc.gov/media/4212/vqeg_hdtv_final_report_version_2.0.zip

  63. [71]

    Tan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2024. DisCo: Disentangled Control for Realistic Human Dance Generation. http://arxiv.org/abs/2307. 00040 arXiv:2307.00040 [cs]

  64. [72]

    Xuanyu Wang, Weizhan Zhang, Christian Sandor, and Hongbo Fu. 2024. Real-and-Present: Investigating the Use of Life-Size 2D Video Avatars in HMD-Based AR Teleconferencing. https://doi.org/10.48550/arXiv.2401.02171 arXiv:2401.02171 [cs]

  65. [73]

    Bovik, H.R

    Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13, 4 (April 2004), 600–612. https://doi.org/10.1109/TIP. 2003.819861

  66. [74]

    Wang, E.P

    Z. Wang, E.P. Simoncelli, and A.C. Bovik. 2003. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 . IEEE, Pacific Grove, CA, USA, 1398–1402. https://doi.org/10.1109/ACSSC.2003.1292216

  67. [75]

    Florian Weidner, Gerd Boettcher, Stephanie Arévalo Arboleda, Chenyao Diao, Luljeta Sinani, Christian Kunert, Christoph Gerhardt, Wolfgang Broll, and Alexander Raake. 2023. A Systematic Review on the Visualization of Avatars and Agents in AR & VR displayed using Head-Mounted Di...

  68. [76]

    Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. 2024. VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time. http://arxiv.org/abs/2404.10667 arXiv:2404.10667 [cs]

  69. [77]

    Yuelang Xu, Benwang Chen, Zhe Li, Hongwen Zhang, Lizhen Wang, Zerong Zheng, and Yebin Liu. 2024. Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians. In CVPR

  70. [78]

    Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, and Haoqiang Fan. 2024. MegActor: Harness the Power of Raw Video for Vivid Portrait Animation. http://arxiv.org/abs/2405.20851 arXiv:2405.20851 [cs]

  71. [79]

    Williamson, Fei Chen, Fuzheng Yang, and Shidong Shang

    Gaoxiong Yi, Wei Xiao, Yiming Xiao, Babak Naderi, Sebastian Möller, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen, Fuzheng Yang, and Shidong Shang. 2022. ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Asse...

  72. [80]

    Kevin Yu, Gleb Gorbachev, Ulrich Eck, Frieder Pankratz, Nassir Navab, and Daniel Roth. 2021. Avatars for Teleconsulta- tion: Effects of Avatar Embodiment Techniques on User Perception in 3D Asymmetric Telepresence. IEEE Transactions on Visualization and Computer Graphics 27, 1...

  73. [81]

    Eduard Zell, Carlos Aliaga, Adrian Jarabo, Katja Zibrek, Diego Gutierrez, Rachel McDonnell, and Mario Botsch. 2015. To stylize or not to stylize?: the effect of shape and material stylization on the perception of computer-generated faces. ACM Transactions on Graphics 34, 6 (No...

  74. [82]

    Efros, Eli Shechtman, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Salt Lake City, UT, 586–595. https://doi....

  75. [83]

    Mingyuan Zhou, Rakib Hyder, Ziwei Xuan, and Guojun Qi. 2024. UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures. In CVPR. , Vol. 1, No. 1, Article . Publication date: April 2025. A multidimensional measurement of photorealistic avata...

  76. [2014]

    IEEE Transactions on Multimedia 16, 2 (Feb

    Best Practices for QoE Crowdtesting: QoE Assessment With Crowdsourcing. IEEE Transactions on Multimedia 16, 2 (Feb. 2014), 541–558. https://doi.org/10.1109/TMM.2013.2291663

  77. [2024]

    SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting. In CVPR. http://arxiv.org/abs/2403.05087

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.