REVIEW 3 major objections 3 minor 85 references
A multidimensional measurement of photorealistic avatar quality of experience
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that subjective human ratings are the only reliable way to measure photorealistic avatar quality of experience: standard objective metrics correlate weakly with nine of ten measured dimensions and only moderately with…
desk verdict Solid open-source avatar QoE framework with a credible negative result on objective metrics, but the no-uncanny-valley claim exceeds what the data can support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the open-source crowdsourced subjective test framework, an implementation of the P.910 subjective video quality assessment standard extended to avatars, with rater qualification, display calibration, gold clips, trapping items, and repeated items to control data quality. It uses two survey templates: Template A asks for agreement with eight statements about the avatar alone, and Template B shows the avatar side-by-side with the real person to rate resemblance and emotion accuracy. The framework's output is a per-clip and per-condition mean opinion score for each dimension, and its validation against laboratory video studies and an expert panel is what gives the correlations their weight. The analytic load is carried by inter-dimension Pearson correlations, a principal component analysis that finds two components explaining 94% of the variance, and a linear fit of affinity against realism whose R²=0.966 is the basis for the no-uncanny-valley conclusion.
What would settle it
Have separate groups of raters each judge only one dimension of the same avatar clips, then compare those single-dimension scores with the all-items-at-once scores. If trust, comfort, and appropriateness no longer track realism when realism is not among the questions asked — or if the average inter-item correlation drops well below 0.996 — then the reported high correlations are a halo or order effect, and the paper's dimensionality reduction and no-uncanny-valley linearity would not survive independent measurement.
Extended reading notes
Core claim
The central discovery is that the quality of experience of a photorealistic avatar is a multidimensional subjective quantity that standard objective metrics do not capture. Using a validated crowdsourced survey, the paper measures ten dimensions — realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness (reported as not creepy), formality, affinity, resemblance to the person, and emotion accuracy — across 19 avatar models. The inter-item correlations among the eight Template A dimensions average PCC=0.996, and for avatars with realism above 2 on the 1–5 scale they average 0.999, which the paper interprets as a strong redundancy that permits reducing the ten dimensions to three via principal component analysis. Against PSNR, SSIM, LPIPS, FID, and FVD, the subjective dimensions show weak correlation, with only emotion accuracy reaching moderate correlation (e.g., PSNR=0.63, LPIPS=-0.67). In addition, affinity versus realism is well described by a straight line (R²=0.966), which the paper takes as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.
Load-bearing premise
The ten survey items are treated as independently meaningful dimensions, but the near-perfect inter-item correlations (average PCC=0.996) are exactly what a single global impression would look like, so if raters are not actually distinguishing the dimensions, the dimensionality reduction and the linear affinity–realism relation would be artifacts of rating co-movement rather than evidence about the underlying quality structure.
Editorial extensions
If this is right
- Avatar developers who optimize only PSNR, SSIM, LPIPS, FID, or FVD will leave most of the quality-of-experience space unmeasured, so subjective testing becomes necessary for usability claims.
- For avatars with realism above 2 on the 1–5 scale, one Template A item (e.g., realism) plus the two Template B items (emotion accuracy and resemblance) is enough, shrinking the survey from ten items to three.
- Avatars that are less realistic than real video will score lower on trust, comfort using, comfort interacting, work appropriateness, formality, and affinity, and higher on creepiness; matching real video requires matching its realism.
- Because affinity rises linearly with realism (R²=0.966), designers of telecommunication avatars should maximize realism rather than stop short of a presumed uncanny valley.
- Some state-of-the-art photorealistic avatars approach real video on these subjective dimensions, but none reaches real-video realism, so the gap between avatars and real video remains measurable and large.
Reading between the lines
- A testable extension the paper leaves implicit: if the near-perfect inter-item correlations are a halo effect, single-dimension ratings from independent groups will diverge from the all-items survey; running that comparison is a direct check on whether the ten dimensions are really separate constructs.
- The no-uncanny-valley conclusion is tied to the stimulus set used here — 2D talking-head clips in telecommunication poses; the same survey applied to full-body avatars, VR/AR displays, or avatars with motion artifacts could still find a dip the paper does not test.
- Because emotion accuracy and resemblance show moderate-to-strong correlation with PSNR, SSIM, and LPIPS when computed on the face region, a learned perceptual metric trained on subjective avatar ratings could plausibly close much of the objective-subjective gap, a direction the paper mentions as future work.
- The framework is a passive-viewing test; an interactive variant in which a rater converses with a live avatar could change trust and comfort ratings, since the paper's own reproducibility and accuracy claims rest on passive observation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an open-source crowdsourced test framework, based on the authors' P.910 implementation, for measuring ten subjective dimensions of photorealistic avatar quality of experience: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance, and emotion accuracy. The framework is validated against VQEG laboratory video-quality data and against a panel of experts, and its run-to-run reproducibility is reported. Using 19 avatar models and real-video baselines, the authors find that the subjective dimensions are weakly correlated with common objective metrics (PSNR, SSIM, LPIPS, FID, FVD), that eight Template A dimensions become nearly collinear for avatars with realism above a threshold, and that affinity is linearly related to realism, which they interpret as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.
Significance. If the claims hold, this is a useful contribution: an open, validated, reproducible subjective testing methodology for avatar quality of experience, a public test set, and evidence that objective metrics currently used in avatar development do not capture the usability dimensions that matter. The framework validation is genuinely strong: run-to-run clip-level correlations are high for Template A (Table 8, average PCC 0.962), model-level reproducibility is above 0.98 (Table 7), accuracy against an expert panel averages PCC 0.904 (Table 6), and the crowdsourcing implementation is validated against external VQEG lab data (Tables 2-4). The main risk is that the two most prominent perceptual conclusions, the dimensionality reduction and the absence of an uncanny valley, rest on near-perfect inter-item correlations that may reflect a general impression halo rather than distinct underlying constructs, and the no-uncanny-valley conclusion is drawn from a linear fit on clustered clip-level points without a direct test for a local dip.
major comments (3)
- [§5.5, Fig. 10] I would ask the authors to add an avatar-random-effect analysis at the model level, a direct test for a local dip in the near-photorealistic range, confidence intervals on the fit, or to substantially soften the conclusion and restrict it to the range actually sampled.
- [§5.3, Tables 9-11] Since the realism > 2 threshold is chosen post hoc and the claim directly feeds the developer guideline that 'only 3 dimensions need to be measured' (§6.2), the paper should provide stronger evidence of discriminant validity, for example a confirmatory factor analysis with a specified measurement model, or should reframe the claim as a practical shortcut rather than a statement about the underlying psychological dimensions.
- [§5.4, Tables 12-13] Please report confidence intervals or bootstrap intervals for Table 12, and adjust the abstract and conclusions to reflect the face-limited results.
minor comments (3)
- [§4.3] The text says the average clip-level run-run PCC is 0.968, but Table 8 gives Template A averages of 0.962 and a Template B emotion-accuracy average of 0.870; the text should match the table, and the lower emotion-accuracy reproducibility should be acknowledged as a caveat to the 'highly reproducible' claim.
- [§6] The first sentence lists 'PSNR, SSIM, LPIPS, FIV, FVD'; this should be FID, not FIV.
- [§3.1] The open-source tool is named P.910, identical to the ITU-T Recommendation P.910. This naming is likely to confuse readers; consider renaming the repository or adding a clarifying sentence in Section 3.1.
Circularity Check
No significant circularity: the framework's ratings are primary data and are validated against external VQEG lab data and an expert panel; the authors' self-citation of their own P.910 tool is not load-bearing.
full rationale
The central claims are built on new crowdsourced subjective ratings of avatar clips, not on fitted parameters or on prior results by the same authors. The P.910 crowdsourcing implementation is the authors' own tool, but its validity is established against external VQEG HDTV laboratory data (Section 3.1.9, PCC=0.952 per sequence) and the avatar extension is compared to an independent expert panel (Table 6, average PCC=0.904). Reproducibility is demonstrated across five mutually exclusive MTurk runs (Tables 7-8). The weak objective-metric correlations (Table 12) are computed from the collected MOS and standard objective metrics, so they are empirical findings rather than constructed equivalences. The most conceptually vulnerable claim is the no-uncanny-valley conclusion (Section 5.5), because the near-perfect inter-item correlations (Table 9, average PCC=0.996) suggest a strong halo effect; a linear affinity-realism fit (R2=0.966) may partly reflect rating co-movement rather than independent psychological dimensions. However, the paper openly reports these correlations and the limitations that the test is passive and 2D, and the linearity claim is an empirical inference, not a reduction of the conclusion to its inputs by definition. No step in the derivation chain equates a predicted quantity to a fitted parameter or imports a uniqueness theorem from the authors' prior work. Hence the circularity score is low.
Assumptions & free parameters
free parameters (3)
- Affinity-realism linear fit slope =
not reported
- Affinity-realism linear fit intercept =
not reported
- Realism threshold for subgroup analysis =
2 on a 1-5 Likert scale
assumptions (5)
- domain assumption Passive viewing of short 2D video clips by crowdsourced raters is a valid proxy for avatar quality of experience in telecommunication scenarios.
- domain assumption The 5-expert panel provides an accurate ground truth for the subjective dimensions.
- domain assumption The ten survey items measure distinct psychological constructs rather than a single overall impression.
- standard math PCA with two retained components (94% variance) is a valid linear reduction of the ten items.
- domain assumption Objective metrics computed on background-removed, face-aligned original videos are a fair basis for comparing avatar output to the source.
Cite this review
Pith. "Pith review of A multidimensional measurement of photorealistic avatar quality of experience." pith.science (2026). https://pith.science/paper/XEYRNSWY
@misc{pith2026241109066,
author = {Pith},
title = {Pith review of: A multidimensional measurement of photorealistic avatar quality of experience},
year = {2026},
howpublished = {\url{https://pith.science/paper/XEYRNSWY}},
note = {Machine review of arXiv:2411.09066}
}
read the original abstract
Photorealistic avatars are human avatars that look, move, and talk like real people. The performance of photorealistic avatars has significantly improved recently based on objective metrics such as PSNR, SSIM, LPIPS, FID, and FVD. However, recent photorealistic avatar publications do not provide subjective tests of the avatars to measure human usability factors. We provide an open source test framework to subjectively measure photorealistic avatar performance in ten dimensions: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance to the person, and emotion accuracy. Using telecommunication scenarios, we show that the correlation of nine of these subjective metrics with PSNR, SSIM, LPIPS, FID, and FVD is weak, and moderate for emotion accuracy. The crowdsourced subjective test framework is highly reproducible and accurate when compared to a panel of experts. We analyze a wide range of avatars from photorealistic to cartoon-like and show that some photorealistic avatars are approaching real video performance based on these dimensions. We also find that for avatars above a certain level of realism, eight of these measured dimensions are strongly correlated. This means that avatars that are not as realistic as real video will have lower trust, comfortableness using, comfortableness interacting with, appropriateness for work, formality, and affinity, and higher creepiness compared to real video. In addition, because there is a strong linear relationship between avatar affinity and realism, there is no uncanny valley effect for photorealistic avatars in the telecommunication scenario. We suggest several extensions of this test framework for future work and discuss design implications for telecommunication systems. The test framework is available at https://github.com/microsoft/P.910.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Jeremy N. Bailenson. 2021. Nonverbal overload: A theoretical argument for the causes of Zoom fatigue. Technology, Mind, and Behavior 2, 1 (Feb. 2021). https://doi.org/10.1037/tmb0000030
-
[2]
Biehl, Daniel Avrahami, and Anthony Dunnigan
Jacob T. Biehl, Daniel Avrahami, and Anthony Dunnigan. 2015. Not Really There: Understanding Embodied Communication Affordances in Team Perception and Participation. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing . ACM, Vancouver BC Canada, 1567–1575. https: //doi.org/10.1145/2675133.2675220
arXiv 2015
-
[3]
Ryan Canales, Doug Roble, and Michael Neff. 2024. The Impact of Avatar Stylization on Trust. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR) . IEEE, Orlando, FL, USA, 418–428. https://doi.org/10.1109/VR58804.2024. 00063
arXiv 2024
-
[4]
Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, and Jason Saragih. 2022. Authentic volumetric avatars from a phone scan. ACM Transactions on Graphics 41, 4 (July 2022), 163:1–163:19. https://doi.org/10.1145/3528223.3530143
arXiv 2022
-
[5]
Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Bo Qiu, Stefano Zanetti, Ioannis Katsavounidis, and Alan C. Bovik. 2024. Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality. http://arxiv.org/abs/2408.07041 arXiv:2408.07041 [eess]
work page Pith review arXiv 2024
-
[6]
JH Clark. 1924. The Ishihara test for color blindness. American Journal of Physiological Optics (1924)
1924
-
[7]
Christophe Croux and Catherine Dehon. 2010. Influence functions of the Spearman and Kendall correlation measures. Statistical methods & applications 19 (2010), 497–515
2010
-
[8]
Yu Deng, Duomin Wang, Xiaohang Ren, Xingyu Chen, and Baoyuan Wang. 2024. Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data. In CVPR
2024
Show all 85 references
-
[9]
Macdorman
Alexander Diel, Sarah Weigelt, and Karl F. Macdorman. 2022. A Meta-analysis of the Uncanny Valley’s Independent and Dependent Variables. ACM Transactions on Human-Robot Interaction 11, 1 (March 2022), 1–33. https://doi.org/10. 1145/3470742
2022
-
[10]
Geraldine Fauville, Mufan Luo, Anna C. M. Queiroz, Jeremy N. Bailenson, and Jeff Hancock. 2021. Nonverbal Mechanisms Predict Zoom Fatigue and Explain Why Women Experience Higher Levels than Men. SSRN Electronic Journal (2021). https://doi.org/10.2139/ssrn.3820035
2021 doi
-
[11]
Aisha Frampton-Clerk and Oyewole Oyekoya. 2022. Investigating the Perceived Realism of the Other User’s Look-Alike Avatars. In Proceedings of the 28th ACM Symposium on Virtual Reality Software and Technology . ACM, Tsukuba Japan, 1–5. https://doi.org/10.1145/3562939.3565636
2022
-
[12]
Guo Freeman and Divine Maloney. 2021. Body, Avatar, and Me: The Presentation and Perception of Self in Social Virtual Reality. Proceedings of the ACM on Human-Computer Interaction 4, CSCW3 (Jan. 2021), 1–27. https://doi.org/ 10.1145/3432938
2021 doi
-
[13]
Chapman, and Adrian Smallbone
Anthony Gale, Graham Spratt, Antony J. Chapman, and Adrian Smallbone. 1975. EEG correlates of eye contact and interpersonal distance. Biological Psychology 3, 4 (Dec. 1975), 237–245. https://doi.org/10.1016/0301-0511(75)90023-X
1975 doi
-
[14]
Maia Garau, Mel Slater, Vinoba Vinayagamoorthy, Andrea Brogni, Anthony Steed, and M Angela Sasse. 2003. The Impact of Avatar Realism and Eye Gaze Control on Perceived Quality of Communication in a Shared Immersive Virtual Environment. NEW HORIZONS (2003)
2003
-
[15]
Cristina Gasch, Alireza Javanmardi, Azucena Garcia-Palacios, and Alain Pagani. 2024. Avatar quality: A study on presence and user preference. In 2024 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR). IEEE, Los Angeles, CA, USA, ...
2024
-
[16]
Gonzalez and R
R. Gonzalez and R. Woods. 2006. Digital image processing (3rd ed.). Prentice Hall
2006
-
[17]
Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. 2024. Live- Portrait: Efficient Portrait Animation with Stitching and Retargeting Control. http://arxiv.org/abs/2407.03168 , Vol. 1, No. 1, Article . Publication date: April 2025. ...
2024 arXiv
-
[18]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural In- formation Processing Systems , Vol. 30. Curran Associates, Inc. https:/...
2017
-
[19]
MacDorman
Chin-Chang Ho and Karl F. MacDorman. 2017. Measuring the Uncanny Valley Effect: Refinements to Indices for Perceived Humanness, Attractiveness, and Eeriness. International Journal of Social Robotics 9, 1 (Jan. 2017), 129–139. https://doi.org/10.1007/s12369-016-0380-9
2017 doi
-
[20]
Tobias Hossfeld, Christian Keimel, Matthias Hirth, Bruno Gardlo, Julian Habigt, Klaus Diepold, and Phuoc Tran-Gia
-
[21]
Ziyao Huang, Fan Tang, Yong Zhang, Xiaodong Cun, Juan Cao, Jintao Li, and Tong-Yee Lee. 2024. Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework. In CVPR
2024
-
[22]
Inkpen and Mara Sedlins
Kori M. Inkpen and Mara Sedlins. 2011. Me and my avatar: exploring users’ comfort with avatars for workplace communication. In Proceedings of the ACM 2011 conference on Computer supported cooperative work . ACM, Hangzhou China, 383–386. https://doi.org/10.1145/1958824.1958883
2011
-
[23]
ISO Ophthalmic optics and instruments. 2017. Ophthalmic optics — Visual acuity testing — Standard and clinical optotypes and their presentation . Standard ISO 8596:2017. International Organization for Standardization, Geneva, CH. https://www.iso.org/standard/69042.html
2017
-
[24]
ITU-R Recommendation BT.500-14. 2019. Methodologies for the subjective assessment of the quality of television images . International Telecommunication Union, Geneva
2019
-
[25]
ITU-T PSTR-CROWDS. 2018. Subjective evaluation of media quality using a crowdsourcing approach . International Telecommunication Union, Geneva
2018
-
[26]
ITU-T Recommendation P.1320. 2022. Quality of experience assessment of extended reality meetings . International Telecommunication Union, Geneva
2022
-
[27]
ITU-T Recommendation P.808. 2021. Subjective evaluation of speech quality with a crowdsourcing approach. International Telecommunication Union, Geneva
2021
-
[28]
{ITU-T Recommendation P.910}. 2021. {Subjective video quality assessment methods for multimedia applications} . International Telecommunication Union, Geneva, Switzerland
2021
-
[29]
ITU-T Recommendation P.910. 2021. Subjective video quality assessment methods for multimedia applications . Interna- tional Telecommunication Union, Geneva
2021
-
[30]
ITU-T Recommendation P.913. 2021. Methods for the subjective assessment of video quality, audio quality and audiovisual quality of Internet video and distribution quality television in any environment . International Telecommunication Union, Geneva
2021
-
[31]
ITU-T Recommendation P.918. 2021. Dimension-based subjective quality evaluation for video content . International Telecommunication Union, Geneva
2021
-
[32]
J. Jung, M. Wien, and V. Baroncini. 2021. ISO/IEC JTC 1/SC 29/AG 5: Guidelines for remote experts viewing sessions
2021
-
[33]
Sasa Junuzovic, Kori Inkpen, John Tang, Mara Sedlins, and Kristie Fisher. 2012. To see or not to see: a study comparing four-way avatar, video, and audio conferencing for work. In Proceedings of the 17th ACM international conference on Supporting group work. ACM, Sanibel Islan...
2012
-
[34]
Christian Keimel, Julian Habigt, Clemens Horch, and Klaus Diepold. 2012. QualityCrowd-A Framework for Crowd-based Quality Evaluation. In Picture coding symposium. https://doi.org/10.1109/PCS.2012.6213338
2012
-
[35]
M. G. Kendall. 1938. A New Measure of Rank Correlation. Biometrika 30, 1/2 (1938), 81–93. https://doi.org/10.2307/ 2332226 Publisher: [Oxford University Press, Biometrika Trust]
1938
-
[36]
Tobias Kirschstein, Simon Giebenhain, and Matthias Nießner. 2024. DiffusionAvatars: Deferred Diffusion for High- fidelity 3D Head Avatars. In CVPR
2024
-
[37]
Jari Kätsyri, Beatrice de Gelder, and Tapio Takala. 2019. Virtual Faces Evoke Only a Weak Uncanny Valley Effect: An Empirical Investigation With Controlled Virtual Face Images. Perception 48, 10 (Oct. 2019), 968–991. https: //doi.org/10.1177/0301006619869134 Publisher: SAGE Pu...
2019 doi
-
[38]
Jinghuai Lin, Johrine Cronjé, Carolin Wienrich, Paul Pauli, and Marc Erich Latoschik. 2023. Visual Indicators Represent- ing Avatars’ Authenticity in Social Virtual Reality and Their Impacts on Perceived Trustworthiness. IEEE Transactions on Visualization and Computer Graphics...
2023
-
[39]
Jinlin Liu, Kai Yu, Mengyang Feng, Xiefan Guo, and Miaomiao Cui. 2024. Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation. http://arxiv.org/abs/2405.16393 arXiv:2405.16393 [cs]
2024 arXiv
-
[40]
Abrar Majeedi, Babak Naderi, Yasaman Hosseinkashi, Juhee Cho, Ruben Alvarez Martinez, and Ross Cutler. 2023. Full Reference Video Quality Assessment for Machine Learning-Based Video Codecs. http://arxiv.org/abs/2309.00769 , Vol. 1, No. 1, Article . Publication date: April 2025...
2023 arXiv
-
[41]
Maya Mathur and David Reichling. 2016. Navigating a social world with robot partners: A quantitative cartography of the Uncanny Valley. Cognition 146 (Jan. 2016), 22–32
2016
-
[42]
Bülthoff
Rachel McDonnell, Martin Breidt, and Heinrich H. Bülthoff. 2012. Render me real?: investigating the effect of render style on the perception of animated virtual humans. ACM Transactions on Graphics 31, 4 (Aug. 2012), 1–11. https://doi.org/10.1145/2185520.2185587
2012
-
[43]
Masahiro Mori, Karl MacDorman, and Norri Kageki. 2012. The Uncanny Valley [From the Field]. IEEE Robotics & Automation Magazine 19, 2 (June 2012), 98–100. https://doi.org/10.1109/MRA.2012.2192811
2012
-
[44]
Babak Naderi and Ross Cutler. 2020. An Open Source Implementation of ITU-T Recommendation P.808 with Validation. INTERSPEECH (Oct. 2020), 2862–2866. https://doi.org/10.21437/Interspeech.2020-2665
2020 doi
-
[45]
Babak Naderi and Ross Cutler. 2023. A crowdsourcing approach to video quality assessment. http://arxiv.org/abs/ 2204.06784 arXiv:2204.06784
2023 arXiv
-
[46]
Babak Naderi and Ross Cutler. 2024. A crowdsourcing approach to video quality assessment. In ICASSP
2024
-
[47]
Babak Naderi and Sebastian Möller. 2020. Application of just-noticeable difference in quality as environment suitability test for crowdsourcing speech quality assessment task. In2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX). IEEE, 1–6
2020
-
[48]
Nguyen and John Canny
David T. Nguyen and John Canny. 2007. Multiview: improving trust in group video conferencing through spatial faithfulness. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI ’07 . ACM Press, San Jose, California, USA, 1465–1474. https://doi.org...
2007
-
[49]
Nguyen and John Canny
David T. Nguyen and John Canny. 2009. More than Face-to-Face: Empathy Effects of Video Framing
2009
-
[50]
ITU-T Recommendation P.912. 2016. Subjective video quality assessment methods for recognition tasks
2016
-
[51]
Vrushank Phadnis, Kristin Moore, and Mar Gonzalez Franco. 2023. Avatars in Work Meetings: Correlation Between Photorealism and Appeal. http://arxiv.org/abs/2304.01405 arXiv:2304.01405 [cs]
2023 arXiv
-
[52]
Margaret H Pinson and Lucjan Janowski. 2014. A new subjective audiovisual & video quality testing recommendation. An Era of Change (2014), 51
2014
-
[53]
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao. 2021. Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates. http://arxiv.org/abs/2108.08020 arXiv:2108.08020 [cs]
2021 arXiv
-
[54]
Benjamin Rainer, Markus Waltl, and Christian Timmerer. 2013. A web based subjective evaluation platform. In 2013 Fifth International Workshop on Quality of Multimedia Experience (QoMEX) . 24–25. https://doi.org/10.1109/QoMEX. 2013.6603196
2013
-
[55]
Rakesh Rao Ramachandra Rao, Steve Göring, and Alexander Raake. 2021. Towards High Resolution Video Quality Assessment in the Crowd. In 2021 13th International Conference on Quality of Multimedia Experience (QoMEX) . 1–6. https://doi.org/10.1109/QoMEX51781.2021.9465425 ISSN: 2472-7814
2021
-
[56]
Rabindra Ratan and Béatrice Hasler. 2014. Playing well with virtual classmates: Relating avatar design to group satisfaction. Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing (Feb. 2014), 564–573. https://doi.org/10.1145/2531602.2531732
2014
-
[57]
Christos Sagonas, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic. 2013. A Semi-automatic Methodology for Facial Landmark Annotation. In 2013 IEEE Conference on Computer Vision and Pattern Recognition Workshops . IEEE, OR, USA, 896–903. https://doi.org/10.1109/CVPR...
2013 doi
-
[58]
Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. 2024. Relightable Gaussian Codec Avatars. In CVPR
2024
-
[59]
Falk Schiffner and Sebastian Möller. 2017. Defining the relevant perceptual quality space for video and video-telephony. In 2017 Ninth International Conference on Quality of Multimedia Experience (QoMEX) . IEEE, 1–3
2017
-
[60]
Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang
-
[61]
Haejung Suk and Teemu H. Laine. 2023. Influence of Avatar Facial Appearance on Users’ Perceived Embodiment and Presence in Immersive Virtual Reality. Electronics 12, 3 (Jan. 2023), 583. https://doi.org/10.3390/electronics12030583 Number: 3 Publisher: Multidisciplinary Digital ...
2023 doi
-
[62]
Michela Testolina and Touradj Ebrahimi. 2021. Review of subjective quality assessment methodologies and standards for compressed images evaluation. In Applications of Digital Image Processing XLIV , Vol. 11842. SPIE, 302–315. https: //doi.org/10.1117/12.2597813
2021 doi
-
[63]
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. 2024. EMO: Emote Portrait Alive – Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions. http://arxiv.org/abs/2402.17485 arXiv:2402.17485 [cs]
2024 arXiv
-
[64]
Charlton
Angela Tinwell, Deborah Abdel Nabi, and John P. Charlton. 2013. Perception of psychopathy and the Uncanny Valley in virtual characters. Computers in Human Behavior 29, 4 (July 2013), 1617–1625. https://doi.org/10.1016/j.chb.2013.01.008 , Vol. 1, No. 1, Article . Publication da...
2013 doi
-
[65]
Toshiko Tominaga, Takanori Hayashi, Jun Okamoto, and Akira Takahashi. 2010. Performance comparisons of subjective quality assessment methods for mobile video. In2010 Second International Workshop on Quality of Multimedia Experience (QoMEX). 82–87. https://doi.org/10.1109/QOMEX...
2010
-
[66]
Nicholas Toothman and Michael Neff. 2019. The Impact of Avatar Tracking Errors on User Experience in VR. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR) . IEEE, Osaka, Japan, 756–766. https://doi.org/10.1109/ VR.2019.8798108
2019
-
[67]
Bernhard Treutwein. 1995. Adaptive psychophysical procedures. Vision research 35, 17 (1995), 2503–2522
1995
-
[68]
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2019. FVD: A new metric for video generation. ICLR (2019)
2019
-
[69]
Evgeniy Upenik, Michela Testolina, Joao Ascenso, Fernando Pereira, and Touradj Ebrahimi (Eds.). 2021. Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image Coding. IEEE Visual Communications and Image Processing (VCIP 2021) (2021)
2021
-
[70]
Video Quality Experts Group. 2010. Report on the validation of video quality models for high definition video content. (2010). https://www.its.bldrdoc.gov/media/4212/vqeg_hdtv_final_report_version_2.0.zip
2010
-
[71]
Tan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2024. DisCo: Disentangled Control for Realistic Human Dance Generation. http://arxiv.org/abs/2307. 00040 arXiv:2307.00040 [cs]
2024 arXiv
- [72]
-
[73]
Bovik, H.R
Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13, 4 (April 2004), 600–612. https://doi.org/10.1109/TIP. 2003.819861
2004
-
[74]
Wang, E.P
Z. Wang, E.P. Simoncelli, and A.C. Bovik. 2003. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 . IEEE, Pacific Grove, CA, USA, 1398–1402. https://doi.org/10.1109/ACSSC.2003.1292216
2003 arXiv
-
[75]
Florian Weidner, Gerd Boettcher, Stephanie Arévalo Arboleda, Chenyao Diao, Luljeta Sinani, Christian Kunert, Christoph Gerhardt, Wolfgang Broll, and Alexander Raake. 2023. A Systematic Review on the Visualization of Avatars and Agents in AR & VR displayed using Head-Mounted Di...
2023
-
[76]
Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. 2024. VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time. http://arxiv.org/abs/2404.10667 arXiv:2404.10667 [cs]
2024 arXiv
-
[77]
Yuelang Xu, Benwang Chen, Zhe Li, Hongwen Zhang, Lizhen Wang, Zerong Zheng, and Yebin Liu. 2024. Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians. In CVPR
2024
-
[78]
Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, and Haoqiang Fan. 2024. MegActor: Harness the Power of Raw Video for Vivid Portrait Animation. http://arxiv.org/abs/2405.20851 arXiv:2405.20851 [cs]
2024 arXiv
-
[79]
Williamson, Fei Chen, Fuzheng Yang, and Shidong Shang
Gaoxiong Yi, Wei Xiao, Yiming Xiao, Babak Naderi, Sebastian Möller, Wafaa Wardah, Gabriel Mittag, Ross Cutler, Zhuohuang Zhang, Donald S. Williamson, Fei Chen, Fuzheng Yang, and Shidong Shang. 2022. ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Asse...
2022
-
[80]
Kevin Yu, Gleb Gorbachev, Ulrich Eck, Frieder Pankratz, Nassir Navab, and Daniel Roth. 2021. Avatars for Teleconsulta- tion: Effects of Avatar Embodiment Techniques on User Perception in 3D Asymmetric Telepresence. IEEE Transactions on Visualization and Computer Graphics 27, 1...
2021
-
[81]
Eduard Zell, Carlos Aliaga, Adrian Jarabo, Katja Zibrek, Diego Gutierrez, Rachel McDonnell, and Mario Botsch. 2015. To stylize or not to stylize?: the effect of shape and material stylization on the perception of computer-generated faces. ACM Transactions on Graphics 34, 6 (No...
2015
-
[82]
Efros, Eli Shechtman, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Salt Lake City, UT, 586–595. https://doi....
2018
-
[83]
Mingyuan Zhou, Rakib Hyder, Ziwei Xuan, and Guojun Qi. 2024. UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures. In CVPR. , Vol. 1, No. 1, Article . Publication date: April 2025. A multidimensional measurement of photorealistic avata...
2024
-
[2014]
IEEE Transactions on Multimedia 16, 2 (Feb
Best Practices for QoE Crowdtesting: QoE Assessment With Crowdsourcing. IEEE Transactions on Multimedia 16, 2 (Feb. 2014), 541–558. https://doi.org/10.1109/TMM.2013.2291663
2014
-
[2024]
SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting. In CVPR. http://arxiv.org/abs/2403.05087
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.