{"id":"4d5a73c5-437d-4481-b009-ba6f1a2c7f3c","arxiv_id":"2411.09066","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A validated open-source subjective framework measures ten avatar quality dimensions and finds standard objective metrics miss most of what people perceive, with no uncanny valley for photorealistic avatars in telecom clips.","lead":"Photorealistic avatars are usually judged by image-quality scores like PSNR and FID, which this paper shows do not match how people actually feel about them. The authors release an open-source, crowdsourced test that rates avatars on ten human-centered dimensions, and find no uncanny valley for photorealistic avatars in telecom-style clips.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-uncanny-valley conclusion is underdetermined: Figure 10 fits independent-looking clip points from clustered avatar models with sparse near-photorealistic sampling, while Tables 9–10 show the dimensions are near-collinear, so the linear affinity–realism result may be a halo artifact rather than…","rationale":"I read the paper as making two intertwined claims: (1) the open-source crowdsourced framework accurately and reproducibly measures avatar QoE and that standard objective metrics miss most of it; and (2) a stronger perceptual claim that there is no uncanny valley for photorealistic avatars in telecommunication scenarios, based on the linear affinity–realism relationship. The first claim is well supported: the framework is validated against expert panels (average PCC=0.904, Table 6) and across repeated runs (average PCC=0.987 at the condition level, Table 7), and the objective-metric correlations are weak across eight models, though the sample is small. The second claim is where the argument is least secure. The reader's weakest-assumption analysis correctly identifies the near-perfect inter-item correlations as evidence of a possible halo effect. I agree that this threatens the multidimensional interpretation, but I think the most load-bearing consequence is the no-uncanny-valley inference: if affinity and realism are both driven by a single overall impression, their linear relationship is expected and cannot falsify a localized dip. In addition, the clip-level scatter in Figure 10 overcounts independence, and the near-photorealistic region is likely under-sampled. A targeted model-level, mixed-effects re-analysis with a quadratic term and explicit sampling-density reporting would settle whether the linearity is robust. Because the paper already receives a CONDITIONAL verdict and my concern is a reason to keep that conditional status rather than to reject the framework, I recommend UNCHANGED. I have not identified an internal inconsistency or evidence of bad faith; the issue is the strength of a strong empirical inference relative to the data actually presented.","tokens_in":25690,"tokens_out":6682,"duration_ms":77480,"concrete_test":"Recompute the Section 5.5 analysis at the avatar-model level (average clips per model) using a linear mixed-effects model with a random intercept per avatar model and a quadratic term in realism; report the number of distinct avatar models in each 0.5 MOS realism bin between 3 and 5. If the quadratic term is insignificant and the confidence band excludes a dip of the size predicted by uncanny-valley meta-analyses, the no-uncanny-valley claim survives. Additionally, re-run the analysis on an independent set of at least 10 avatars spanning realism 3–5 with at least 3 clips each, so clip-level clustering cannot drive the apparent linearity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing inference is the Section 5.5 claim that there is no uncanny valley for photorealistic avatars, stated in the abstract and repeated in the developer guidelines. Figure 10 plots affinity versus realism for the clips used in Figure 7 and reports R^2=0.966 for a linear fit. Those are clip-level points, so multiple clips from the same avatar model are treated as independent; a line through clustered points can be dominated by a few models, and it cannot rule out a localized dip if the near-photorealistic range (roughly realism 3.5–4.5 on the 1–5 MOS scale) is sparsely sampled. The conclusion is also entangled with the halo effect the paper itself documents: Table 9 shows Template A dimensions correlating at average PCC=0.996, and Table 10 shows PCC=0.999 for realism>2. If raters anchor on a single overall impression, affinity and realism are near-collinear by construction, so a linear affinity–realism relationship is expected and supplies little independent evidence about Mori's dip. The paper's own limitations (passive 2D testing, N=19 avatar models) further narrow the scope, yet the abstract and developer guidelines state the no-uncanny-valley result without those caveats. This concern does not undercut the framework's demonstrated reproducibility or the weak objective-metric correlations, but it does undercut the strongest perceptual claim in the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an open-source crowdsourced test framework, based on the authors' P.910 implementation, for measuring ten subjective dimensions of photorealistic avatar quality of experience: realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness, formality, affinity, resemblance, and emotion accuracy. The framework is validated against VQEG laboratory video-quality data and against a panel of experts, and its run-to-run reproducibility is reported. Using 19 avatar models and real-video baselines, the authors find that the subjective dimensions are weakly correlated with common objective metrics (PSNR, SSIM, LPIPS, FID, FVD), that eight Template A dimensions become nearly collinear for avatars with realism above a threshold, and that affinity is linearly related to realism, which they interpret as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.","tokens_in":25922,"tokens_out":7258,"duration_ms":70013,"significance":"If the claims hold, this is a useful contribution: an open, validated, reproducible subjective testing methodology for avatar quality of experience, a public test set, and evidence that objective metrics currently used in avatar development do not capture the usability dimensions that matter. The framework validation is genuinely strong: run-to-run clip-level correlations are high for Template A (Table 8, average PCC 0.962), model-level reproducibility is above 0.98 (Table 7), accuracy against an expert panel averages PCC 0.904 (Table 6), and the crowdsourcing implementation is validated against external VQEG lab data (Tables 2-4). The main risk is that the two most prominent perceptual conclusions, the dimensionality reduction and the absence of an uncanny valley, rest on near-perfect inter-item correlations that may reflect a general impression halo rather than distinct underlying constructs, and the no-uncanny-valley conclusion is drawn from a linear fit on clustered clip-level points without a direct test for a local dip.","major_comments":[{"comment":"I would ask the authors to add an avatar-random-effect analysis at the model level, a direct test for a local dip in the near-photorealistic range, confidence intervals on the fit, or to substantially soften the conclusion and restrict it to the range actually sampled.","section":"§5.5, Fig. 10"},{"comment":"Since the realism > 2 threshold is chosen post hoc and the claim directly feeds the developer guideline that 'only 3 dimensions need to be measured' (§6.2), the paper should provide stronger evidence of discriminant validity, for example a confirmatory factor analysis with a specified measurement model, or should reframe the claim as a practical shortcut rather than a statement about the underlying psychological dimensions.","section":"§5.3, Tables 9-11"},{"comment":"Please report confidence intervals or bootstrap intervals for Table 12, and adjust the abstract and conclusions to reflect the face-limited results.","section":"§5.4, Tables 12-13"}],"minor_comments":[{"comment":"The text says the average clip-level run-run PCC is 0.968, but Table 8 gives Template A averages of 0.962 and a Template B emotion-accuracy average of 0.870; the text should match the table, and the lower emotion-accuracy reproducibility should be acknowledged as a caveat to the 'highly reproducible' claim.","section":"§4.3"},{"comment":"The first sentence lists 'PSNR, SSIM, LPIPS, FIV, FVD'; this should be FID, not FIV.","section":"§6"},{"comment":"The open-source tool is named P.910, identical to the ITU-T Recommendation P.910. This naming is likely to confuse readers; consider renaming the repository or adding a clarifying sentence in Section 3.1.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The framework validation is solid and the open-source release is a genuine strength. The main risk is that the abstract and developer guidelines overstate the no-uncanny-valley and dimensionality-reduction conclusions relative to the statistical evidence, which is underdetermined by the clustered, near-collinear data. If the authors add the requested model-level and nonlinearity analyses and adjust the abstract to match the evidence, the paper would be publishable; without those changes, the strongest perceptual claim is not supported by the current analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ross,\n\nThis one is worth reading for the framework alone. The authors ship an open-source crowdsourced implementation of P.910 with rater, environment, hardware, and network qualifications, validate it against VQEG lab data and an expert panel, and show run-to-run correlations around 0.96-0.99. That is real, reproducible infrastructure. The finding that PSNR, SSIM, LPIPS, FID, and FVD correlate weakly with nine of ten subjective dimensions, with only emotion accuracy reaching moderate correlation, is a useful empirical result that argues against optimizing avatars with those metrics alone. The inclusion of a real-video baseline and 19 avatar models, some photorealistic, is a step up from most prior avatar studies.\n\nThe soft spot is the uncanny valley conclusion. The abstract and developer guidelines state there is no uncanny valley for photorealistic avatars in telecommunication scenarios, but the evidence is underdetermined. Figure 10 fits a line through clip-level points from a small number of avatar models; multiple clips from the same model are not independent, the near-photorealistic range is sparsely sampled, and a linear fit cannot rule out a localized dip. More importantly, the paper itself documents near-collinearity among the Template A dimensions (average PCC=0.996, 0.999 for realism>2). That is exactly what a halo effect would produce, so the strong linear affinity-realism relationship may simply reflect raters anchoring on one overall impression rather than independent evidence about Mori's dip. The dimensionality reduction to three items has the same problem. The authors mention they tested three survey formats and saw the same high correlations, but that does not rule out halo; it suggests the effect is robust across presentation modes.\n\nThe objective-metric correlations are also computed on only eight models with no confidence intervals, which makes the exact magnitudes fragile even if the weak-correlation direction is believable.\n\nThis is not a rejection. The framework is a genuine contribution and the weak objective-subjective correlation result deserves to be published. But the strong perceptual claim should be tempered or re-analyzed with mixed-effects models, per-model clustering, and denser sampling of the near-photorealistic range. I'd send it to review and ask for those changes.\n\nFor you: cite the framework if you do QoE or avatar work; bring it to reading group if you want a good discussion of halo effects in subjective testing.\n\nBest,","headline":"Solid open-source avatar QoE framework with a credible negative result on objective metrics, but the no-uncanny-valley claim exceeds what the data can support.","tokens_in":26542,"tokens_out":2003,"would_cite":true,"duration_ms":20156,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that subjective human ratings are the only reliable way to measure photorealistic avatar quality of experience: standard objective metrics correlate weakly with nine of ten measured dimensions and only moderately with…","keywords":["photorealistic avatars","quality of experience","subjective testing","crowdsourcing","uncanny valley","objective metrics","telecommunication","avatar realism"],"falsifier":"Have separate groups of raters each judge only one dimension of the same avatar clips, then compare those single-dimension scores with the all-items-at-once scores. If trust, comfort, and appropriateness no longer track realism when realism is not among the questions asked — or if the average inter-item correlation drops well below 0.996 — then the reported high correlations are a halo or order effect, and the paper's dimensionality reduction and no-uncanny-valley linearity would not survive independent measurement.","tokens_in":25411,"feed_emoji":"🧑💻","tokens_out":11849,"duration_ms":103500,"temperature":0.7,"pith_summary":"Photorealistic avatars are usually developed and compared using pixel- and feature-based metrics, but this paper argues those metrics cannot tell you whether an avatar can be trusted, feels comfortable to work with, or creeps people out. It contributes an open-source, crowdsourced test framework that measures avatar quality of experience along ten human-centered dimensions, and shows the framework is accurate against a panel of experts and reproducible across repeated runs. Applied to 19 avatar models spanning photorealistic to cartoon-like, the measurements reveal that PSNR, SSIM, LPIPS, FID, and FVD correlate weakly with nine of the ten dimensions and only moderately with emotion accuracy. The paper also reports that for avatars above a realism threshold, eight dimensions move together, so the survey can be shortened from ten items to three. Finally, it finds affinity grows linearly with realism and concludes that for photorealistic avatars in telecommunication there is no uncanny valley; lower-realism avatars are simply worse on trust, comfort, appropriateness, formality, and creepiness than real video.","feed_headline":"Objective metrics miss 9 of 10 avatar quality dimensions","feed_subtitle":"A crowdsourced test of 19 avatars shows PSNR, SSIM, and LPIPS cannot track trust, creepiness, or work appropriateness; realism drives…","key_machinery":"The load-bearing mechanism is the open-source crowdsourced subjective test framework, an implementation of the P.910 subjective video quality assessment standard extended to avatars, with rater qualification, display calibration, gold clips, trapping items, and repeated items to control data quality. It uses two survey templates: Template A asks for agreement with eight statements about the avatar alone, and Template B shows the avatar side-by-side with the real person to rate resemblance and emotion accuracy. The framework's output is a per-clip and per-condition mean opinion score for each dimension, and its validation against laboratory video studies and an expert panel is what gives the correlations their weight. The analytic load is carried by inter-dimension Pearson correlations, a principal component analysis that finds two components explaining 94% of the variance, and a linear fit of affinity against realism whose R²=0.966 is the basis for the no-uncanny-valley conclusion.","core_discovery":"The central discovery is that the quality of experience of a photorealistic avatar is a multidimensional subjective quantity that standard objective metrics do not capture. Using a validated crowdsourced survey, the paper measures ten dimensions — realism, trust, comfortableness using, comfortableness interacting with, appropriateness for work, creepiness (reported as not creepy), formality, affinity, resemblance to the person, and emotion accuracy — across 19 avatar models. The inter-item correlations among the eight Template A dimensions average PCC=0.996, and for avatars with realism above 2 on the 1–5 scale they average 0.999, which the paper interprets as a strong redundancy that permits reducing the ten dimensions to three via principal component analysis. Against PSNR, SSIM, LPIPS, FID, and FVD, the subjective dimensions show weak correlation, with only emotion accuracy reaching moderate correlation (e.g., PSNR=0.63, LPIPS=-0.67). In addition, affinity versus realism is well described by a straight line (R²=0.966), which the paper takes as evidence that there is no uncanny valley for photorealistic avatars in the telecommunication scenario.","pith_inferences":["A testable extension the paper leaves implicit: if the near-perfect inter-item correlations are a halo effect, single-dimension ratings from independent groups will diverge from the all-items survey; running that comparison is a direct check on whether the ten dimensions are really separate constructs.","The no-uncanny-valley conclusion is tied to the stimulus set used here — 2D talking-head clips in telecommunication poses; the same survey applied to full-body avatars, VR/AR displays, or avatars with motion artifacts could still find a dip the paper does not test.","Because emotion accuracy and resemblance show moderate-to-strong correlation with PSNR, SSIM, and LPIPS when computed on the face region, a learned perceptual metric trained on subjective avatar ratings could plausibly close much of the objective-subjective gap, a direction the paper mentions as future work.","The framework is a passive-viewing test; an interactive variant in which a rater converses with a live avatar could change trust and comfort ratings, since the paper's own reproducibility and accuracy claims rest on passive observation."],"forward_implications":["Avatar developers who optimize only PSNR, SSIM, LPIPS, FID, or FVD will leave most of the quality-of-experience space unmeasured, so subjective testing becomes necessary for usability claims.","For avatars with realism above 2 on the 1–5 scale, one Template A item (e.g., realism) plus the two Template B items (emotion accuracy and resemblance) is enough, shrinking the survey from ten items to three.","Avatars that are less realistic than real video will score lower on trust, comfort using, comfort interacting, work appropriateness, formality, and affinity, and higher on creepiness; matching real video requires matching its realism.","Because affinity rises linearly with realism (R²=0.966), designers of telecommunication avatars should maximize realism rather than stop short of a presumed uncanny valley.","Some state-of-the-art photorealistic avatars approach real video on these subjective dimensions, but none reaches real-video realism, so the gap between avatars and real video remains measurable and large."],"supporting_citations":[{"why":"Supplies the standardized subjective video quality methodology on which the crowdsourced framework is built and which it extends with rater, environment, hardware, and network qualifications.","marker":"[28]"},{"why":"Provides the original avatar evaluation dimensions (formality, non-creepiness, realism, resemblance, appropriateness, comfortable using and interacting) that Template A extends.","marker":"[22]"},{"why":"Defines the quality-of-experience dimensions for extended-reality meetings, including realism, trust, and facial expression, that motivate the added survey items.","marker":"[26]"},{"why":"Describes the crowdsourcing implementation backbone, including gold clips, trapping items, and result parsing, on which the avatar survey is built.","marker":"[45]"},{"why":"Supplies the laboratory HDTV subjective datasets used to validate that the crowdsourced implementation is accurate and reproducible.","marker":"[70]"},{"why":"States the uncanny valley hypothesis that the paper tests through the affinity-versus-realism relation.","marker":"[43]"},{"why":"Defines LPIPS, one of the objective perceptual metrics the paper finds weakly correlated with nine of the ten subjective dimensions.","marker":"[82]"},{"why":"Defines FID, another objective metric whose correlation with the subjective dimensions is tested and found weak except in the face-only region.","marker":"[18]"},{"why":"Defines FVD, the video-level objective metric included in the correlation analysis with the subjective scores.","marker":"[68]"}],"fun_headline_variants":["9 of 10 avatar quality metrics elude PSNR and SSIM","Photorealistic avatars: no uncanny valley in telecom","Subjective scores beat objective metrics for avatar quality","Trust, creepiness, affinity: objective metrics miss them all","Avatar realism drives 8 quality dimensions in one shot"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ten survey items are treated as independently meaningful dimensions, but the near-perfect inter-item correlations (average PCC=0.996) are exactly what a single global impression would look like, so if raters are not actually distinguishing the dimensions, the dimensionality reduction and the linear affinity–realism relation would be artifacts of rating co-movement rather than evidence about the underlying quality structure.","fun_headline_variants_meta":{"raw":{"variants":["9 of 10 avatar quality metrics elude PSNR and SSIM","Photorealistic avatars: no uncanny valley in telecom","Subjective scores beat objective metrics for avatar quality","Trust, creepiness, affinity: objective metrics miss them all","Avatar realism drives 8 quality dimensions in one shot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1462,"prompt_tokens":1132,"completion_tokens":330,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":748,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":748,"tokens_out":330,"duration_ms":3754,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:06:15.676274+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have separate groups of raters each judge only one dimension of the same avatar clips, then compare those single-dimension scores with the all-items-at-once scores. If trust, comfort, and appropriateness no longer track realism when realism is not among the questions asked — or if the average inter-item correlation drops well below 0.996 — then the reported high correlations are a halo or order effect, and the paper's dimensionality reduction and no-uncanny-valley linearity would not survive independent measurement.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the standardized subjective video quality methodology on which the crowdsourced framework is built and which it extends with rater, environment, hardware, and network qualifications."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the quality-of-experience dimensions for extended-reality meetings, including realism, trust, and facial expression, that motivate the added survey items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the laboratory HDTV subjective datasets used to validate that the crowdsourced implementation is accurate and reproducible."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines FID, another objective metric whose correlation with the subjective dimensions is tested and found weak except in the face-only region."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines FVD, the video-level objective metric included in the correlation analysis with the subjective scores."}],"review_version":1}