{"id":"8bc59852-77ae-4c94-a016-1de714426b60","arxiv_id":"2411.18047","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"The study claims that specific eye and mouth proportions based on the baby schema effect make virtual agent faces look more trustworthy to older adults, but the reported optimal values are internally inconsistent.","lead":"This study generated 729 child-like virtual faces with different eye and mouth proportions and had older adults rate their trustworthiness. The authors propose a design recipe for trustworthy virtual caregivers for older adults, though the supporting data analysis has serious flaws.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The winning parameter vector is not uniquely defined: abstract, §4.3, and §4.4 report different optimal eye-spacing, mouth-height, and smile-arc values, so the central claim has no stable target.","rationale":"I choose the parameter-vector inconsistency as the load-bearing concern because it attacks the central claim before any statistical reanalysis is needed. The abstract, §4.3, and §4.4 give different optimal values for at least three of the six parameters; the §4.4 vector, which is the actual design paradigm, disagrees with the abstract's headline numbers. This is not a minor typo, since the disagreements are systematic across main effects, interactions, and the final verification, and §4.3's narrative about mouth height contradicts §4.3.1's table. The reader's weakest assumption about non-independence of ratings is also serious: with 729 faces each rated by only two participants, the ANOVA denominator df of 1457 is not credible, and a hierarchical reanalysis could change rankings. But even if that statistical issue were fixed, the paper would still need to state which parameter vector is the claimed optimum. I therefore keep the reader's REJECT verdict unchanged.","tokens_in":20678,"tokens_out":6055,"duration_ms":51471,"concrete_test":"Build an internal-consistency table from the manuscript text listing every passage that names a highest-credibility parameter vector (abstract, §4.2.2 interaction maxima, §4.3.1 main effects, §4.3.2 interactions, §4.4 final face). If the vectors do not agree component-wise, the central claim lacks a unique referent. A stronger check is to request the raw 1,458 ratings and recompute the top-ranked 3^6 combination under a mixed-effects model with participant and face-combination random effects; if the winning vector changes, the proposed paradigm is not reproducible from the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript's central claim is that a single set of proportions produces the highest facial credibility, but the text reports at least three incompatible 'winning' settings. The abstract selects eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H. Section 4.3.1 reports the highest credibility at eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H, and §4.3.2's three-way interaction repeats the 0.062H/0.43W/0.74H combination as highest. Section 4.4, which presents the 'face with the highest level of credibility' and the proposed design paradigm, instead uses eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. These are not cosmetic differences: they involve the specific output values that the design paradigm is meant to prescribe. Section 4.3's own text even states that a mouth height of 0.74H shows lower credibility than 0.77H and 0.79H, contradicting its own main-effect summary. Since no single vector is consistently identified as the optimum, the headline result is not stably defined and cannot be tested or applied as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests whether facial proportions associated with the baby schema affect older adults' trust in virtual humanoid agents. The authors generated 729 child-like faces by varying six features (eye size, eye height, eye spacing, mouth size, mouth height, smile arc) at three levels each, and 162 older adults rated subsets of nine faces on a five-item credibility scale. The abstract claims a single optimal proportion set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye distance 0.43W, mouth height 0.74H, smile arc 0.043H) and proposes this as a design paradigm. The body of the paper, however, reports different winning combinations across sections, and the statistical analysis treats nested ratings as independent observations, so the central claim is not stably defined or supported.","tokens_in":20926,"tokens_out":4268,"duration_ms":36610,"significance":"If the central claim were sound, the paper would provide practically useful design guidance for trustworthy virtual agents aimed at older adults, a population that is growing in smart-home and care contexts. The study has genuine strengths: a full-factorial stimulus design (729 faces), a targeted older adult sample, and an additional validation rating of perceived juvenility. These are commendable. However, the contradictory statements of the optimal parameters and the invalid treatment of the clustered data mean the paper currently does not establish any single proportion set as most trustworthy, and the proposed design paradigm is effectively a post-hoc description of one observed cell rather than a tested prediction.","major_comments":[{"comment":"The paper does not identify a single, stable winning combination. The abstract states that the highest credibility occurred at eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H. Section 4.3.1 states that the highest credibility was associated with eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H, and §4.3.2 repeats the 0.062H/0.43W/0.74H combination as the highest three-way interaction. Section 4.4, which presents the 'face with the highest level of credibility' and the design paradigm, instead uses eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. Because the headline result is a specific proportion vector, these discrepancies are load-bearing: they leave the central claim undefined and prevent a reader from testing or applying the proposed paradigm.","section":"Abstract; §4.3.1; §4.3.2; §4.4"},{"comment":"The main-effect summary for mouth height is internally contradictory. The text first states that 'the highest credibility was associated with a mouth height of 0.74H,' then later states that 'faces having a mouth height of 0.74H exhibiting lower credibility than faces with mouth heights of 0.77H and 0.79H. H5 was confirmed.' A single analysis cannot support both claims, and H5 pre-specified 0.77H, not 0.74H. This contradiction affects the interpretation of the mouth-height result and the confirmation reported for H5.","section":"§4.3.1"},{"comment":"The ANCOVA treats the 1,458 ratings as independent observations, with degrees of freedom near 1457. According to §3.4, each of the 162 participants rated nine faces, and each of the 729 face combinations was rated by two participants. Ratings are therefore nested within participants and within stimuli. The reported F-tests and p-values for the six-way ANCOVA and its interactions are not valid under this design; participant-level variance and stimulus-level clustering are ignored. This is a fundamental problem for all main-effect and interaction claims in Section 4, including the eye-spacing, mouth-height, and smile-arc effects that drive the conclusion.","section":"§3.4; §4.2; §4.3; Table 3"},{"comment":"The proposed design paradigm is not an independent validation but a restatement of the empirically highest-scoring cell from the same ANCOVA. The hypotheses pre-specify some factor levels (e.g., H4 at 0.41W, H5 at 0.77H, H6 at 0.043H), but the final §4.4 face also includes eye height 0.64H, which was not pre-specified (H3 was 0.62H), and the selection is made by inspecting the same data used to estimate the effects. As a result, the claimed 'design paradigm' overfits the sample and has no out-of-sample predictive test, so its generality is not established.","section":"§4.4; Hypotheses H1–H6"}],"minor_comments":[{"comment":"The sentence 'HI needed to be supported' should read 'H1 was not supported.'","section":"§4.2.1"},{"comment":"The table lists 'Eye width (EW)' as a factor, but the text and hypotheses refer to eye spacing; the relationship between 'eye width' and 'eye spacing' should be clarified to avoid confusion.","section":"Table 3"},{"comment":"The F-statistics for eye size and eye height are reported as identical (F(2,1455) = 14.32) despite different means and standard errors; this is likely a copy-paste error and should be corrected.","section":"§4.1"},{"comment":"References [42] and [43] appear to be the same source, and entries [43] and [44] duplicate each other; the reference list should be deduplicated and renumbered.","section":"References"},{"comment":"The caption contains the typo 'Da Mouth' and should be corrected to 'Large mouth.'","section":"Figure 9 caption"}],"recommendation":"reject","confidential_remarks":"The manuscript's central claim is not stably defined and the reported statistical analysis is inappropriate for the clustered design. These are load-bearing problems that would require a new experimental design or at least a full mixed-model reanalysis, which is beyond a routine revision. I therefore recommend rejection, although the topic and stimulus generation effort may provide a useful basis for future work if the design is revised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the study collects a substantial new dataset—729 child-like virtual faces, parameterized by six eye/mouth features, rated by 162 older adults—and the idea of calibrating baby-schema features for this population is worth pursuing. But as written, the paper's central claim is not stably defined. The abstract crowns eye distance 0.43W, mouth height 0.74H, smile arc 0.043H; §4.3.1 awards highest credibility to mouth height 0.74H and smile intensity 0.062H; §4.4's \"face with the highest credibility\" uses eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. §4.3's own narrative even says mouth height 0.74H scored lower than 0.77H and 0.79H. These are the precise values the proposed design paradigm is supposed to hand to practitioners, so the mismatch is load-bearing, not cosmetic.\n\nThe statistical analysis is the second soft spot. The text says each of 162 participants saw 9 random images, with 2 participants per subset of the 729 combinations. The ANCOVA tables report F-tests with error df around 1457, which treats the 1458 ratings as independent. Ratings are nested within participants and within the 729 cells; the \"highest credibility\" cell means rest on n=2. Without a mixed model or at least participant-level clustering, the reported p-values and the winning-cell selection are not reliable. Also missing: stimuli, data, code, and a non-baby-schema baseline condition. The design paradigm in §4.4 is a restatement of the empirically winning cell, not an independent prediction; the hypotheses pre-specify values that do not match the results sections.\n\nCredit where due: the full factorial sweep, the older-adult sample, the manipulation checks, the internal consistency (α=0.954), and the small juvenility validation study are all reasonable pieces of work. The research question is legitimate and the problem is real for HCI. But the internal contradictions alone would force a major revision, and the statistical treatment would need to be redone with proper random effects.\n\nRecommendation: reject the current version, but tell the authors the underlying study is worth resubmitting after (a) reconciling the conflicting optima, (b) reanalyzing with mixed-effects models, and (c) releasing stimuli and data. This deserves a serious referee, not a desk reject, because the dataset and question are genuinely useful.","headline":"A useful dataset and a real question, but the winning proportions shift between abstract, results, and design paradigm, and the statistics pool dependent ratings; needs major revision before it can be used.","tokens_in":21466,"tokens_out":4148,"would_cite":false,"duration_ms":34729,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that baby-schema eye and mouth proportions can be tuned to maximize older adults' perceived trustworthiness in virtual humanoid agents, and proposes a concrete design paradigm for those features.","keywords":["baby schema effect","virtual humanoid agents","trustworthiness","older adults","facial feature design","eye design","mouth design","human-agent interaction"],"falsifier":"Run a mixed-effects reanalysis with participant and face combination as random factors: if the mouth-size, eye-spacing, mouth-height, and smile-arc effects disappear, or the winning combination moves outside the reported range, the fixed-effect ANCOVA result is an artifact. Alternatively, present the Section 4.4 highest-credibility face and the abstract's competing face to a fresh sample in pairwise forced choice; the paradigm is falsified if the Section 4.4 face does not win significantly more often than chance.","tokens_in":20385,"feed_emoji":"🤖","tokens_out":7510,"duration_ms":62240,"temperature":0.7,"pith_summary":"This paper tries to establish that the \"baby schema\" effect, the cluster of infant-like facial proportions that make faces seem cute and trustworthy, can be carried over to the design of virtual humanoid agents for older adults, and that the eyes and mouth are the features that matter most. To test this, the author built a base child-like virtual face, varied six eye and mouth dimensions (eye size, eye height, eye spacing, mouth size, mouth height, and smile arc), each at three levels, and had 162 older adults rate the credibility of 729 resulting faces. The results point to a specific winning combination of proportions, proposed as a design paradigm for trusted virtual caregivers. One caution is that the paper states the winning combination differently in the abstract (eye spacing 0.43W, mouth height 0.74H) and in Section 4.4 (eye spacing 0.41W, mouth height 0.77H), so the exact paradigm is not stably pinned down. The significance, if the finding holds, is a concrete, parameter-level guideline for making assistive agents more acceptable to an aging population.","feed_headline":"Baby-schema face proportions boost trust in virtual caregivers","feed_subtitle":"A 162-person study pinpoints eye and mouth ratios that maximize credibility for older adults.","key_machinery":"The machinery is a controlled six-factor stimulus ladder built on a base 4 to 6 year-old female face whose facial width-to-height ratio (1.90) was checked against the baby schema standard (1.96). Each of six features, eye size, eye height, eye spacing, mouth size, mouth height, and smile arc, varied over three levels expressed as fractions of face width W or face height H, generating $3^{6}$ = 729 faces. Older adults rated each face on a five-item credibility scale, and an ANCOVA with gender and smartphone experience as covariates was used to rank the levels and identify the best combination. The proposed \"facial design paradigm\" is the set of winning normalized proportions, intended to be used directly by designers of virtual humanoid agents.","core_discovery":"On the paper's own terms, the central discovery is that perceived credibility of a child-like virtual agent among older adults is not flat across facial proportions: mouth size, eye spacing, mouth height, and smile arc each produced significant main effects in a six-way ANCOVA, while eye size and eye height did not. Interactions further shaped the ratings, with the most credible face in Section 4.4 having eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye spacing 0.41W, mouth height 0.77H, and smile arc 0.043H, on a face whose width-to-height ratio already matches the baby schema. The abstract reports a different winning set (eye spacing 0.43W, mouth height 0.74H), so the paper's concrete recommendation is internally inconsistent even though its directional claim, that baby-schema-consistent eyes and mouth raise trust, is consistent throughout. The study then validates that the top-rated faces are perceived as more childlike than the lowest-rated ones, tying the trust result to the baby schema mechanism.","pith_inferences":["Beyond the paper: because the abstract and Section 4.4 disagree on the two most salient parameters, the safest reading is that the broad direction (baby-schema eyes and mouth raise trust) is supported while the precise numbers are not; a pre-registered replication that tests the two competing faces head-to-head would settle the paradigm.","Beyond the paper: the design could be tested in a real smart-home interaction, such as medication reminders, to see whether the trust advantage in ratings changes willingness to follow advice; rating studies can overstate first-impression effects.","Beyond the paper: since the sample is Chinese older adults with varied smartphone experience, the same stimulus set could be run with younger adults or other cultural groups to decompose the baby-schema response from age-specific and culture-specific preferences."],"forward_implications":["If the winning proportions hold, designers of virtual caregivers for older adults can start from a concrete parameter set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, mouth height near 0.77H, smile arc 0.043H) instead of relying on intuition.","Mouth dimensions do more work than eye dimensions: mouth size and mouth height had strong main effects, whereas eye size did not, suggesting that design effort should go to the mouth first.","Because eye size mattered mainly in interaction with other features, the paper implies that facial features should not be tuned independently; a configural approach to face design is more promising.","The validation results imply that \"looks like a child\" and \"looks trustworthy\" move together for older adults, so perceived juvenility can serve as a proxy target in agent design."],"supporting_citations":[{"why":"It defines the baby schema effect that motivates the facial proportions manipulated in the study.","marker":"Lorenz, 1943"},{"why":"It shows that facial features of virtual characters alter viewer perception and moral choices, justifying feature-level manipulation.","marker":"Ferstl, 2017"},{"why":"It establishes facial anthropomorphic trustworthiness in social robots, the framework the study extends to older adults.","marker":"Song, 2021"},{"why":"It supplies the five-item credibility scale used as the dependent measure in the experiment.","marker":"Gorn, 2008"},{"why":"It provides the child-face image source and evidence that baby schema evokes trust in different age groups.","marker":"Luo, 2020"},{"why":"It provides background for the eye-size hypothesis, showing that eye size affects cuteness perceptions across expressions and ages.","marker":"Yao, 2022"},{"why":"It shows that eyes and mouth are critical regions for facial emotion and trust judgments, focusing the manipulation on those features.","marker":"Adolphs, R., 2002"},{"why":"It describes validated virtual human facial expressions used in building the stimulus set.","marker":"Garcia, 2020"}],"fun_headline_variants":["Baby-schema face proportions boost trust in virtual agents for seniors","Childlike eyes and mouth on AI agents earn older adults' trust","Baby-faced virtual agents make seniors trust them more","Baby-schema facial cues lift trust in virtual caregivers for older adults"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The statistical analysis treats each of the 1,458 ratings as an independent observation even though only two participants saw each face combination, so the reported F-tests assume that ratings do not cluster by participant or by the nine-face subset.","fun_headline_variants_meta":{"raw":{"variants":["Baby-schema face proportions boost trust in virtual agents for seniors","Childlike eyes and mouth on AI agents earn older adults' trust","Baby-faced virtual agents make seniors trust them more","Baby-schema facial cues lift trust in virtual caregivers for older adults"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000881,"raw_usage":{"total_tokens":3870,"prompt_tokens":1070,"completion_tokens":2800,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":2730}},"tokens_in":686,"tokens_out":2800,"duration_ms":18198,"temperature":1.0,"reasoning_tokens":2730,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:34:19.102657+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a mixed-effects reanalysis with participant and face combination as random factors: if the mouth-size, eye-spacing, mouth-height, and smile-arc effects disappear, or the winning combination moves outside the reported range, the fixed-effect ANCOVA result is an artifact. Alternatively, present the Section 4.4 highest-credibility face and the abstract's competing face to a fresh sample in pairwise forced choice; the paradigm is falsified if the Section 4.4 face does not win significantly more often than chance.","supporting_citations":[],"review_version":1}