REVIEW 4 major objections 5 minor 2 references
The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that baby-schema eye and mouth proportions can be tuned to maximize older adults' perceived trustworthiness in virtual humanoid agents, and proposes a concrete design paradigm for those features.
desk verdict A useful dataset and a real question, but the winning proportions shift between abstract, results, and design paradigm, and the statistics pool dependent ratings; needs major revision before it can be used. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a controlled six-factor stimulus ladder built on a base 4 to 6 year-old female face whose facial width-to-height ratio (1.90) was checked against the baby schema standard (1.96). Each of six features, eye size, eye height, eye spacing, mouth size, mouth height, and smile arc, varied over three levels expressed as fractions of face width W or face height H, generating $3^{6}$ = 729 faces. Older adults rated each face on a five-item credibility scale, and an ANCOVA with gender and smartphone experience as covariates was used to rank the levels and identify the best combination. The proposed "facial design paradigm" is the set of winning normalized proportions, intended to be used directly by designers of virtual humanoid agents.
What would settle it
Run a mixed-effects reanalysis with participant and face combination as random factors: if the mouth-size, eye-spacing, mouth-height, and smile-arc effects disappear, or the winning combination moves outside the reported range, the fixed-effect ANCOVA result is an artifact. Alternatively, present the Section 4.4 highest-credibility face and the abstract's competing face to a fresh sample in pairwise forced choice; the paradigm is falsified if the Section 4.4 face does not win significantly more often than chance.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that perceived credibility of a child-like virtual agent among older adults is not flat across facial proportions: mouth size, eye spacing, mouth height, and smile arc each produced significant main effects in a six-way ANCOVA, while eye size and eye height did not. Interactions further shaped the ratings, with the most credible face in Section 4.4 having eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye spacing 0.41W, mouth height 0.77H, and smile arc 0.043H, on a face whose width-to-height ratio already matches the baby schema. The abstract reports a different winning set (eye spacing 0.43W, mouth height 0.74H), so the paper's concrete recommendation is internally inconsistent even though its directional claim, that baby-schema-consistent eyes and mouth raise trust, is consistent throughout. The study then validates that the top-rated faces are perceived as more childlike than the lowest-rated ones, tying the trust result to the baby schema mechanism.
Load-bearing premise
The statistical analysis treats each of the 1,458 ratings as an independent observation even though only two participants saw each face combination, so the reported F-tests assume that ratings do not cluster by participant or by the nine-face subset.
Editorial extensions
If this is right
- If the winning proportions hold, designers of virtual caregivers for older adults can start from a concrete parameter set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, mouth height near 0.77H, smile arc 0.043H) instead of relying on intuition.
- Mouth dimensions do more work than eye dimensions: mouth size and mouth height had strong main effects, whereas eye size did not, suggesting that design effort should go to the mouth first.
- Because eye size mattered mainly in interaction with other features, the paper implies that facial features should not be tuned independently; a configural approach to face design is more promising.
- The validation results imply that "looks like a child" and "looks trustworthy" move together for older adults, so perceived juvenility can serve as a proxy target in agent design.
Reading between the lines
- Beyond the paper: because the abstract and Section 4.4 disagree on the two most salient parameters, the safest reading is that the broad direction (baby-schema eyes and mouth raise trust) is supported while the precise numbers are not; a pre-registered replication that tests the two competing faces head-to-head would settle the paradigm.
- Beyond the paper: the design could be tested in a real smart-home interaction, such as medication reminders, to see whether the trust advantage in ratings changes willingness to follow advice; rating studies can overstate first-impression effects.
- Beyond the paper: since the sample is Chinese older adults with varied smartphone experience, the same stimulus set could be run with younger adults or other cultural groups to decompose the baby-schema response from age-specific and culture-specific preferences.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper tests whether facial proportions associated with the baby schema affect older adults' trust in virtual humanoid agents. The authors generated 729 child-like faces by varying six features (eye size, eye height, eye spacing, mouth size, mouth height, smile arc) at three levels each, and 162 older adults rated subsets of nine faces on a five-item credibility scale. The abstract claims a single optimal proportion set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye distance 0.43W, mouth height 0.74H, smile arc 0.043H) and proposes this as a design paradigm. The body of the paper, however, reports different winning combinations across sections, and the statistical analysis treats nested ratings as independent observations, so the central claim is not stably defined or supported.
Significance. If the central claim were sound, the paper would provide practically useful design guidance for trustworthy virtual agents aimed at older adults, a population that is growing in smart-home and care contexts. The study has genuine strengths: a full-factorial stimulus design (729 faces), a targeted older adult sample, and an additional validation rating of perceived juvenility. These are commendable. However, the contradictory statements of the optimal parameters and the invalid treatment of the clustered data mean the paper currently does not establish any single proportion set as most trustworthy, and the proposed design paradigm is effectively a post-hoc description of one observed cell rather than a tested prediction.
major comments (4)
- [Abstract; §4.3.1; §4.3.2; §4.4] The paper does not identify a single, stable winning combination. The abstract states that the highest credibility occurred at eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H. Section 4.3.1 states that the highest credibility was associated with eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H, and §4.3.2 repeats the 0.062H/0.43W/0.74H combination as the highest three-way interaction. Section 4.4, which presents the 'face with the highest level of credibility' and the design paradigm, instead uses eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. Because the headline result is a specific proportion vector, these discrepancies are load-bearing: they leave the central claim undefined and prevent a reader from testing or applying the proposed paradigm.
- [§4.3.1] The main-effect summary for mouth height is internally contradictory. The text first states that 'the highest credibility was associated with a mouth height of 0.74H,' then later states that 'faces having a mouth height of 0.74H exhibiting lower credibility than faces with mouth heights of 0.77H and 0.79H. H5 was confirmed.' A single analysis cannot support both claims, and H5 pre-specified 0.77H, not 0.74H. This contradiction affects the interpretation of the mouth-height result and the confirmation reported for H5.
- [§3.4; §4.2; §4.3; Table 3] The ANCOVA treats the 1,458 ratings as independent observations, with degrees of freedom near 1457. According to §3.4, each of the 162 participants rated nine faces, and each of the 729 face combinations was rated by two participants. Ratings are therefore nested within participants and within stimuli. The reported F-tests and p-values for the six-way ANCOVA and its interactions are not valid under this design; participant-level variance and stimulus-level clustering are ignored. This is a fundamental problem for all main-effect and interaction claims in Section 4, including the eye-spacing, mouth-height, and smile-arc effects that drive the conclusion.
- [§4.4; Hypotheses H1–H6] The proposed design paradigm is not an independent validation but a restatement of the empirically highest-scoring cell from the same ANCOVA. The hypotheses pre-specify some factor levels (e.g., H4 at 0.41W, H5 at 0.77H, H6 at 0.043H), but the final §4.4 face also includes eye height 0.64H, which was not pre-specified (H3 was 0.62H), and the selection is made by inspecting the same data used to estimate the effects. As a result, the claimed 'design paradigm' overfits the sample and has no out-of-sample predictive test, so its generality is not established.
minor comments (5)
- [§4.2.1] The sentence 'HI needed to be supported' should read 'H1 was not supported.'
- [Table 3] The table lists 'Eye width (EW)' as a factor, but the text and hypotheses refer to eye spacing; the relationship between 'eye width' and 'eye spacing' should be clarified to avoid confusion.
- [§4.1] The F-statistics for eye size and eye height are reported as identical (F(2,1455) = 14.32) despite different means and standard errors; this is likely a copy-paste error and should be corrected.
- [References] References [42] and [43] appear to be the same source, and entries [43] and [44] duplicate each other; the reference list should be deduplicated and renumbered.
- [Figure 9 caption] The caption contains the typo 'Da Mouth' and should be corrected to 'Large mouth.'
Circularity Check
No circular derivation: the paper's central claim is an empirically selected result, not a prediction that reduces to its own inputs; the reported inconsistencies across sections are statistical and reporting problems, not circularity.
full rationale
This paper contains no formal derivation chain in which an output is constructed from its own inputs. The baby-schema parameter values used to build the stimuli are taken from external literature (e.g., Lorenz 1943; Luo 2020; Ferstl 2017), not from the present experiment, so there is no self-definitional loop in which 'baby schema' is defined by the credibility scores. The hypotheses H1-H6 are stated before analysis in Section 3.1, and the results are obtained from ANCOVA on participant ratings. The later 'design paradigm' in Section 4.4 is a restatement of the empirically selected winning face, which is ordinary empirical inference rather than a fitted parameter renamed as an independent prediction. There are no load-bearing self-citations and no imported uniqueness theorem. The paper does contain a serious internal inconsistency: the abstract reports the winner as eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H; Section 4.3.1 and 4.3.2 report the highest credibility at eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H; and Section 4.4 selects eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. This instability, along with the contradictory statement in Section 4.3 that 0.74H mouth height has lower credibility than 0.77H and 0.79H, undermines the reliability and interpretability of the headline result. However, inconsistency and post-hoc selection are not the same as circularity, and no equation or constructed equivalence can be exhibited that makes the result equal to its input by definition. The appropriate circularity verdict is therefore no significant circularity, while the manuscript's correctness risk is high.
Assumptions & free parameters
free parameters (6)
- Eye size optimum =
0.25W
- Mouth size optimum =
0.27W
- Eye height optimum =
0.64H
- Eye distance optimum =
0.43W (abstract) / 0.41W (Section 4.4)
- Mouth height optimum =
0.74H (abstract) / 0.77H (Section 4.4)
- Smile arc optimum =
0.043H (abstract) / 0.062H (Section 4.3.1)
assumptions (5)
- domain assumption The baby schema effect causes increased perceived trustworthiness in older adults
- domain assumption The six geometric ratios capture the relevant aspects of baby schema for trust
- ad hoc to paper The frontal view of the 3D character faithfully represents the intended feature proportions
- domain assumption Five Likert items on reliability, sincerity, honesty, trustworthiness, and persuasiveness measure facial credibility
- ad hoc to paper The 1,458 ratings from 162 participants can be treated as independent observations in the ANCOVA
Cite this review
Pith. "Pith review of The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness." pith.science (2026). https://pith.science/paper/V6WV72IQ
@misc{pith2026241118047,
author = {Pith},
title = {Pith review of: The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness},
year = {2026},
howpublished = {\url{https://pith.science/paper/V6WV72IQ}},
note = {Machine review of arXiv:2411.18047}
}
read the original abstract
The increasing proportion of the older adult population has made the smart home care industry one of the critical markets for virtual human-like agents. It is crucial to effectively promote a trustworthy human-computer partnership with older adults, enhancing service acceptance and effectiveness. However, few studies have focused on the facial features of the agents themselves, where the "baby schema" effect plays a vital role in enhancing trustworthiness. The eyes and mouth, in particular, attract most of the audience's attention and are especially significant. This study explores the impact of eye and mouth design on users' perception of trustworthiness. Specifically, a virtual humanoid agents model was developed, and based on this, 729 virtual facial images of children were designed. Participants (N=162) were asked to evaluate the impact of variations in the size and positioning of the eyes and mouth regions on the perceived credibility of these virtual agents. The results revealed that when the facial aspect ratio (width and height denoted as W and H, respectively) aligned with the "baby schema" effect (eye size at 0.25W, mouth size at 0.27W, eye height at 0.64H, eye distance at 0.43W, mouth height at 0.74H, and smile arc at 0.043H), the virtual agents achieved the highest facial credibility. This study proposes a design paradigm for the main facial features of virtual humanoid agents, which can increase the trust of older adults during interactions and significantly contribute to the research on the trustworthiness of virtual humanoid agents.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[20]
L., Hogstrom, J., Sundstrom, K., Frick, A., Serlachius, E
Kleberg, J. L., Hogstrom, J., Sundstrom, K., Frick, A., Serlachius, E. (2021). Delayed gaze shifts away from others' eyes in children and adolescents with social anxiety disorder. Journal of Affective Disorders,278,280-287. [21] Kleisner, K., Kocnar, T., Rubesova, A., Flegr, J. (2010). Eye color predicts but does not directly influence perceived dominance...
work page 2021
-
[48]
Establishing and maintaining long-term human-computer relationships
Venturoso, L., Gabrieli, G., Truzzi, A., Azhari, A., Setoh, P., Bornstein, M. H., Esposito, G. (2019). Effects of baby schema and mere exposure on explicit and implicit face processing. Frontiers in psychology,10,2649. [49] Wang, C.-H., Wu, C.-L. (2022). Bridging the digital divide: the smart tv as a platform for digital literacy among older adults. Behav...
work page 2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.