Pith. sign in

REVIEW 4 major objections 5 minor 2 references

The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that baby-schema eye and mouth proportions can be tuned to maximize older adults' perceived trustworthiness in virtual humanoid agents, and proposes a concrete design paradigm for those features.

desk verdict A useful dataset and a real question, but the winning proportions shift between abstract, results, and design paradigm, and the statistics pool dependent ratings; needs major revision before it can be used. read the letter →

arxiv 2411.18047 v1 pith:V6WV72IQ submitted 2024-11-27 cs.HC

classification cs.HC
keywords babyschemaeffectvirtualhumanoidagentstrustworthinessolderadultsfacialfeaturedesigneyemouthhuman-agentinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the "baby schema" effect, the cluster of infant-like facial proportions that make faces seem cute and trustworthy, can be carried over to the design of virtual humanoid agents for older adults, and that the eyes and mouth are the features that matter most. To test this, the author built a base child-like virtual face, varied six eye and mouth dimensions (eye size, eye height, eye spacing, mouth size, mouth height, and smile arc), each at three levels, and had 162 older adults rate the credibility of 729 resulting faces. The results point to a specific winning combination of proportions, proposed as a design paradigm for trusted virtual caregivers. One caution is that the paper states the winning combination differently in the abstract (eye spacing 0.43W, mouth height 0.74H) and in Section 4.4 (eye spacing 0.41W, mouth height 0.77H), so the exact paradigm is not stably pinned down. The significance, if the finding holds, is a concrete, parameter-level guideline for making assistive agents more acceptable to an aging population.

What carries the argument

The machinery is a controlled six-factor stimulus ladder built on a base 4 to 6 year-old female face whose facial width-to-height ratio (1.90) was checked against the baby schema standard (1.96). Each of six features, eye size, eye height, eye spacing, mouth size, mouth height, and smile arc, varied over three levels expressed as fractions of face width W or face height H, generating $3^{6}$ = 729 faces. Older adults rated each face on a five-item credibility scale, and an ANCOVA with gender and smartphone experience as covariates was used to rank the levels and identify the best combination. The proposed "facial design paradigm" is the set of winning normalized proportions, intended to be used directly by designers of virtual humanoid agents.

What would settle it

Run a mixed-effects reanalysis with participant and face combination as random factors: if the mouth-size, eye-spacing, mouth-height, and smile-arc effects disappear, or the winning combination moves outside the reported range, the fixed-effect ANCOVA result is an artifact. Alternatively, present the Section 4.4 highest-credibility face and the abstract's competing face to a fresh sample in pairwise forced choice; the paradigm is falsified if the Section 4.4 face does not win significantly more often than chance.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that perceived credibility of a child-like virtual agent among older adults is not flat across facial proportions: mouth size, eye spacing, mouth height, and smile arc each produced significant main effects in a six-way ANCOVA, while eye size and eye height did not. Interactions further shaped the ratings, with the most credible face in Section 4.4 having eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye spacing 0.41W, mouth height 0.77H, and smile arc 0.043H, on a face whose width-to-height ratio already matches the baby schema. The abstract reports a different winning set (eye spacing 0.43W, mouth height 0.74H), so the paper's concrete recommendation is internally inconsistent even though its directional claim, that baby-schema-consistent eyes and mouth raise trust, is consistent throughout. The study then validates that the top-rated faces are perceived as more childlike than the lowest-rated ones, tying the trust result to the baby schema mechanism.

Load-bearing premise

The statistical analysis treats each of the 1,458 ratings as an independent observation even though only two participants saw each face combination, so the reported F-tests assume that ratings do not cluster by participant or by the nine-face subset.

Editorial extensions

If this is right

  • If the winning proportions hold, designers of virtual caregivers for older adults can start from a concrete parameter set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, mouth height near 0.77H, smile arc 0.043H) instead of relying on intuition.
  • Mouth dimensions do more work than eye dimensions: mouth size and mouth height had strong main effects, whereas eye size did not, suggesting that design effort should go to the mouth first.
  • Because eye size mattered mainly in interaction with other features, the paper implies that facial features should not be tuned independently; a configural approach to face design is more promising.
  • The validation results imply that "looks like a child" and "looks trustworthy" move together for older adults, so perceived juvenility can serve as a proxy target in agent design.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: because the abstract and Section 4.4 disagree on the two most salient parameters, the safest reading is that the broad direction (baby-schema eyes and mouth raise trust) is supported while the precise numbers are not; a pre-registered replication that tests the two competing faces head-to-head would settle the paradigm.
  • Beyond the paper: the design could be tested in a real smart-home interaction, such as medication reminders, to see whether the trust advantage in ratings changes willingness to follow advice; rating studies can overstate first-impression effects.
  • Beyond the paper: since the sample is Chinese older adults with varied smartphone experience, the same stimulus set could be run with younger adults or other cultural groups to decompose the baby-schema response from age-specific and culture-specific preferences.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper tests whether facial proportions associated with the baby schema affect older adults' trust in virtual humanoid agents. The authors generated 729 child-like faces by varying six features (eye size, eye height, eye spacing, mouth size, mouth height, smile arc) at three levels each, and 162 older adults rated subsets of nine faces on a five-item credibility scale. The abstract claims a single optimal proportion set (eye size 0.25W, mouth size 0.27W, eye height 0.64H, eye distance 0.43W, mouth height 0.74H, smile arc 0.043H) and proposes this as a design paradigm. The body of the paper, however, reports different winning combinations across sections, and the statistical analysis treats nested ratings as independent observations, so the central claim is not stably defined or supported.

Significance. If the central claim were sound, the paper would provide practically useful design guidance for trustworthy virtual agents aimed at older adults, a population that is growing in smart-home and care contexts. The study has genuine strengths: a full-factorial stimulus design (729 faces), a targeted older adult sample, and an additional validation rating of perceived juvenility. These are commendable. However, the contradictory statements of the optimal parameters and the invalid treatment of the clustered data mean the paper currently does not establish any single proportion set as most trustworthy, and the proposed design paradigm is effectively a post-hoc description of one observed cell rather than a tested prediction.

major comments (4)
  1. [Abstract; §4.3.1; §4.3.2; §4.4] The paper does not identify a single, stable winning combination. The abstract states that the highest credibility occurred at eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H. Section 4.3.1 states that the highest credibility was associated with eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H, and §4.3.2 repeats the 0.062H/0.43W/0.74H combination as the highest three-way interaction. Section 4.4, which presents the 'face with the highest level of credibility' and the design paradigm, instead uses eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. Because the headline result is a specific proportion vector, these discrepancies are load-bearing: they leave the central claim undefined and prevent a reader from testing or applying the proposed paradigm.
  2. [§4.3.1] The main-effect summary for mouth height is internally contradictory. The text first states that 'the highest credibility was associated with a mouth height of 0.74H,' then later states that 'faces having a mouth height of 0.74H exhibiting lower credibility than faces with mouth heights of 0.77H and 0.79H. H5 was confirmed.' A single analysis cannot support both claims, and H5 pre-specified 0.77H, not 0.74H. This contradiction affects the interpretation of the mouth-height result and the confirmation reported for H5.
  3. [§3.4; §4.2; §4.3; Table 3] The ANCOVA treats the 1,458 ratings as independent observations, with degrees of freedom near 1457. According to §3.4, each of the 162 participants rated nine faces, and each of the 729 face combinations was rated by two participants. Ratings are therefore nested within participants and within stimuli. The reported F-tests and p-values for the six-way ANCOVA and its interactions are not valid under this design; participant-level variance and stimulus-level clustering are ignored. This is a fundamental problem for all main-effect and interaction claims in Section 4, including the eye-spacing, mouth-height, and smile-arc effects that drive the conclusion.
  4. [§4.4; Hypotheses H1–H6] The proposed design paradigm is not an independent validation but a restatement of the empirically highest-scoring cell from the same ANCOVA. The hypotheses pre-specify some factor levels (e.g., H4 at 0.41W, H5 at 0.77H, H6 at 0.043H), but the final §4.4 face also includes eye height 0.64H, which was not pre-specified (H3 was 0.62H), and the selection is made by inspecting the same data used to estimate the effects. As a result, the claimed 'design paradigm' overfits the sample and has no out-of-sample predictive test, so its generality is not established.
minor comments (5)
  1. [§4.2.1] The sentence 'HI needed to be supported' should read 'H1 was not supported.'
  2. [Table 3] The table lists 'Eye width (EW)' as a factor, but the text and hypotheses refer to eye spacing; the relationship between 'eye width' and 'eye spacing' should be clarified to avoid confusion.
  3. [§4.1] The F-statistics for eye size and eye height are reported as identical (F(2,1455) = 14.32) despite different means and standard errors; this is likely a copy-paste error and should be corrected.
  4. [References] References [42] and [43] appear to be the same source, and entries [43] and [44] duplicate each other; the reference list should be deduplicated and renumbered.
  5. [Figure 9 caption] The caption contains the typo 'Da Mouth' and should be corrected to 'Large mouth.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the paper's central claim is an empirically selected result, not a prediction that reduces to its own inputs; the reported inconsistencies across sections are statistical and reporting problems, not circularity.

full rationale

This paper contains no formal derivation chain in which an output is constructed from its own inputs. The baby-schema parameter values used to build the stimuli are taken from external literature (e.g., Lorenz 1943; Luo 2020; Ferstl 2017), not from the present experiment, so there is no self-definitional loop in which 'baby schema' is defined by the credibility scores. The hypotheses H1-H6 are stated before analysis in Section 3.1, and the results are obtained from ANCOVA on participant ratings. The later 'design paradigm' in Section 4.4 is a restatement of the empirically selected winning face, which is ordinary empirical inference rather than a fitted parameter renamed as an independent prediction. There are no load-bearing self-citations and no imported uniqueness theorem. The paper does contain a serious internal inconsistency: the abstract reports the winner as eye distance 0.43W, mouth height 0.74H, and smile arc 0.043H; Section 4.3.1 and 4.3.2 report the highest credibility at eye spacing 0.43W, mouth height 0.74H, and smile intensity 0.062H; and Section 4.4 selects eye spacing 0.41W, mouth height 0.77H, and smile intensity 0.043H. This instability, along with the contradictory statement in Section 4.3 that 0.74H mouth height has lower credibility than 0.77H and 0.79H, undermines the reliability and interpretability of the headline result. However, inconsistency and post-hoc selection are not the same as circularity, and no equation or constructed equivalence can be exhibited that makes the result equal to its input by definition. The appropriate circularity verdict is therefore no significant circularity, while the manuscript's correctness risk is high.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

This paper does not derive results from first principles. It relies on the established baby-schema literature, a hand-built base model, and a set of six geometric proportions whose optimal values are effectively the maximum observed means in the data. The statistical analysis depends on an independence assumption that is likely violated by the nested design.

free parameters (6)
  • Eye size optimum = 0.25W
    Reported as the level with highest credibility in the abstract and in the H1 confirmation; not significantly different from other eye sizes in ANCOVA (F(2,1457)=1.00, p=0.37).
  • Mouth size optimum = 0.27W
    Reported as more trustworthy and supported by ANCOVA; however, the interaction analysis reports a higher mean for mouth size 0.34W with eye spacing 0.43W, so the optimum depends on other features.
  • Eye height optimum = 0.64H
    Stated in abstract as optimal, but ANCOVA found no significant eye-height effect (F(2,1457)=0.56, p=0.57) and H3 predicted 0.62H.
  • Eye distance optimum = 0.43W (abstract) / 0.41W (Section 4.4)
    The abstract and Section 4.3.1 say 0.43W, while Section 4.4's highest credibility face uses 0.41W; H4 predicted 0.41W.
  • Mouth height optimum = 0.74H (abstract) / 0.77H (Section 4.4)
    Abstract and Section 4.3.1 say 0.74H, Section 4.4 says 0.77H; H5 predicted 0.77H.
  • Smile arc optimum = 0.043H (abstract) / 0.062H (Section 4.3.1)
    Abstract and Section 4.4 say 0.043H, but Section 4.3.1 states highest credibility for smile intensity 0.062H; H6 predicted 0.043H.
assumptions (5)
  • domain assumption The baby schema effect causes increased perceived trustworthiness in older adults
    Taken from prior literature and used to motivate the hypotheses; not re-tested against a non-baby-schema baseline.
  • domain assumption The six geometric ratios capture the relevant aspects of baby schema for trust
    Section 3.1 defines the base model and feature variations using these ratios; other facial features (nose, cheeks, hair) are held constant or ignored.
  • ad hoc to paper The frontal view of the 3D character faithfully represents the intended feature proportions
    Figure 3 caption notes that due to three-dimensionality, values may not show linear changes, but the frontal-view effect is assumed consistent; this is necessary for the ratios to be meaningful.
  • domain assumption Five Likert items on reliability, sincerity, honesty, trustworthiness, and persuasiveness measure facial credibility
    Section 3.3 adapts the Gorn (2008) babyface trait-inference scale to a 7-point Likert; no additional validation of the scale for older adults is reported.
  • ad hoc to paper The 1,458 ratings from 162 participants can be treated as independent observations in the ANCOVA
    Section 3.4 assigns each participant to 9 of 729 combinations, so ratings are clustered by participant and each cell has n=2; the ANCOVA in Section 4 uses df near 1457, assuming independence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness." pith.science (2026). https://pith.science/paper/V6WV72IQ

@misc{pith2026241118047,
  author       = {Pith},
  title        = {Pith review of: The Trusted Caregiver: The Influence of Eye and Mouth Design Incorporating the Baby Schema Effect in Virtual Humanoid Agents on Older Adults Users' Perception of Trustworthiness},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V6WV72IQ}},
  note         = {Machine review of arXiv:2411.18047}
}
read the original abstract

The increasing proportion of the older adult population has made the smart home care industry one of the critical markets for virtual human-like agents. It is crucial to effectively promote a trustworthy human-computer partnership with older adults, enhancing service acceptance and effectiveness. However, few studies have focused on the facial features of the agents themselves, where the "baby schema" effect plays a vital role in enhancing trustworthiness. The eyes and mouth, in particular, attract most of the audience's attention and are especially significant. This study explores the impact of eye and mouth design on users' perception of trustworthiness. Specifically, a virtual humanoid agents model was developed, and based on this, 729 virtual facial images of children were designed. Participants (N=162) were asked to evaluate the impact of variations in the size and positioning of the eyes and mouth regions on the perceived credibility of these virtual agents. The results revealed that when the facial aspect ratio (width and height denoted as W and H, respectively) aligned with the "baby schema" effect (eye size at 0.25W, mouth size at 0.27W, eye height at 0.64H, eye distance at 0.43W, mouth height at 0.74H, and smile arc at 0.043H), the virtual agents achieved the highest facial credibility. This study proposes a design paradigm for the main facial features of virtual humanoid agents, which can increase the trust of older adults during interactions and significantly contribute to the research on the trustworthiness of virtual humanoid agents.

Figures

Figures reproduced from arXiv: 2411.18047 by the authors.

Figure 1
Figure 1. Diagram of measurement parameters for infant facial proportion. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Basic facial design based on proportional parameters. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Detailed facial feature indicators for virtual agents. (Due to the three-dimensionality of the face, some values may not show linear changes, but the effect observed from a frontal view is consistent.) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Scene of the experiment site (a) and User experiment interface (b). [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Single factor analysis of each characteristic item. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Bar chart of the two-way interaction results of eye size * mouth smile in facial credibility rating. Error bar represents ± 1SE [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Bar chart of two-way interaction results of mouth size * eye spacing in facial credibility rating. Error bar represents ± 1SE [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Bar chart of two-way interaction results of mouth size * mouth height in facial credibility rating. Error bar represents ± 1SE. 4.3 The effect of feature localization 4.3.1 Main Effects of Feature Position The ANCOVA results in [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Bars chart representing the results of three [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Bar chart showing the results of three-way interaction in facial credibility rating. (I) eye spacing * smile degree and low mouth; (II) eye spacing * smile degree and medium high mouth; (III) eye spacing * smile degree and high mouth. Error bar represents ± 1SE [PITH…
Figure 11
Figure 11. Figure 11: Bar chart showing the results of three-way interaction in facial credibility rating. (I) smile degree * eye spacing and low eyes (II) smile degree * eye spacing and medium height eyes; (III)smile degree * eye spacing and high eyes. Error bar represents ± 1SE. In summa…
Figure 12
Figure 12. Figure 12: The Bar chart shows the results of single factor ANCOVA in the face childishness rating. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: The face with the highest level of credibility. [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 14
Figure 14. Figure 14: Facial Design Paradigm of Virtual Human Agent. 5 DISCUSSION This study explored the influence of primary facial features on credibility perception among older adults. A full￾factor mixed design experiment provided preliminary indicators regarding the impact of differe…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages

  1. [20]

    L., Hogstrom, J., Sundstrom, K., Frick, A., Serlachius, E

    Kleberg, J. L., Hogstrom, J., Sundstrom, K., Frick, A., Serlachius, E. (2021). Delayed gaze shifts away from others' eyes in children and adolescents with social anxiety disorder. Journal of Affective Disorders,278,280-287. [21] Kleisner, K., Kocnar, T., Rubesova, A., Flegr, J. (2010). Eye color predicts but does not directly influence perceived dominance...

  2. [48]

    Establishing and maintaining long-term human-computer relationships

    Venturoso, L., Gabrieli, G., Truzzi, A., Azhari, A., Setoh, P., Bornstein, M. H., Esposito, G. (2019). Effects of baby schema and mere exposure on explicit and implicit face processing. Frontiers in psychology,10,2649. [49] Wang, C.-H., Wu, C.-L. (2022). Bridging the digital divide: the smart tv as a platform for digital literacy among older adults. Behav...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.