Pith. sign in

REVIEW 4 major objections 4 minor 28 references

Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum

T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Android Andrea's simulated emotions fail to boost visitor acceptance

desk verdict Honest null result for emotion simulation on a museum android, but the day-condition confound means the negative finding is real but not as clean as it looks. read the letter →

arxiv 2607.16428 v1 pith:AT2CXSYJ submitted 2026-07-17 cs.RO cs.AIcs.HCcs.LG

classification cs.ROcs.AIcs.HCcs.LG
keywords human-robotinteractionandroidrobotemotionsimulationmuseumfieldstudytechnologyacceptancemodelTAM2ChatGPTWASABI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports a six-day field deployment of the android robot Andrea in a German museum, where it autonomously conversed with visitors in multiple languages. Three conditions were compared: no emotion simulation, emotions labeled by ChatGPT 4.1, and emotions produced by the WASABI affect-simulation architecture. Based on 73 visitor questionnaires built on the TAM2 technology-acceptance model, the emotional variants did not improve intention to use, perceived usefulness, or perceived ease of use; the only significant pairwise differences favored the non-emotional condition. The emotion manipulation check also failed: visitors did not rate the emotional versions as more emotional. The authors conclude that these two emotion-implementation approaches, as built, provided no measurable acceptance benefit in this public, short-interaction setting.

What carries the argument

The android Andrea, a 52-actuator humanlike robot, runs a software pipeline combining Whisper for speech-to-text, a ChatGPT 4.1 assistant with retrieval-augmented generation for dialogue, XTTS v2 for multilingual speech synthesis, FaceXHubert for lip-sync, and posenet for visitor tracking. The experimental variable is the emotion component: no emotion, ChatGPT-generated emotion labels, or WASABI's computational affect dynamics, all expressed through validated static facial expressions. Acceptance is measured with an extended TAM2 questionnaire (intention to use, perceived usefulness, perceived ease of use) plus a usefulness scale and three perceived-emotionality items.

What would settle it

Compare acceptance ratings from the two no-emotion days (day 1 and day 4): if they differ substantially, day-level confounds are present and the emotion-condition comparison is unreliable. Separately, annotate video of the robot's face per condition; if the ChatGPT and WASABI conditions show no more visible emotional expressions than the no-emotion condition, the emotion manipulation failed and the null result says nothing about emotion simulation.

Watch

Extended reading notes

Core claim

In this ecologically valid but uncontrolled museum setting, equipping Andrea with either of two emotion-simulation pipelines did not make visitors more accepting of the robot, and the simulated emotions were not noticeable enough to move perceived-emotionality ratings. The no-emotion condition was rated at least as useful and easy to use as the emotional conditions, and significantly more so than the ChatGPT-emotion condition on perceived usefulness and perceived ease of use; WASABI showed no differences, but its sample was small (N=14). The authors interpret the negative manipulation check as devaluing the emotion-implementation approach.

Load-bearing premise

That the three emotion conditions differed only in the emotion simulation—rather than in which calendar day they ran, the visitors who happened to be in the museum, how actively experimenters recruited participants, or whether the robot's facial expressions actually showed the intended emotions; the paper's own manipulation-check failure and log inspection make the last point especially uncertain.

Editorial extensions

If this is right

  • For short, task-oriented public interactions, adding simulated emotion to an android's conversation is not automatically an acceptance benefit and may slightly reduce perceived usefulness and ease of use.
  • Emotion simulation that fails a conscious manipulation check cannot be expected to influence acceptance, so field studies of robot emotion need behavioral or physiological verification that the manipulation was perceptible.
  • The non-emotional version scored highest on ease of use and usefulness, suggesting visitors may value clarity and reliability over emotional mimicry in museum information kiosks.
  • The WASABI condition's small sample (N=14) leaves that architecture's effect unresolved; the data cannot support any claim about WASABI specifically.
  • Running a fully autonomous multilingual android in a museum for six days is feasible, and most visitors arrived without knowing the robot was there.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the emotion labels were rarely non-happy (as the paper's log inspection suggests) and emotions were shown only through the face while the robot was not speaking, the emotion signal may have been too weak or too brief to register; a redesign that times emotional expressions to pauses or adds vocal emotion cues might yield a different result.
  • The day-level confounding (conditions on different days, one day excluded, varying experimenter recruiting) means this should be read as a suggestive null result rather than evidence that emotion simulation is useless; a within-day counterbalanced or longer random-order deployment would tighten the test.
  • A testable extension would compare the same android with and without emotion in a repeated one-on-one interaction over minutes, where visitors have time to notice and respond to emotional states, rather than in a one-shot public encounter.
  • The slight preference for the non-emotional robot hints that an android's humanlike appearance may already set social expectations that emotional behavior, if imperfectly executed, does not meet; matching the quality of the emotion display to the realism of the face may matter more than whether emotion is simulated at all.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper describes a field study in which the android robot Andrea interacted autonomously with 73 museum visitors under three emotion conditions: a no-emotion baseline (ChatGPTpure), ChatGPT-generated emotion labels, and the WASABI emotion architecture. Visitors completed an extended TAM2 questionnaire. Kruskal-Wallis tests found no significant differences for Intention to Use or Usefulness Today, a significant difference for Perceived Ease of Use (p=.033), and a near-significant difference for Perceived Usefulness (p=.056); pairwise Dwass-Steel-Critchlow-Fligner tests showed that the ChatGPTpure condition outscored the ChatGPTemotion condition on PU and PEOU. The manipulation-check items on perceived emotionality showed no significant condition differences. The authors conclude that the two emotion-simulation approaches did not improve acceptance and were not consciously detectable.

Significance. If the conclusions were valid, this would be a useful negative result for the design of affective android systems in public settings, and the rich multilingual, real-world deployment data are a genuine contribution. The manuscript is commendably transparent: it reports normality checks, non-parametric tests, effect sizes, box plots, a manipulation check, and the hardware failure that removed day 6. However, the central causal claim about the emotion conditions is compromised by a complete day-condition confound and by the very small WASABI sample. As it stands, the paper provides an honest exploratory field report, but the abstract's strong negative conclusion is not yet supported by the design.

major comments (4)
  1. [§4, §5] Condition is perfectly confounded with calendar day: no-emotion on days 1/4, ChatGPTemotion on days 2/5, WASABI on day 3 (day 6 excluded). The abstract's claim that emotion simulation 'did not yield any positive effects' is a causal statement that requires exchangeability of visitor populations, experimenter recruiting behavior, hardware state, and exhibit context across days. The Method section explicitly states that recruiting style depended on the experimenter on duty, and the robot's facial expressions visibly degraded by day 6. The significant PEOU/PU contrasts between ChatGPTemotion and ChatGPTpure, and the null manipulation check, could each be produced by day-level differences. The authors never mention this confound; a day-level permutation analysis, or at least an explicit argument with supporting data, is needed to support the central negative conclusion.
  2. [Table 1, §6] The WASABI condition has N=14 and is inseparable from day 3. The authors acknowledge that the condition is 'likely underpowered,' but they still include WASABI in the abstract's general conclusion that 'these first two approaches' had no positive effects. With N=14, the null result for WASABI is weakly informative, and the day-3-only assignment makes it impossible to distinguish condition effects from day effects. Either the conclusions should be restricted to the ChatGPTemotion arm (with the confound still unresolved), or the authors should provide a sensitivity/power analysis and day-level evidence.
  3. [Table 2, §5] The reporting of effect sizes is inconsistent as written. Table 2 lists epsilon-squared values of 0.05, 0.08, and 0.09, but §5 describes these as 'moderate effect size (ε > 0.6)'. These numbers do not match; epsilon-squared of 0.09 is generally considered small. In addition, pairwise comparisons for PU were conducted despite the overall Kruskal-Wallis p = .056, which is not below α = .05. The authors should justify this decision or treat the PU pairwise result as exploratory; otherwise the claims of 'significant' advantages for ChatGPTpure over ChatGPTemotion are overstated.
  4. [Abstract, Table 5] The claim that the emotion manipulation was 'not detectable on a conscious level' rests on null results for three questionnaire items with unequal group sizes and no equivalence testing or sensitivity analysis. A null p-value cannot by itself establish that the conditions were indistinguishable, particularly with N=14 in one arm. The authors should either soften this to 'no evidence of a conscious difference in this sample' or report a sensitivity analysis showing the smallest detectable effect.
minor comments (4)
  1. [§5] Typographical errors: 'responeded' and 'It responeded most often' should be corrected; 'to a very small extend' should be 'extent'; there is a stray 'W ASABI' in the age distribution paragraph.
  2. [§2] The related-work paragraph on previous Andrea results (2023) repeats information given later in §3; consider merging to avoid redundancy.
  3. [Figures 4–7] The box plots would be more informative if they showed individual data points or at least the group N alongside the medians; this is especially relevant given the uneven group sizes.
  4. [Table 2] The epsilon-squared values are labelled 'ε²' but the text uses 'ε'; unify the notation and report confidence intervals for the effect sizes where possible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical comparison uses independent TAM2 and emotion-perception questionnaire outcomes; the self-citations are technical context, not load-bearing.

full rationale

The derivation chain runs from an implemented manipulation (no emotion vs ChatGPT-assigned emotion labels vs WASABI dynamics) to visitor questionnaire responses. The outcome variables — ITU, PU, PEOU, UT, and the three perceived-emotionality items — are standard or separately worded survey items, not values computed from the emotion-generation parameters. No equation fits a parameter to the reported acceptance or emotion data and then presents it as a prediction; Kruskal-Wallis and Dwass-Steel-Critchlow-Flinger tests are applied directly to raw ratings. The manipulation check is a distinct set of items, so its null result is an empirical finding rather than an artifact of the definition of the emotion conditions. The paper cites the authors' prior work for the robot software ([10,11]), the WASABI simulator ([2]), validated facial expressions ([15]), and the voice ([19,20]), but these citations support the construction of the stimulus, not the conclusion that the emotion conditions failed to improve acceptance or were not consciously detected. Even the paper's own limitation paragraph (WASABI N=14, degraded hardware, limited gestures) and the day-level confounding between condition and calendar day are threats to internal validity and statistical power, not circularity: they concern whether the observed null results are trustworthy, not whether the results are equivalent to their inputs by construction. No self-definitional, fitted-input-as-prediction, uniqueness-import, or ansatz-smuggling pattern is present, so the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No new entities. The paper's empirical claims rest on the deployment pipeline (LLM, TTS, face animation, emotion modules) as a given system, on the TAM2 instrument, and on the day-condition assignment being exchangeable. The dominant risk is the day confound and the small WASABI sample, not free parameters or invented constructs.

free parameters (1)
  • None in the statistical claims
    No model parameters are fitted to the survey data beyond descriptive statistics and non-parametric tests. The emotion thresholds, linear decay, and WASABI internals are inherited from cited prior work [2, 15, 20], not fitted to this study's outcome.
assumptions (4)
  • domain assumption ChatGPT 4.1's emotion label generation and the WASABI valence-prompting pipeline function as an adequate 'emotion simulation' implementation.
    Section 3 describes the two pipelines; the study's interpretation that the manipulation was implemented correctly rests on this assumption, and the failed manipulation check leaves room for an implementation-fidelity issue.
  • domain assumption The TAM2 questionnaire, translated from [23]/[26], measures the intended latent constructs in a museum context with visitors of mixed ages and nationalities.
    Section 4.2 lists the items; no additional validation of the German/English instrument in this deployment is reported.
  • domain assumption Visitors interacted with only one condition (between-groups validity).
    Section 4 states the between-groups design; return visits or prior exposure to the 2023 deployment are not excluded and could contaminate comparisons.
  • ad hoc to paper Day-assignment confounding is harmless (visitor populations are exchangeable across days).
    Section 4 assigns conditions to days 1-6; Section 5 reports varying recruitment styles by experimenters, making day-level confounds plausible. This is the weakest structural assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum." pith.science (2026). https://pith.science/paper/AT2CXSYJ

@misc{pith2026260716428,
  author       = {Pith},
  title        = {Pith review of: Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AT2CXSYJ}},
  note         = {Machine review of arXiv:2607.16428}
}
read the original abstract

For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with visitors, fully autonomously. Building on previously gathered qualitative results, the robot was now capable of engaging in multi-lingual conversation with the visitors about the museum context. The robot was prepared with context information about the museum in general and its surrounding exhibits this time. The robot featured a slightly artificial sounding voice that was previously evaluated as congruent with its gender-ambiguous but very humanlike design. Three experimental conditions were implemented, in which either (1) the robot simulated no emotions, (2) the robots emotions were determined by ChatGPT 4.1, or (3) the WASABI emotion simulation architecture simulated the robot's emotion dynamics. An extended version of the TAM2 questionnaire was employed to let 73 visitors report on several factors of their opinion about the android robot Andrea after having experienced it. In result, the statistical analysis suggests that these first two approaches to implementing emotions into the chat architecture of our android robot Andrea did not yield any positive effects on the subjective evaluations by the visitors and were not detectable on a conscious level.

Figures

Figures reproduced from arXiv: 2607.16428 by the authors.

Figure 1
Figure 1. Android Andrea sitting on a bench in the museum and conversing with visitors [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The android robot placed on a bench in the museum [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Overview of the software architecture of the conversational functions of the android Andrea. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Box plot showing the distribution of intention to use (ITU) by experimental condition with indicated median values. 2 4 6 ChatGPTemotion ChatGPTpure WASABI Experimental Condition PerceivedUsefulness [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: Box plot showing the distribution of perceived ease of use (PEOU) by experimental condition with indicated me￾dian values. 0.0 2.5 5.0 7.5 10.0 ChatGPTemotion ChatGPTpure WASABI Experimental Condition How useful do you think ...? [PITH_FULL_IMAGE:figures/full_fig_p006…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

28 extracted references · 6 canonical work pages

  1. [1]

    Evangelia Baka, Nidhi Mishra, Emmanouil Sylligardos, and Nadia Magnenat- Thalmann. 2022. Social Robots and Digital Humans as Job Interviewers: A Study of Human Reactions Towards a More Naturalistic Interaction. InHuman- Computer Interaction. Technological Innovation, Masaaki Kurosu (Ed.). Springer International Publishing, Cham, 455–474.doi:10.1007/978-3-...

  2. [2]

    Becker-Asano

    C. Becker-Asano. 2014. WASABI for affect simulation in human-computer in- teraction. In Proc. on Emotion Representations and Modelling for HCI Systems . Springer, Sydney, Australia

  3. [3]

    Christian Becker-Asano, Kohei Ogawa, Shuichi Nishio, and Hiroshi Ishiguro

  4. [4]

    YouScareMe

    KarstenBernsandAshitaAshok.2024. “YouScareMe”:TheEffectsofHumanoid Robot Appearance, Emotion, and Interaction Skills on Uncanny Valley Phenom- enon. Actuators 13, 10 (Oct. 2024), 419.doi:10.3390/act13100419

  5. [5]

    Felix Carros, Berenike Bürvenich, Ryan Browne, Yoshio Matsumoto, Gabriele Trovato, Mehrbod Manavi, Keiko Homma, Toshimi Ogawa, Rainer Wieching, and Volker Wulf. 2022. Not that Uncanny After All? An Ethnographic Study on Android Robots Perception of Older Adults in Germany and Japan. InSocial Robotics. Vol. 13818. Springer Nature Switzerland, Cham, 574–586...

  6. [6]

    Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Lo- gan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, and Ju- lian Weber. 2024. XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model. InInterspeech 2024. 4978–4982. doi:10.21437/Interspeech.2024-2016

  7. [7]

    Haozhe Chen, Run Chen, and Julia Hirschberg. 2024. EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control. arXiv:2410.00316 [cs.CL] https: //arxiv.org/abs/2410.00316

  8. [8]

    Norina Gasteiger, Mehdi Hellou, and Ho Seok Ahn. 2021. Deploying social robotsinmuseumsettings:Aquasi-systematicreviewexploringpurposeandac- ceptability. International Journal of Advanced Robotic Systems 18, 6 (Nov. 2021), 17298814211066740. doi:10.1177/17298814211066740 Publisher: SAGE Publica- tions

Show all 28 references
  1. [9]

    Kazi Injamamul Haque and Zerrin Yumak. 2023. FaceXHuBERT: Text- less Speech-driven E(X)pressive 3D Facial Animation Synthesis Using Self- Supervised Speech Representation Learning. InProceedings of the 25th Interna- tional Conference on Multimodal Interaction (ICMI ’23) . Asso...

  2. [10]

    Marcel Heisler and Christian Becker-Asano. 2023. An Android Robot Head as Embodied Conversational Agent. InISR Europe 2023; 56th International Sympo- sium on Robotics. 93–99. https://ieeexplore.ieee.org/document/10363058

  3. [11]

    Marcel Heisler and Christian Becker-Asano. 2025. Conversations with An- drea: Visitors’ Opinions on Android Robots in a Museum. In2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO- MAN). IEEE, Eindhoven, Netherlands, 112–119. doi:10.1109/...

  4. [12]

    Marcel Heisler, Stefan Kopp, and Christian Becker-Asano. 2023. Making an Android Robot Head Talk. In2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . 1837–1842. doi:10.1109/RO- MAN57019.2023.10309532 ISSN: 1944-9437

  5. [13]

    Mehdi Hellou, JongYoon Lim, Norina Gasteiger, Minsu Jang, and Ho Seok Ahn

  6. [14]

    Hangyeol Kang, Thiago Freitas, Maher Ben Moussa, and Nadia Magnenat Thal- mann. 2026. Affective and Conversational Predictors of Re-Engagement in Human–Robot Interactions: A Student-Centered Study with a Humanoid Social Robot. Int J of Soc Robotics18,3(March2026),41. doi:10.10...

  7. [15]

    Amelie Kassner and Christian Becker-Asano. 2023. Comparing an android head with its digital twin regarding the dynamic expression of emotions. In2023 11th International Conference on Affective Computing and Intelligent Interaction Work- shops and Demos (ACIIW)

  8. [16]

    IntelligentConversational Android ERICA Applied to Attentive Listening and Job Interview.http://arxiv

    TatsuyaKawahara,KojiInoue,andDiveshLala.2021. IntelligentConversational Android ERICA Applied to Attentive Listening and Job Interview.http://arxiv. org/abs/2105.00403

  9. [17]

    Dacher Keltner and Daniel T. Cordaro. 2017. Understanding Multimodal Emo- tional Expressions: Recent Advances in Basic Emotion Theory. InThe Science of Facial Expression, James A. Russell and Jose Miguel Fernandez Dols (Eds.). Ox- ford University Press.doi:10.1093/acprof:oso/9...

  10. [18]

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. 2015. PoseNet: A Convo- lutional Network for Real-Time 6-DOF Camera Relocalization. In2015 IEEE Intl. Conf. on Computer Vision (ICCV) . 2938–2946. doi:10.1109/ICCV.2015.336 ISSN: 2380-7504

  11. [19]

    YourRobot,MyVoice: Enhancing Android Robot Likability through Personalization by Cloning the User’s Voice

    Johanna Magdalena Kuch, Marcel Heisler, Stina Klein, Silvan Mertes, Lennart Eing,ElisabethAndré,andChristianBecker-Asano.2025. YourRobot,MyVoice: Enhancing Android Robot Likability through Personalization by Cloning the User’s Voice. In2025 34th IEEE International Conference o...

  12. [20]

    EvaluatingGenderAmbigu- ity,NoveltyandAnthropomorphisminHummingandTalkingVoicesforRobots

    Johanna Magdalena Kuch, Jauwairia Nasir, Silvan Mertes, Ruben Schlagowski, ChristianBecker-Asano,andElisabethAndré.2024. EvaluatingGenderAmbigu- ity,NoveltyandAnthropomorphisminHummingandTalkingVoicesforRobots. In 2024 33rd IEEE International Conference on Robot and Human Inte...

  13. [21]

    Martina Mara and Markus Appel. 2015. Science fiction reduces the eeriness of androidrobots:Afieldexperiment. Computers in Human Behavior48(July2015), 156–162. doi:10.1016/j.chb.2015.01.007

  14. [22]

    Now Loading…

    Shushi Namba, Wataru Sato, Saori Namba, Alexander Diel, Carlos Ishi, and Takashi Minato. 2024. How an Android Expresses “Now Loading…”: Examin- ing the Properties of Thinking Faces. Int J of Soc Robotics 16, 8 (Aug. 2024), 1861–1877. doi:10.1007/s12369-024-01163-9

  15. [23]

    Thomas Olbrecht. 2010. Akzeptanz von E-Learning. Eine Auseinandersetzung mit dem Technologieakzeptanzmodell zur Analyse individueller und sozialer Ein- flussfaktoren. Univ., Jena. 215 S. pages.http://www.db-thueringen.de/servlets/ DerivateServlet/Derivate-21996/Olbrecht/Disser...

  16. [24]

    Astrid M Rosenthal-von der Pütten, Nicole C Krämer, Christian Becker-Asano, Kohei Ogawa, Shuichi Nishio, and Hiroshi Ishiguro. 2014. The uncanny in the wild. Analysis of unscripted human–android interaction in the field.Interna- tional Journal of Social Robotics 6 (2014), 67–83

  17. [25]

    Nadinethe Social Robot: Three Case Studies in Everyday Life

    NadiaMagnenatThalmann,NidhiMishra,andGauriTulsulkar.2021. Nadinethe Social Robot: Three Case Studies in Everyday Life. InSocial Robotics. Springer International Publishing, Cham, 107–116.doi:10.1007/978-3-030-90525-5_10

  18. [26]

    Viswanath Venkatesh and Fred D. Davis. 2000. A Theoretical Extension of the Technology Acceptance Model: Four Longitudinal Field Studies.Management Science 46, 2 (2000), 186–204

  19. [2010]

    In Proceedings of IADIS International conference interfaces and human computer interaction

    Exploring the uncanny valley with Geminoid HI-1 in a real-world ap- plication. In Proceedings of IADIS International conference interfaces and human computer interaction. 121–128

  20. [2022]

    2022), 1767–1786.doi:10.1007/ s12369-022-00904-y

    Technical Methods for Social Robots in Museum Settings: An Overview of the Literature.Int J of Soc Robotics 14, 8 (Oct. 2022), 1767–1786.doi:10.1007/ s12369-022-00904-y

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.