REVIEW 4 major objections 4 minor 28 references
Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Android Andrea's simulated emotions fail to boost visitor acceptance
desk verdict Honest null result for emotion simulation on a museum android, but the day-condition confound means the negative finding is real but not as clean as it looks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The android Andrea, a 52-actuator humanlike robot, runs a software pipeline combining Whisper for speech-to-text, a ChatGPT 4.1 assistant with retrieval-augmented generation for dialogue, XTTS v2 for multilingual speech synthesis, FaceXHubert for lip-sync, and posenet for visitor tracking. The experimental variable is the emotion component: no emotion, ChatGPT-generated emotion labels, or WASABI's computational affect dynamics, all expressed through validated static facial expressions. Acceptance is measured with an extended TAM2 questionnaire (intention to use, perceived usefulness, perceived ease of use) plus a usefulness scale and three perceived-emotionality items.
What would settle it
Compare acceptance ratings from the two no-emotion days (day 1 and day 4): if they differ substantially, day-level confounds are present and the emotion-condition comparison is unreliable. Separately, annotate video of the robot's face per condition; if the ChatGPT and WASABI conditions show no more visible emotional expressions than the no-emotion condition, the emotion manipulation failed and the null result says nothing about emotion simulation.
Extended reading notes
Core claim
In this ecologically valid but uncontrolled museum setting, equipping Andrea with either of two emotion-simulation pipelines did not make visitors more accepting of the robot, and the simulated emotions were not noticeable enough to move perceived-emotionality ratings. The no-emotion condition was rated at least as useful and easy to use as the emotional conditions, and significantly more so than the ChatGPT-emotion condition on perceived usefulness and perceived ease of use; WASABI showed no differences, but its sample was small (N=14). The authors interpret the negative manipulation check as devaluing the emotion-implementation approach.
Load-bearing premise
That the three emotion conditions differed only in the emotion simulation—rather than in which calendar day they ran, the visitors who happened to be in the museum, how actively experimenters recruited participants, or whether the robot's facial expressions actually showed the intended emotions; the paper's own manipulation-check failure and log inspection make the last point especially uncertain.
Editorial extensions
If this is right
- For short, task-oriented public interactions, adding simulated emotion to an android's conversation is not automatically an acceptance benefit and may slightly reduce perceived usefulness and ease of use.
- Emotion simulation that fails a conscious manipulation check cannot be expected to influence acceptance, so field studies of robot emotion need behavioral or physiological verification that the manipulation was perceptible.
- The non-emotional version scored highest on ease of use and usefulness, suggesting visitors may value clarity and reliability over emotional mimicry in museum information kiosks.
- The WASABI condition's small sample (N=14) leaves that architecture's effect unresolved; the data cannot support any claim about WASABI specifically.
- Running a fully autonomous multilingual android in a museum for six days is feasible, and most visitors arrived without knowing the robot was there.
Reading between the lines
- If the emotion labels were rarely non-happy (as the paper's log inspection suggests) and emotions were shown only through the face while the robot was not speaking, the emotion signal may have been too weak or too brief to register; a redesign that times emotional expressions to pauses or adds vocal emotion cues might yield a different result.
- The day-level confounding (conditions on different days, one day excluded, varying experimenter recruiting) means this should be read as a suggestive null result rather than evidence that emotion simulation is useless; a within-day counterbalanced or longer random-order deployment would tighten the test.
- A testable extension would compare the same android with and without emotion in a repeated one-on-one interaction over minutes, where visitors have time to notice and respond to emotional states, rather than in a one-shot public encounter.
- The slight preference for the non-emotional robot hints that an android's humanlike appearance may already set social expectations that emotional behavior, if imperfectly executed, does not meet; matching the quality of the emotion display to the realism of the face may matter more than whether emotion is simulated at all.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a field study in which the android robot Andrea interacted autonomously with 73 museum visitors under three emotion conditions: a no-emotion baseline (ChatGPTpure), ChatGPT-generated emotion labels, and the WASABI emotion architecture. Visitors completed an extended TAM2 questionnaire. Kruskal-Wallis tests found no significant differences for Intention to Use or Usefulness Today, a significant difference for Perceived Ease of Use (p=.033), and a near-significant difference for Perceived Usefulness (p=.056); pairwise Dwass-Steel-Critchlow-Fligner tests showed that the ChatGPTpure condition outscored the ChatGPTemotion condition on PU and PEOU. The manipulation-check items on perceived emotionality showed no significant condition differences. The authors conclude that the two emotion-simulation approaches did not improve acceptance and were not consciously detectable.
Significance. If the conclusions were valid, this would be a useful negative result for the design of affective android systems in public settings, and the rich multilingual, real-world deployment data are a genuine contribution. The manuscript is commendably transparent: it reports normality checks, non-parametric tests, effect sizes, box plots, a manipulation check, and the hardware failure that removed day 6. However, the central causal claim about the emotion conditions is compromised by a complete day-condition confound and by the very small WASABI sample. As it stands, the paper provides an honest exploratory field report, but the abstract's strong negative conclusion is not yet supported by the design.
major comments (4)
- [§4, §5] Condition is perfectly confounded with calendar day: no-emotion on days 1/4, ChatGPTemotion on days 2/5, WASABI on day 3 (day 6 excluded). The abstract's claim that emotion simulation 'did not yield any positive effects' is a causal statement that requires exchangeability of visitor populations, experimenter recruiting behavior, hardware state, and exhibit context across days. The Method section explicitly states that recruiting style depended on the experimenter on duty, and the robot's facial expressions visibly degraded by day 6. The significant PEOU/PU contrasts between ChatGPTemotion and ChatGPTpure, and the null manipulation check, could each be produced by day-level differences. The authors never mention this confound; a day-level permutation analysis, or at least an explicit argument with supporting data, is needed to support the central negative conclusion.
- [Table 1, §6] The WASABI condition has N=14 and is inseparable from day 3. The authors acknowledge that the condition is 'likely underpowered,' but they still include WASABI in the abstract's general conclusion that 'these first two approaches' had no positive effects. With N=14, the null result for WASABI is weakly informative, and the day-3-only assignment makes it impossible to distinguish condition effects from day effects. Either the conclusions should be restricted to the ChatGPTemotion arm (with the confound still unresolved), or the authors should provide a sensitivity/power analysis and day-level evidence.
- [Table 2, §5] The reporting of effect sizes is inconsistent as written. Table 2 lists epsilon-squared values of 0.05, 0.08, and 0.09, but §5 describes these as 'moderate effect size (ε > 0.6)'. These numbers do not match; epsilon-squared of 0.09 is generally considered small. In addition, pairwise comparisons for PU were conducted despite the overall Kruskal-Wallis p = .056, which is not below α = .05. The authors should justify this decision or treat the PU pairwise result as exploratory; otherwise the claims of 'significant' advantages for ChatGPTpure over ChatGPTemotion are overstated.
- [Abstract, Table 5] The claim that the emotion manipulation was 'not detectable on a conscious level' rests on null results for three questionnaire items with unequal group sizes and no equivalence testing or sensitivity analysis. A null p-value cannot by itself establish that the conditions were indistinguishable, particularly with N=14 in one arm. The authors should either soften this to 'no evidence of a conscious difference in this sample' or report a sensitivity analysis showing the smallest detectable effect.
minor comments (4)
- [§5] Typographical errors: 'responeded' and 'It responeded most often' should be corrected; 'to a very small extend' should be 'extent'; there is a stray 'W ASABI' in the age distribution paragraph.
- [§2] The related-work paragraph on previous Andrea results (2023) repeats information given later in §3; consider merging to avoid redundancy.
- [Figures 4–7] The box plots would be more informative if they showed individual data points or at least the group N alongside the medians; this is especially relevant given the uneven group sizes.
- [Table 2] The epsilon-squared values are labelled 'ε²' but the text uses 'ε'; unify the notation and report confidence intervals for the effect sizes where possible.
Circularity Check
No significant circularity: the empirical comparison uses independent TAM2 and emotion-perception questionnaire outcomes; the self-citations are technical context, not load-bearing.
full rationale
The derivation chain runs from an implemented manipulation (no emotion vs ChatGPT-assigned emotion labels vs WASABI dynamics) to visitor questionnaire responses. The outcome variables — ITU, PU, PEOU, UT, and the three perceived-emotionality items — are standard or separately worded survey items, not values computed from the emotion-generation parameters. No equation fits a parameter to the reported acceptance or emotion data and then presents it as a prediction; Kruskal-Wallis and Dwass-Steel-Critchlow-Flinger tests are applied directly to raw ratings. The manipulation check is a distinct set of items, so its null result is an empirical finding rather than an artifact of the definition of the emotion conditions. The paper cites the authors' prior work for the robot software ([10,11]), the WASABI simulator ([2]), validated facial expressions ([15]), and the voice ([19,20]), but these citations support the construction of the stimulus, not the conclusion that the emotion conditions failed to improve acceptance or were not consciously detected. Even the paper's own limitation paragraph (WASABI N=14, degraded hardware, limited gestures) and the day-level confounding between condition and calendar day are threats to internal validity and statistical power, not circularity: they concern whether the observed null results are trustworthy, not whether the results are equivalent to their inputs by construction. No self-definitional, fitted-input-as-prediction, uniqueness-import, or ansatz-smuggling pattern is present, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- None in the statistical claims
assumptions (4)
- domain assumption ChatGPT 4.1's emotion label generation and the WASABI valence-prompting pipeline function as an adequate 'emotion simulation' implementation.
- domain assumption The TAM2 questionnaire, translated from [23]/[26], measures the intended latent constructs in a museum context with visitors of mixed ages and nationalities.
- domain assumption Visitors interacted with only one condition (between-groups validity).
- ad hoc to paper Day-assignment confounding is harmless (visitor populations are exchangeable across days).
Cite this review
Pith. "Pith review of Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum." pith.science (2026). https://pith.science/paper/AT2CXSYJ
@misc{pith2026260716428,
author = {Pith},
title = {Pith review of: Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum},
year = {2026},
howpublished = {\url{https://pith.science/paper/AT2CXSYJ}},
note = {Machine review of arXiv:2607.16428}
}
read the original abstract
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with visitors, fully autonomously. Building on previously gathered qualitative results, the robot was now capable of engaging in multi-lingual conversation with the visitors about the museum context. The robot was prepared with context information about the museum in general and its surrounding exhibits this time. The robot featured a slightly artificial sounding voice that was previously evaluated as congruent with its gender-ambiguous but very humanlike design. Three experimental conditions were implemented, in which either (1) the robot simulated no emotions, (2) the robots emotions were determined by ChatGPT 4.1, or (3) the WASABI emotion simulation architecture simulated the robot's emotion dynamics. An extended version of the TAM2 questionnaire was employed to let 73 visitors report on several factors of their opinion about the android robot Andrea after having experienced it. In result, the statistical analysis suggests that these first two approaches to implementing emotions into the chat architecture of our android robot Andrea did not yield any positive effects on the subjective evaluations by the visitors and were not detectable on a conscious level.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Evangelia Baka, Nidhi Mishra, Emmanouil Sylligardos, and Nadia Magnenat- Thalmann. 2022. Social Robots and Digital Humans as Job Interviewers: A Study of Human Reactions Towards a More Naturalistic Interaction. InHuman- Computer Interaction. Technological Innovation, Masaaki Kurosu (Ed.). Springer International Publishing, Cham, 455–474.doi:10.1007/978-3-...
-
[2]
Becker-Asano
C. Becker-Asano. 2014. WASABI for affect simulation in human-computer in- teraction. In Proc. on Emotion Representations and Modelling for HCI Systems . Springer, Sydney, Australia
2014
-
[3]
Christian Becker-Asano, Kohei Ogawa, Shuichi Nishio, and Hiroshi Ishiguro
-
[4]
KarstenBernsandAshitaAshok.2024. “YouScareMe”:TheEffectsofHumanoid Robot Appearance, Emotion, and Interaction Skills on Uncanny Valley Phenom- enon. Actuators 13, 10 (Oct. 2024), 419.doi:10.3390/act13100419
-
[5]
Felix Carros, Berenike Bürvenich, Ryan Browne, Yoshio Matsumoto, Gabriele Trovato, Mehrbod Manavi, Keiko Homma, Toshimi Ogawa, Rainer Wieching, and Volker Wulf. 2022. Not that Uncanny After All? An Ethnographic Study on Android Robots Perception of Older Adults in Germany and Japan. InSocial Robotics. Vol. 13818. Springer Nature Switzerland, Cham, 574–586...
2022
-
[6]
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Lo- gan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, and Ju- lian Weber. 2024. XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model. InInterspeech 2024. 4978–4982. doi:10.21437/Interspeech.2024-2016
-
[7]
Haozhe Chen, Run Chen, and Julia Hirschberg. 2024. EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control. arXiv:2410.00316 [cs.CL] https: //arxiv.org/abs/2410.00316
arXiv 2024
-
[8]
Norina Gasteiger, Mehdi Hellou, and Ho Seok Ahn. 2021. Deploying social robotsinmuseumsettings:Aquasi-systematicreviewexploringpurposeandac- ceptability. International Journal of Advanced Robotic Systems 18, 6 (Nov. 2021), 17298814211066740. doi:10.1177/17298814211066740 Publisher: SAGE Publica- tions
Show all 28 references
-
[9]
Kazi Injamamul Haque and Zerrin Yumak. 2023. FaceXHuBERT: Text- less Speech-driven E(X)pressive 3D Facial Animation Synthesis Using Self- Supervised Speech Representation Learning. InProceedings of the 25th Interna- tional Conference on Multimodal Interaction (ICMI ’23) . Asso...
2023
-
[10]
Marcel Heisler and Christian Becker-Asano. 2023. An Android Robot Head as Embodied Conversational Agent. InISR Europe 2023; 56th International Sympo- sium on Robotics. 93–99. https://ieeexplore.ieee.org/document/10363058
2023
-
[11]
Marcel Heisler and Christian Becker-Asano. 2025. Conversations with An- drea: Visitors’ Opinions on Android Robots in a Museum. In2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO- MAN). IEEE, Eindhoven, Netherlands, 112–119. doi:10.1109/...
2025
-
[12]
Marcel Heisler, Stefan Kopp, and Christian Becker-Asano. 2023. Making an Android Robot Head Talk. In2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . 1837–1842. doi:10.1109/RO- MAN57019.2023.10309532 ISSN: 1944-9437
2023
-
[13]
Mehdi Hellou, JongYoon Lim, Norina Gasteiger, Minsu Jang, and Ho Seok Ahn
-
[14]
Hangyeol Kang, Thiago Freitas, Maher Ben Moussa, and Nadia Magnenat Thal- mann. 2026. Affective and Conversational Predictors of Re-Engagement in Human–Robot Interactions: A Student-Centered Study with a Humanoid Social Robot. Int J of Soc Robotics18,3(March2026),41. doi:10.10...
2026 doi
-
[15]
Amelie Kassner and Christian Becker-Asano. 2023. Comparing an android head with its digital twin regarding the dynamic expression of emotions. In2023 11th International Conference on Affective Computing and Intelligent Interaction Work- shops and Demos (ACIIW)
2023
-
[16]
IntelligentConversational Android ERICA Applied to Attentive Listening and Job Interview.http://arxiv
TatsuyaKawahara,KojiInoue,andDiveshLala.2021. IntelligentConversational Android ERICA Applied to Attentive Listening and Job Interview.http://arxiv. org/abs/2105.00403
2021 arXiv
-
[17]
Dacher Keltner and Daniel T. Cordaro. 2017. Understanding Multimodal Emo- tional Expressions: Recent Advances in Basic Emotion Theory. InThe Science of Facial Expression, James A. Russell and Jose Miguel Fernandez Dols (Eds.). Ox- ford University Press.doi:10.1093/acprof:oso/9...
2017
-
[18]
Alex Kendall, Matthew Grimes, and Roberto Cipolla. 2015. PoseNet: A Convo- lutional Network for Real-Time 6-DOF Camera Relocalization. In2015 IEEE Intl. Conf. on Computer Vision (ICCV) . 2938–2946. doi:10.1109/ICCV.2015.336 ISSN: 2380-7504
2015 doi
-
[19]
YourRobot,MyVoice: Enhancing Android Robot Likability through Personalization by Cloning the User’s Voice
Johanna Magdalena Kuch, Marcel Heisler, Stina Klein, Silvan Mertes, Lennart Eing,ElisabethAndré,andChristianBecker-Asano.2025. YourRobot,MyVoice: Enhancing Android Robot Likability through Personalization by Cloning the User’s Voice. In2025 34th IEEE International Conference o...
2025
-
[20]
EvaluatingGenderAmbigu- ity,NoveltyandAnthropomorphisminHummingandTalkingVoicesforRobots
Johanna Magdalena Kuch, Jauwairia Nasir, Silvan Mertes, Ruben Schlagowski, ChristianBecker-Asano,andElisabethAndré.2024. EvaluatingGenderAmbigu- ity,NoveltyandAnthropomorphisminHummingandTalkingVoicesforRobots. In 2024 33rd IEEE International Conference on Robot and Human Inte...
2024
-
[21]
Martina Mara and Markus Appel. 2015. Science fiction reduces the eeriness of androidrobots:Afieldexperiment. Computers in Human Behavior48(July2015), 156–162. doi:10.1016/j.chb.2015.01.007
2015 doi
-
[22]
Now Loading…
Shushi Namba, Wataru Sato, Saori Namba, Alexander Diel, Carlos Ishi, and Takashi Minato. 2024. How an Android Expresses “Now Loading…”: Examin- ing the Properties of Thinking Faces. Int J of Soc Robotics 16, 8 (Aug. 2024), 1861–1877. doi:10.1007/s12369-024-01163-9
2024 doi
-
[23]
Thomas Olbrecht. 2010. Akzeptanz von E-Learning. Eine Auseinandersetzung mit dem Technologieakzeptanzmodell zur Analyse individueller und sozialer Ein- flussfaktoren. Univ., Jena. 215 S. pages.http://www.db-thueringen.de/servlets/ DerivateServlet/Derivate-21996/Olbrecht/Disser...
2010
-
[24]
Astrid M Rosenthal-von der Pütten, Nicole C Krämer, Christian Becker-Asano, Kohei Ogawa, Shuichi Nishio, and Hiroshi Ishiguro. 2014. The uncanny in the wild. Analysis of unscripted human–android interaction in the field.Interna- tional Journal of Social Robotics 6 (2014), 67–83
2014
-
[25]
Nadinethe Social Robot: Three Case Studies in Everyday Life
NadiaMagnenatThalmann,NidhiMishra,andGauriTulsulkar.2021. Nadinethe Social Robot: Three Case Studies in Everyday Life. InSocial Robotics. Springer International Publishing, Cham, 107–116.doi:10.1007/978-3-030-90525-5_10
2021 doi
-
[26]
Viswanath Venkatesh and Fred D. Davis. 2000. A Theoretical Extension of the Technology Acceptance Model: Four Longitudinal Field Studies.Management Science 46, 2 (2000), 186–204
2000
-
[2010]
In Proceedings of IADIS International conference interfaces and human computer interaction
Exploring the uncanny valley with Geminoid HI-1 in a real-world ap- plication. In Proceedings of IADIS International conference interfaces and human computer interaction. 121–128
-
[2022]
2022), 1767–1786.doi:10.1007/ s12369-022-00904-y
Technical Methods for Social Robots in Museum Settings: An Overview of the Literature.Int J of Soc Robotics 14, 8 (Oct. 2022), 1767–1786.doi:10.1007/ s12369-022-00904-y
2022
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.