REVIEW 5 major objections 4 minor 45 references
Existential Crisis: A Social Robot's Reason for Being
T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that a humorous personality in a medical questionnaire robot raises users' positive affect and their perception of the robot's anthropomorphism, likeability, and intelligence, based on a 12-person within-subject pilot.
desk verdict Honest course-pilot with a solid systems appendix; the conclusion overreaches the descriptive data and the order counterbalancing is contradictory, but the paper would be a fair short-paper revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central setup is a within-subject comparison of two dialogue policies for the same small humanoid robot, both generated by a locally deployed large language model from different system prompts. The personality prompt instructed a polite, pleasant tone with jokes tied to the user's answers; the control prompt instructed a formal, direct, monotone style with no additional commentary. The robot's spoken output is produced by a text-to-speech module, and the user's spoken answers are transcribed by a speech-to-text model, with the robot signaling that it is listening by changing its eye color. The outcome machinery is two standardized questionnaires, the PANAS mood scale and the Godspeed robot-perception scale, filled out after each condition.
What would settle it
Run a between-subjects version with at least 30 participants per condition, a strictly randomized order, and pre-registered inferential statistics, plus a control that delivers the same jokes in a monotone, task-only style. If the flat-joke robot produces the same high likeability and positive-affect scores as the personality robot, or if the personality advantage disappears when order is balanced, the paper's claim is refuted.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that a visible, humorous personality in a medical questionnaire robot raises self-reported user experience in a short, single-session interaction. The quantitative basis is the gap in mean scores: personality-condition means exceed control means by roughly 1.3 to 1.75 points on the 1-to-5 PANAS positive-affect items, and by similarly visible margins on the Godspeed dimensions of anthropomorphism, likeability, and intelligence. Qualitative responses confirm the manipulation: participants described the personality robot as playful, conversational, funny, and enjoyable, and the control robot as static, boring, and form-like. The authors conclude that users felt more comfortable and more engaged and open during question answering in the personality condition, which they interpret as confirming the hypothesis that contextual humor creates a better experience.
Load-bearing premise
The load-bearing premise is that the higher mean scores in the personality condition are caused by the robot's personality (its jokes and informal tone) rather than by which condition came first, by participants expecting to like the more social robot, or by the fact that the humorous robot also acknowledged each user's answer, none of which were controlled for in this pilot.
Editorial extensions
If this is right
- A medical questionnaire robot that shows humor and an informal tone produces higher self-reported positive affect and higher perceived anthropomorphism, likeability, and intelligence than the same robot in a neutral, task-only mode.
- Users are able to clearly detect the difference between the two interaction styles, and they report feeling more comfortable, relaxed, and open with the humorous robot.
- The humorous robot changed user behavior: participants gave shorter, more concise answers to the formal robot and became more expansive with the humorous one.
- Humor's effect is not uniformly positive: some participants felt the jokes were inappropriate for a medical setting, and one participant responded negatively to a culturally insensitive joke.
- The effect of humor depends on context and user, indicating that robot personality should be tuned to the individual patient rather than applied one-size-fits-all.
Reading between the lines
- Because the two conditions differed in multiple ways at once (tone, jokes, responsiveness to each answer, and perceived engagement), the study cannot isolate humor as the active ingredient; a condition that delivers the same jokes in a flat monotone would separate 'personality' from 'acknowledging the user'.
- The reported mean differences of roughly 1.3 to 1.75 points on a 1-to-5 scale are large enough that a moderately powered replication with proper counterbalancing and inferential statistics could confirm or refute the pattern; the descriptive-only analysis leaves the conclusion open.
- The paper's internal inconsistency about the order in which the two groups encountered the conditions (Section 2.2 says the first group got the neutral robot first; Section 2.4 says the first group started with the personality condition) means order effects could not have been properly controlled even within the within-subject design.
- For healthcare robotics, a testable extension would be an adaptive robot that gauges a patient's reaction to early jokes and dials humor up or down accordingly, potentially capturing the engagement gain while avoiding the reported downside of jokes felt to be inappropriate for a medical setting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a within-subject pilot study comparing a personality-driven NAO robot (with humor and informal tone) against a task-oriented control robot in a medical questionnaire interaction. Twelve students interacted with both versions, and the authors measured emotional state with an abbreviated PANAS and robot perception with Godspeed subscales, supplemented by qualitative responses. The central claim is that the personality condition improved user experience, and the paper concludes that the original hypothesis is confirmed. The manuscript includes full experimental instructions, questionnaires, LLM prompts, and an ethical self-check, and it explicitly labels the study as a pilot with a small sample.
Significance. If the reported effect is real, the result is a modest but useful pilot contribution to humor and personality in human-robot interaction for healthcare questionnaires. The study's strengths are its use of standardized instruments (PANAS, Godspeed), a transparent appendix with questionnaires and prompts, explicit ethical consent, and a reproducible open-source LLM pipeline. The qualitative findings on perceived playfulness versus formality are plausible and consistent with prior HRI work. However, the confirmatory conclusion is not supported by the evidence presented: the study reports only descriptive statistics from twelve participants, the counterbalancing order is described inconsistently, and the manipulation is a bundle of tone, responsiveness, and content differences. These issues are load-bearing for the paper's central causal claim.
major comments (5)
- [§5 versus §2.5] The conclusion states that the results are 'confirming our original hypothesis,' but §2.5 explicitly says the sample is too small for meaningful inferential tests, and no inferential statistics, effect sizes, or confidence intervals are reported anywhere. The quantitative basis in Table 1 and Figures 3–4 is a set of mean differences from N=12, which cannot support a confirmatory causal claim without accounting for within-subject variability, multiple comparisons, and order. Please either report appropriate inferential statistics (or at least effect sizes and confidence intervals) and reclassify the study as exploratory, or substantially temper the wording of the conclusion.
- [§2.2 and §2.4] The description of condition order is internally contradictory. §2.2 states that the first group interacted with the robot without personality first and the other half interacted with the personality condition first. §2.4 states that the first group started with the personality condition followed by the control condition, and only 'after the first six participants' did the second group begin in the control condition. These cannot both be true. Because order effects (practice, fatigue, and contrast) are a real threat in within-subject designs, the actual order must be stated precisely and, ideally, checked in the data for order-by-condition interactions.
- [§2.2–§2.4, manipulation confound] The two conditions differ along multiple dimensions at once: the personality condition includes jokes and witty follow-ups, informal tone, and possibly different LLM-generated content, while the control condition uses formal tone and no additional commentary. Consequently, the observed differences in PANAS positive affect and Godspeed ratings cannot be attributed specifically to 'personality' or 'humor'; they could reflect any component of this bundle, including response relevance, verbosity, or perceived attentiveness. Please either decompose the manipulation or reframe the research question as comparing two holistic interaction styles, and adjust the causal language in §5 accordingly.
- [Appendix C, Part 3] The manipulation check question 'Did you perceive Robot A as more task-oriented and Robot B as more social and engaging?' is leading because it states the expected distinction and is combined with preference questions. Under these conditions, the qualitative finding that all participants detected a difference is unsurprising and does not independently validate the manipulation. A more neutral manipulation check, such as coding free descriptions, or indirect behavioral measures would strengthen the claim that participants perceived the intended manipulation without being cued.
- [§4.1 and §3.2] The limitations section correctly acknowledges the small, homogeneous sample, the participants' prior familiarity with NAO robots, the short interaction duration, and occasional nonsensical LLM outputs. These limitations, combined with the post-hoc self-report nature of the qualitative data, further undermine the strength of the confirmatory statement in §5. The qualitative quotes are informative as illustrations, but they should not be used as primary evidence for a causal effect of personality on user experience.
minor comments (4)
- [§2.3, references] The text attributes the claim about reducing perceived waiting time to 'D. Kang et al.', but reference [18] is Pelikan and Hofstetter; please correct the citation or add the intended reference.
- [Throughout] There are several typos and formatting issues: 'Thereareexistingoccurrences' in the introduction, 'T able 1' and '0 .669' / '1 .138' in Table 1, 'likability' versus 'likeability', 'ideologues' in §4.1 (likely 'idiosyncrasies'), and 'non-nonsensical' in §4.1 (likely 'nonsensical'). These should be corrected.
- [Throughout] The robot name is written inconsistently as both 'NAO' and 'Nao'; please standardize to one form.
- [Figures 3 and 4] Figures 3 and 4 are referenced in the text, but the full text does not show the actual plots; please ensure the figures are included with axis labels and legends, and add explicit captions.
Circularity Check
No circularity: the study is an empirical user study with externally standardized outcome measures, and none of its conclusions are derived by construction from its inputs.
full rationale
The paper makes no formal derivation, fits no model to the outcome data, and validates no prediction against a fitted quantity. Its outcome measures are the standardized PANAS and Godspeed questionnaires (References 15 and 9), which are external instruments and are filled out by participants after each interaction; the reported means in Table 1 and Figures 3-4 are descriptive statistics of those independent measurements. The only calibrated parameters mentioned are the audio-level threshold and pause duration used for speech segmentation (Section 2.3), which are unrelated to the PANAS or Godspeed scores and do not enter the reported results. There is no cited 'uniqueness theorem', no prior-work citation that is load-bearing for the empirical claim, and no renamed known result: the paper compares two concrete robot behaviors and reports participant ratings. The conclusion in Section 5 that users felt more comfortable and engaged in the personality condition is an empirical generalization from the descriptive data and qualitative responses; whether that inference is statistically warranted is a methodological validity concern (no inferential statistics, inconsistent counterbalancing description), not a circularity concern. The hypothesis is not defined in terms of the outcome measure, and the outcome measure is not constructed from the hypothesis. Thus the derivation chain is not circular: no claim reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (2)
- Voice activity detection amplitude threshold =
350 (unitless amplitude level)
- Silence duration threshold =
1.5 seconds
assumptions (3)
- domain assumption PANAS and Godspeed are valid measures of emotional state and robot perception in this setting
- domain assumption The two conditions differ only in personality and humor, not in other interaction qualities
- domain assumption Participants' self-reports reflect the robot's behavior rather than order or social desirability
Cite this review
Pith. "Pith review of Existential Crisis: A Social Robot's Reason for Being." pith.science (2026). https://pith.science/paper/BK33V2UA
@misc{pith2026250103376,
author = {Pith},
title = {Pith review of: Existential Crisis: A Social Robot's Reason for Being},
year = {2026},
howpublished = {\url{https://pith.science/paper/BK33V2UA}},
note = {Machine review of arXiv:2501.03376}
}
read the original abstract
As Robots become ever more important in our daily lives there's growing need for understanding how they're perceived by people. This study aims to investigate how the user perception of robots is influenced by displays of personality. Using LLMs and speech to text technology, we designed a within-subject study to compare two conditions: a personality-driven robot and a purely task-oriented, personality-neutral robot. Twelve participants, recruited from Socially Intelligent Robotics course at Vrije Universiteit Amsterdam, interacted with a robot Nao tasked with asking them a set of medical questions under both conditions. After completing both interactions, the participants completed a user experience questionnaire measuring their emotional states and robot perception using standardized questionnaires from the SRI and Psychology literature.
Figures
Reference graph
Works this paper leans on
-
[1]
OpenAI. (2023). GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 . https://doi.org/10.48550/arXiv.2303.08774
-
[2]
Dubey, Abhimanyu & Jauhri, Abhinav & Pandey, Abhinav & Kadian, Abhishek & Al-Dahle, Ahmad & Letman, Aiesha & Mathur, Akhil & Schelten, Alan & Yang, Amy & Fan, Angela & Goyal, Anirudh & Hartshorn, Anthony & Yang, Aobo & Mitra, Archi & Sravankumar, Archie & Korenev, Artem & Hinsvark, Arthur & Rao, Arun & Zhang, Aston & Zhao, Zhiwei. (2024). The Llama 3 Herd...
-
[3]
leng2024longcontextragperformance, title=Long Context RAG Performance of Large Language Models, author=Quinn Leng and Jacob Portes and Sam Havens and Matei Zaharia and Michael Carbin, year=2024, eprint=2411.03538, archivePrefix=arXiv, primaryClass=cs.LG, https://arxiv.org/abs/2411.03538
arXiv 2024
-
[4]
Aldebaran. (2014). Technical site. https://corporate-internal-prod. aldebaran.com/en/pepper
work page 2014
-
[5]
Mileounis, A., Cuijpers, R.H., Barakova, E.I. (2015). Creating Robots with Per- sonality: The Effect of Personality on Social Intelligence. In: Ferrández Vicente, J., Álvarez-Sánchez, J., de la Paz López, F., Toledo-Moreo, F., Adeli, H. (eds) Artificial Computation in Biology and Medicine. IWINAC 2015. Lecture Notes in Computer Science(), vol 9107. Spring...
work page 2015
-
[6]
Minsky, M. (2006).The Emotion Machine: Commonsense Thinking, Artificial In- telligence, and the Future of the Human Mind. Simon and Schuster
work page 2006
-
[7]
Picard, Rosalind W. (1997). Affective Computing. MIT Press, Cambridge, MA. ISBN: 978-0-262-16170-2
work page 1997
-
[8]
Mehrabian, A., & Ferris, S. R. (1967). Inference of attitudes from nonverbal com- munication in two channels. Journal of Consulting Psychology, 31(3), 248–252. https://doi.org/10.1037/h0024648
doi:10.1037/h0024648 1967
Show all 45 references
-
[9]
Bartneck, C., Kulić, D., Croft, E., & Zoghbi, S. (2009). Measurement instru- ments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots.International Journal of Social Robotics, 1(1), 71–81. https://doi.org/10.1007/s12369-008-0001-3
2009 doi
-
[10]
M., Wyman, A
Carpinella, C. M., Wyman, A. B., Perez, M. A., & Stroessner, S. J. (2017). The robotic social attributes scale (RoSAS): Development and validation. In Pro- ceedings of the 2017 ACM/IEEE International Conference on Human-Robot In- teraction (HRI ’17) (pp. 254–262). Vienna, Aust...
2017
-
[11]
Heerink, M., Kröse, B., Evers, V., & Wielinga, B. (2010). Assessing acceptance of assistive social agent technology by older adults: The Almere Model. Inter- national Journal of Social Robotics, 2(4), 361–375. https://doi.org/10.1007/ s12369-010-0068-5
2010
-
[12]
S., & Scassellati, B
Admoni, H., Dragan, A., Srinivasa, S. S., & Scassellati, B. (2014). Deliberate delays during robot-to-human handovers improve compliance with gaze communication. In 2014 9th ACM/IEEE International Conference on Human-Robot Interaction (HRI) (pp. 49–56). Bielefeld, Germany.http...
2014 doi
-
[13]
Thomas Orden, Mike EU Ligthart, Koen Hindriks, Thomas Wiggers, and Karen Chiang. [n. d.]. The Social Interaction Cloud (SIC) https://socialrobotics. atlassian.net/wiki/spaces/CBSR/overview
-
[14]
Picard, R. W. (1997).Affective Computing. MIT Press. Existential Crisis: A Social Robot’s Reason for Being 13
1997
-
[15]
A., & Tellegen, A
Watson, D., Clark, L. A., & Tellegen, A. (1988). Development and validation of brief measures of positive and negative affect: The PANAS scales.Journal of Personality and Social Psychology, 54(6), 1063–1070.https://doi.org/10.1037/ 0022-3514.54.6.1063
1988
-
[16]
Broadbent, E., Stafford, R., & MacDonald, B. (2013). Acceptance of healthcare robotsfortheolderpopulation:Reviewandfuturedirections. International Journal of Social Robotics, 5(3), 319–330.https://doi.org/10.1007/s12369-013-0196-4
2013 doi
- [17]
-
[18]
Hannah Pelikan and Emily Hofstetter. 2023. Managing Delays in Human-Robot Interaction. ACM Trans. Comput.-Hum. Interact. 30, 4, Article 50 (August 2023), 42 pages. https://doi.org/10.1145/3569890
2023 doi
-
[19]
H., & Torreira, F
Kendrick, K. H., & Torreira, F. (2015). The timing and construction of preference: A quantitative study. Discourse Processes, 52(4), 255-289
2015
-
[20]
John C. Meyer, Humor as a Double-Edged Sword: Four Functions of Humor in Communication, Communication Theory, Volume 10, Issue3, 1 August 2000,Pages 310–331 https://doi.org/10.1111/j.1468-2885.2000.tb00194.x 14 Medgyesy et al. A Appendix: Participant instruction Title of the S...
-
[21]
I understand that my participation is voluntary and that I can withdraw at any time without penalty
-
[22]
I understand that my responses will be anonymous and that no personal data will be collected that could identify me
-
[23]
I understand that the data collected will be used solely for research purposes and may be published in academic journals
-
[24]
Consent By signing below, you indicate that you have read and understood the informa- tion provided and agree to participate in this study
I have been informed about the nature of the study and the procedures involved. Consent By signing below, you indicate that you have read and understood the informa- tion provided and agree to participate in this study. Participant Name: Participant Signature: Date: Thank you ...
-
[25]
In your own words, describe your overall experience with the robot
-
[26]
What stood out to you the most during your interaction with the robot?
-
[27]
How did the robot’s behavior affect your engagement during the interaction?
-
[28]
Part 2: Numerical Questions – User Experience and Perception A
Did you feel the robot understood you? Please explain. Part 2: Numerical Questions – User Experience and Perception A. How Do You Feel Instructions: Indicate to what extent you feel each emotion right now, after interacting with the robot. Use a scale from 1 (Very slightly or ...
-
[29]
Interested ............................................. (1–5)
-
[30]
Excited ............................................... (1–5)
-
[31]
Upset ................................................. (1–5)
-
[32]
Strong ................................................ (1–5)
-
[33]
Enthusiastic ........................................... (1–5)
-
[34]
Distressed ............................................. (1–5)
-
[35]
Determined ........................................... (1–5)
-
[36]
Nervous ............................................... (1–5)
-
[37]
Alert .................................................. (1–5)
-
[38]
Inspired ............................................... (1–5) B. How Do You View the Robot Instructions: For each pair of adjectives, select a number from 1 to 5 that best describes your perception of the robot. 16 Medgyesy et al
-
[39]
Anthropomorphism – Fake ..................1 — 2 — 3 — 4 — 5 ..................Natural – Machinelike .............1 — 2 — 3 — 4 — 5 .............Humanlike – Unconscious .............1 — 2 — 3 — 4 — 5 .............Conscious – Artificial ................1 — 2 — 3 — 4 — 5 ...........
-
[40]
1 — 2 — 3 — 4 — 5
Likeability – Dislike .................. 1 — 2 — 3 — 4 — 5 .................. Like – Unfriendly ...............1 — 2 — 3 — 4 — 5 ...............Friendly – Unkind ..................1 — 2 — 3 — 4 — 5 ..................Kind – Unpleasant ..............1 — 2 — 3 — 4 — 5 ..............
-
[41]
1 — 2 — 3 — 4 — 5
Intelligence – Incompetent ............1 — 2 — 3 — 4 — 5 ............Competent – Ignorant ............ 1 — 2 — 3 — 4 — 5 ............ Knowledgeable – Irresponsible ............1 — 2 — 3 — 4 — 5 ............Responsible – Unintelligent ............ 1 — 2 — 3 — 4 — 5 ...............
-
[42]
Did you perceive Robot A as more task-oriented and Robot B as more social and engaging? □ Yes □ No □ Unsure
-
[43]
Which robot did you prefer interacting with? □ Robot A: Task-oriented □ Robot B: Social and engaging
-
[44]
Please explain why you preferred this robot
-
[45]
question
Would you recommend your preferred robot to others? (1–5) – Scale from 1 (Definitely not) to 5 (Definitely yes) Part 4: Additional Notes Please provide any additional comments or suggestions about your experiences with the robots. [Space for open-ended response] D LLM prompts ...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.