Pith. sign in

REVIEW 5 major objections 4 minor 45 references

Existential Crisis: A Social Robot's Reason for Being

T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that a humorous personality in a medical questionnaire robot raises users' positive affect and their perception of the robot's anthropomorphism, likeability, and intelligence, based on a 12-person within-subject pilot.

desk verdict Honest course-pilot with a solid systems appendix; the conclusion overreaches the descriptive data and the order counterbalancing is contradictory, but the paper would be a fair short-paper revision. read the letter →

arxiv 2501.03376 v1 pith:BK33V2UA submitted 2025-01-06 cs.RO cs.AIcs.HC

classification cs.ROcs.AIcs.HC
keywords socialrobotshuman-robotinteractionrobotpersonalityhumoruserexperiencelargelanguagemodelsmedicalquestionnairewithin-subjectdesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a pilot study asking whether a robot that administers medical questions gives people a better experience when it shows a personality, specifically a sense of humor, than when it behaves as a neutral, task-only interviewer. Twelve participants each talked with the same humanoid robot twice, once in a 'personality' condition where it answered with witty, context-aware remarks and an informal tone, and once in a control condition where it read the same questions formally and added nothing. After each interaction, the participants filled in the PANAS mood questionnaire and the Godspeed robot-perception questionnaire. The paper claims that all users noticed the behavioral difference, felt more comfortable and more engaged in the personality condition, and gave it higher mean scores on positive affect, anthropomorphism, likeability, and perceived intelligence, supporting the hypothesis that personality improves user experience.

What carries the argument

The central setup is a within-subject comparison of two dialogue policies for the same small humanoid robot, both generated by a locally deployed large language model from different system prompts. The personality prompt instructed a polite, pleasant tone with jokes tied to the user's answers; the control prompt instructed a formal, direct, monotone style with no additional commentary. The robot's spoken output is produced by a text-to-speech module, and the user's spoken answers are transcribed by a speech-to-text model, with the robot signaling that it is listening by changing its eye color. The outcome machinery is two standardized questionnaires, the PANAS mood scale and the Godspeed robot-perception scale, filled out after each condition.

What would settle it

Run a between-subjects version with at least 30 participants per condition, a strictly randomized order, and pre-registered inferential statistics, plus a control that delivers the same jokes in a monotone, task-only style. If the flat-joke robot produces the same high likeability and positive-affect scores as the personality robot, or if the personality advantage disappears when order is balanced, the paper's claim is refuted.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that a visible, humorous personality in a medical questionnaire robot raises self-reported user experience in a short, single-session interaction. The quantitative basis is the gap in mean scores: personality-condition means exceed control means by roughly 1.3 to 1.75 points on the 1-to-5 PANAS positive-affect items, and by similarly visible margins on the Godspeed dimensions of anthropomorphism, likeability, and intelligence. Qualitative responses confirm the manipulation: participants described the personality robot as playful, conversational, funny, and enjoyable, and the control robot as static, boring, and form-like. The authors conclude that users felt more comfortable and more engaged and open during question answering in the personality condition, which they interpret as confirming the hypothesis that contextual humor creates a better experience.

Load-bearing premise

The load-bearing premise is that the higher mean scores in the personality condition are caused by the robot's personality (its jokes and informal tone) rather than by which condition came first, by participants expecting to like the more social robot, or by the fact that the humorous robot also acknowledged each user's answer, none of which were controlled for in this pilot.

Editorial extensions

If this is right

  • A medical questionnaire robot that shows humor and an informal tone produces higher self-reported positive affect and higher perceived anthropomorphism, likeability, and intelligence than the same robot in a neutral, task-only mode.
  • Users are able to clearly detect the difference between the two interaction styles, and they report feeling more comfortable, relaxed, and open with the humorous robot.
  • The humorous robot changed user behavior: participants gave shorter, more concise answers to the formal robot and became more expansive with the humorous one.
  • Humor's effect is not uniformly positive: some participants felt the jokes were inappropriate for a medical setting, and one participant responded negatively to a culturally insensitive joke.
  • The effect of humor depends on context and user, indicating that robot personality should be tuned to the individual patient rather than applied one-size-fits-all.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the two conditions differed in multiple ways at once (tone, jokes, responsiveness to each answer, and perceived engagement), the study cannot isolate humor as the active ingredient; a condition that delivers the same jokes in a flat monotone would separate 'personality' from 'acknowledging the user'.
  • The reported mean differences of roughly 1.3 to 1.75 points on a 1-to-5 scale are large enough that a moderately powered replication with proper counterbalancing and inferential statistics could confirm or refute the pattern; the descriptive-only analysis leaves the conclusion open.
  • The paper's internal inconsistency about the order in which the two groups encountered the conditions (Section 2.2 says the first group got the neutral robot first; Section 2.4 says the first group started with the personality condition) means order effects could not have been properly controlled even within the within-subject design.
  • For healthcare robotics, a testable extension would be an adaptive robot that gauges a patient's reaction to early jokes and dials humor up or down accordingly, potentially capturing the engagement gain while avoiding the reported downside of jokes felt to be inappropriate for a medical setting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper reports a within-subject pilot study comparing a personality-driven NAO robot (with humor and informal tone) against a task-oriented control robot in a medical questionnaire interaction. Twelve students interacted with both versions, and the authors measured emotional state with an abbreviated PANAS and robot perception with Godspeed subscales, supplemented by qualitative responses. The central claim is that the personality condition improved user experience, and the paper concludes that the original hypothesis is confirmed. The manuscript includes full experimental instructions, questionnaires, LLM prompts, and an ethical self-check, and it explicitly labels the study as a pilot with a small sample.

Significance. If the reported effect is real, the result is a modest but useful pilot contribution to humor and personality in human-robot interaction for healthcare questionnaires. The study's strengths are its use of standardized instruments (PANAS, Godspeed), a transparent appendix with questionnaires and prompts, explicit ethical consent, and a reproducible open-source LLM pipeline. The qualitative findings on perceived playfulness versus formality are plausible and consistent with prior HRI work. However, the confirmatory conclusion is not supported by the evidence presented: the study reports only descriptive statistics from twelve participants, the counterbalancing order is described inconsistently, and the manipulation is a bundle of tone, responsiveness, and content differences. These issues are load-bearing for the paper's central causal claim.

major comments (5)
  1. [§5 versus §2.5] The conclusion states that the results are 'confirming our original hypothesis,' but §2.5 explicitly says the sample is too small for meaningful inferential tests, and no inferential statistics, effect sizes, or confidence intervals are reported anywhere. The quantitative basis in Table 1 and Figures 3–4 is a set of mean differences from N=12, which cannot support a confirmatory causal claim without accounting for within-subject variability, multiple comparisons, and order. Please either report appropriate inferential statistics (or at least effect sizes and confidence intervals) and reclassify the study as exploratory, or substantially temper the wording of the conclusion.
  2. [§2.2 and §2.4] The description of condition order is internally contradictory. §2.2 states that the first group interacted with the robot without personality first and the other half interacted with the personality condition first. §2.4 states that the first group started with the personality condition followed by the control condition, and only 'after the first six participants' did the second group begin in the control condition. These cannot both be true. Because order effects (practice, fatigue, and contrast) are a real threat in within-subject designs, the actual order must be stated precisely and, ideally, checked in the data for order-by-condition interactions.
  3. [§2.2–§2.4, manipulation confound] The two conditions differ along multiple dimensions at once: the personality condition includes jokes and witty follow-ups, informal tone, and possibly different LLM-generated content, while the control condition uses formal tone and no additional commentary. Consequently, the observed differences in PANAS positive affect and Godspeed ratings cannot be attributed specifically to 'personality' or 'humor'; they could reflect any component of this bundle, including response relevance, verbosity, or perceived attentiveness. Please either decompose the manipulation or reframe the research question as comparing two holistic interaction styles, and adjust the causal language in §5 accordingly.
  4. [Appendix C, Part 3] The manipulation check question 'Did you perceive Robot A as more task-oriented and Robot B as more social and engaging?' is leading because it states the expected distinction and is combined with preference questions. Under these conditions, the qualitative finding that all participants detected a difference is unsurprising and does not independently validate the manipulation. A more neutral manipulation check, such as coding free descriptions, or indirect behavioral measures would strengthen the claim that participants perceived the intended manipulation without being cued.
  5. [§4.1 and §3.2] The limitations section correctly acknowledges the small, homogeneous sample, the participants' prior familiarity with NAO robots, the short interaction duration, and occasional nonsensical LLM outputs. These limitations, combined with the post-hoc self-report nature of the qualitative data, further undermine the strength of the confirmatory statement in §5. The qualitative quotes are informative as illustrations, but they should not be used as primary evidence for a causal effect of personality on user experience.
minor comments (4)
  1. [§2.3, references] The text attributes the claim about reducing perceived waiting time to 'D. Kang et al.', but reference [18] is Pelikan and Hofstetter; please correct the citation or add the intended reference.
  2. [Throughout] There are several typos and formatting issues: 'Thereareexistingoccurrences' in the introduction, 'T able 1' and '0 .669' / '1 .138' in Table 1, 'likability' versus 'likeability', 'ideologues' in §4.1 (likely 'idiosyncrasies'), and 'non-nonsensical' in §4.1 (likely 'nonsensical'). These should be corrected.
  3. [Throughout] The robot name is written inconsistently as both 'NAO' and 'Nao'; please standardize to one form.
  4. [Figures 3 and 4] Figures 3 and 4 are referenced in the text, but the full text does not show the actual plots; please ensure the figures are included with axis labels and legends, and add explicit captions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the study is an empirical user study with externally standardized outcome measures, and none of its conclusions are derived by construction from its inputs.

full rationale

The paper makes no formal derivation, fits no model to the outcome data, and validates no prediction against a fitted quantity. Its outcome measures are the standardized PANAS and Godspeed questionnaires (References 15 and 9), which are external instruments and are filled out by participants after each interaction; the reported means in Table 1 and Figures 3-4 are descriptive statistics of those independent measurements. The only calibrated parameters mentioned are the audio-level threshold and pause duration used for speech segmentation (Section 2.3), which are unrelated to the PANAS or Godspeed scores and do not enter the reported results. There is no cited 'uniqueness theorem', no prior-work citation that is load-bearing for the empirical claim, and no renamed known result: the paper compares two concrete robot behaviors and reports participant ratings. The conclusion in Section 5 that users felt more comfortable and engaged in the personality condition is an empirical generalization from the descriptive data and qualitative responses; whether that inference is statistically warranted is a methodological validity concern (no inferential statistics, inconsistent counterbalancing description), not a circularity concern. The hypothesis is not defined in terms of the outcome measure, and the outcome measure is not constructed from the hypothesis. Thus the derivation chain is not circular: no claim reduces by construction to its own inputs.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

No fitted model or invented constructs; the analysis is descriptive. The only hand-calibrated parameters are audio thresholds, and the outcome comparison depends on the validity of self-report scales and the assumption of a clean personality manipulation.

free parameters (2)
  • Voice activity detection amplitude threshold = 350 (unitless amplitude level)
    Calibrated by the authors to detect the end of user speech; affects turn-taking timing but not the outcome scores.
  • Silence duration threshold = 1.5 seconds
    Calibrated minimum silence before recording stops; same role as above.
assumptions (3)
  • domain assumption PANAS and Godspeed are valid measures of emotional state and robot perception in this setting
    The study uses these standardized instruments without revalidation; standard practice, but scoring subscales are modified (e.g., Godspeed items for animacy and safety are omitted).
  • domain assumption The two conditions differ only in personality and humor, not in other interaction qualities
    The personality condition used a different system prompt with jokes and informal tone; the control used formal prompts. Differences in responsiveness or topical relevance could also drive the results.
  • domain assumption Participants' self-reports reflect the robot's behavior rather than order or social desirability
    Within-subject ordering was meant to control this, but the paper contradicts itself on the order, and participants knew the study was about social robots.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Existential Crisis: A Social Robot's Reason for Being." pith.science (2026). https://pith.science/paper/BK33V2UA

@misc{pith2026250103376,
  author       = {Pith},
  title        = {Pith review of: Existential Crisis: A Social Robot's Reason for Being},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BK33V2UA}},
  note         = {Machine review of arXiv:2501.03376}
}
read the original abstract

As Robots become ever more important in our daily lives there's growing need for understanding how they're perceived by people. This study aims to investigate how the user perception of robots is influenced by displays of personality. Using LLMs and speech to text technology, we designed a within-subject study to compare two conditions: a personality-driven robot and a purely task-oriented, personality-neutral robot. Twelve participants, recruited from Socially Intelligent Robotics course at Vrije Universiteit Amsterdam, interacted with a robot Nao tasked with asking them a set of medical questions under both conditions. After completing both interactions, the participants completed a user experience questionnaire measuring their emotional states and robot perception using standardized questionnaires from the SRI and Psychology literature.

Figures

Figures reproduced from arXiv: 2501.03376 by the authors.

Figure 1
Figure 1. Conversational flow with NAO robot system prompt at the start of the interaction, or the user’s feedback during the interaction. The text-to-speech module of the SIC framework is used to make NAO speak. OpenAI’s whisper [17] is responsible for recording the user input and per￾forming speech-to-text conversion, the output of which is then fed back to the Llama model. The interaction between the user and NAO can be su… view at source ↗
Figure 2
Figure 2. Amplitude of a sample response response. The yellow areas indicate occurrences during the response where the amplitude of the response was: – below a calibrated threshold (In this case 350Hz) – For longer than a calibrated timespan (in this case 1.5 seconds) These parameters could be calibrated at any time in order to enhance a more natural flow of the conversation. If the above requirements were both satisfied, NAO… view at source ↗
Figure 3
Figure 3. Graphical Representation of PANAS Scores [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

45 extracted references · 38 canonical work pages

  1. [1]

    OpenAI. (2023). GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 . https://doi.org/10.48550/arXiv.2303.08774

  2. [2]

    Dubey, Abhimanyu & Jauhri, Abhinav & Pandey, Abhinav & Kadian, Abhishek & Al-Dahle, Ahmad & Letman, Aiesha & Mathur, Akhil & Schelten, Alan & Yang, Amy & Fan, Angela & Goyal, Anirudh & Hartshorn, Anthony & Yang, Aobo & Mitra, Archi & Sravankumar, Archie & Korenev, Artem & Hinsvark, Arthur & Rao, Arun & Zhang, Aston & Zhao, Zhiwei. (2024). The Llama 3 Herd...

  3. [3]

    leng2024longcontextragperformance, title=Long Context RAG Performance of Large Language Models, author=Quinn Leng and Jacob Portes and Sam Havens and Matei Zaharia and Michael Carbin, year=2024, eprint=2411.03538, archivePrefix=arXiv, primaryClass=cs.LG, https://arxiv.org/abs/2411.03538

  4. [4]

    Aldebaran. (2014). Technical site. https://corporate-internal-prod. aldebaran.com/en/pepper

  5. [5]

    Mileounis, A., Cuijpers, R.H., Barakova, E.I. (2015). Creating Robots with Per- sonality: The Effect of Personality on Social Intelligence. In: Ferrández Vicente, J., Álvarez-Sánchez, J., de la Paz López, F., Toledo-Moreo, F., Adeli, H. (eds) Artificial Computation in Biology and Medicine. IWINAC 2015. Lecture Notes in Computer Science(), vol 9107. Spring...

  6. [6]

    (2006).The Emotion Machine: Commonsense Thinking, Artificial In- telligence, and the Future of the Human Mind

    Minsky, M. (2006).The Emotion Machine: Commonsense Thinking, Artificial In- telligence, and the Future of the Human Mind. Simon and Schuster

  7. [7]

    Picard, Rosalind W. (1997). Affective Computing. MIT Press, Cambridge, MA. ISBN: 978-0-262-16170-2

  8. [8]

    Mehrabian, A., & Ferris, S. R. (1967). Inference of attitudes from nonverbal com- munication in two channels. Journal of Consulting Psychology, 31(3), 248–252. https://doi.org/10.1037/h0024648

Show all 45 references
  1. [9]

    Bartneck, C., Kulić, D., Croft, E., & Zoghbi, S. (2009). Measurement instru- ments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots.International Journal of Social Robotics, 1(1), 71–81. https://doi.org/10.1007/s12369-008-0001-3

  2. [10]

    M., Wyman, A

    Carpinella, C. M., Wyman, A. B., Perez, M. A., & Stroessner, S. J. (2017). The robotic social attributes scale (RoSAS): Development and validation. In Pro- ceedings of the 2017 ACM/IEEE International Conference on Human-Robot In- teraction (HRI ’17) (pp. 254–262). Vienna, Aust...

  3. [11]

    Heerink, M., Kröse, B., Evers, V., & Wielinga, B. (2010). Assessing acceptance of assistive social agent technology by older adults: The Almere Model. Inter- national Journal of Social Robotics, 2(4), 361–375. https://doi.org/10.1007/ s12369-010-0068-5

  4. [12]

    S., & Scassellati, B

    Admoni, H., Dragan, A., Srinivasa, S. S., & Scassellati, B. (2014). Deliberate delays during robot-to-human handovers improve compliance with gaze communication. In 2014 9th ACM/IEEE International Conference on Human-Robot Interaction (HRI) (pp. 49–56). Bielefeld, Germany.http...

  5. [13]

    Thomas Orden, Mike EU Ligthart, Koen Hindriks, Thomas Wiggers, and Karen Chiang. [n. d.]. The Social Interaction Cloud (SIC) https://socialrobotics. atlassian.net/wiki/spaces/CBSR/overview

  6. [14]

    Picard, R. W. (1997).Affective Computing. MIT Press. Existential Crisis: A Social Robot’s Reason for Being 13

  7. [15]

    A., & Tellegen, A

    Watson, D., Clark, L. A., & Tellegen, A. (1988). Development and validation of brief measures of positive and negative affect: The PANAS scales.Journal of Personality and Social Psychology, 54(6), 1063–1070.https://doi.org/10.1037/ 0022-3514.54.6.1063

  8. [16]

    Broadbent, E., Stafford, R., & MacDonald, B. (2013). Acceptance of healthcare robotsfortheolderpopulation:Reviewandfuturedirections. International Journal of Social Robotics, 5(3), 319–330.https://doi.org/10.1007/s12369-013-0196-4

  9. [17]

    Radford, Alec & Kim, Jong & Xu, Tao & Brockman, Greg & McLeavey, Chris- tine & Sutskever, Ilya. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. 10.48550/arXiv.2212.04356

  10. [18]

    Hannah Pelikan and Emily Hofstetter. 2023. Managing Delays in Human-Robot Interaction. ACM Trans. Comput.-Hum. Interact. 30, 4, Article 50 (August 2023), 42 pages. https://doi.org/10.1145/3569890

  11. [19]

    H., & Torreira, F

    Kendrick, K. H., & Torreira, F. (2015). The timing and construction of preference: A quantitative study. Discourse Processes, 52(4), 255-289

  12. [20]

    John C. Meyer, Humor as a Double-Edged Sword: Four Functions of Humor in Communication, Communication Theory, Volume 10, Issue3, 1 August 2000,Pages 310–331 https://doi.org/10.1111/j.1468-2885.2000.tb00194.x 14 Medgyesy et al. A Appendix: Participant instruction Title of the S...

  13. [21]

    I understand that my participation is voluntary and that I can withdraw at any time without penalty

  14. [22]

    I understand that my responses will be anonymous and that no personal data will be collected that could identify me

  15. [23]

    I understand that the data collected will be used solely for research purposes and may be published in academic journals

  16. [24]

    Consent By signing below, you indicate that you have read and understood the informa- tion provided and agree to participate in this study

    I have been informed about the nature of the study and the procedures involved. Consent By signing below, you indicate that you have read and understood the informa- tion provided and agree to participate in this study. Participant Name: Participant Signature: Date: Thank you ...

  17. [25]

    In your own words, describe your overall experience with the robot

  18. [26]

    What stood out to you the most during your interaction with the robot?

  19. [27]

    How did the robot’s behavior affect your engagement during the interaction?

  20. [28]

    Part 2: Numerical Questions – User Experience and Perception A

    Did you feel the robot understood you? Please explain. Part 2: Numerical Questions – User Experience and Perception A. How Do You Feel Instructions: Indicate to what extent you feel each emotion right now, after interacting with the robot. Use a scale from 1 (Very slightly or ...

  21. [29]

    Interested ............................................. (1–5)

  22. [30]

    Excited ............................................... (1–5)

  23. [31]

    Upset ................................................. (1–5)

  24. [32]

    Strong ................................................ (1–5)

  25. [33]

    Enthusiastic ........................................... (1–5)

  26. [34]

    Distressed ............................................. (1–5)

  27. [35]

    Determined ........................................... (1–5)

  28. [36]

    Nervous ............................................... (1–5)

  29. [37]

    Alert .................................................. (1–5)

  30. [38]

    Inspired ............................................... (1–5) B. How Do You View the Robot Instructions: For each pair of adjectives, select a number from 1 to 5 that best describes your perception of the robot. 16 Medgyesy et al

  31. [39]

    Anthropomorphism – Fake ..................1 — 2 — 3 — 4 — 5 ..................Natural – Machinelike .............1 — 2 — 3 — 4 — 5 .............Humanlike – Unconscious .............1 — 2 — 3 — 4 — 5 .............Conscious – Artificial ................1 — 2 — 3 — 4 — 5 ...........

  32. [40]

    1 — 2 — 3 — 4 — 5

    Likeability – Dislike .................. 1 — 2 — 3 — 4 — 5 .................. Like – Unfriendly ...............1 — 2 — 3 — 4 — 5 ...............Friendly – Unkind ..................1 — 2 — 3 — 4 — 5 ..................Kind – Unpleasant ..............1 — 2 — 3 — 4 — 5 ..............

  33. [41]

    1 — 2 — 3 — 4 — 5

    Intelligence – Incompetent ............1 — 2 — 3 — 4 — 5 ............Competent – Ignorant ............ 1 — 2 — 3 — 4 — 5 ............ Knowledgeable – Irresponsible ............1 — 2 — 3 — 4 — 5 ............Responsible – Unintelligent ............ 1 — 2 — 3 — 4 — 5 ...............

  34. [42]

    Did you perceive Robot A as more task-oriented and Robot B as more social and engaging? □ Yes □ No □ Unsure

  35. [43]

    Which robot did you prefer interacting with? □ Robot A: Task-oriented □ Robot B: Social and engaging

  36. [44]

    Please explain why you preferred this robot

  37. [45]

    question

    Would you recommend your preferred robot to others? (1–5) – Scale from 1 (Definitely not) to 5 (Definitely yes) Part 4: Additional Notes Please provide any additional comments or suggestions about your experiences with the robots. [Space for open-ended response] D LLM prompts ...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.