{"id":"ae88f61d-051b-4d7f-a5e9-c115267ea0b7","arxiv_id":"2501.03376","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 12-person pilot found higher positive affect and likeability scores for a joking NAO robot than a formal one, but only descriptive statistics support the claim.","lead":"This paper tested whether a NAO robot that tells jokes and uses an informal tone makes people enjoy a medical questionnaire more than a plain, task-only robot. Twelve students tried both versions, and most said the funny robot felt more engaging, but the study was too small to prove the effect.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5's causal 'confirming' claim is unsupported because the paper reports no inferential statistics and §§2.2/2.4 contradict each other on condition order, so order and demand effects remain live.","rationale":"The paper is a transparently reported pilot: system details, prompts, questionnaires, and limitations are included, and the authors correctly note the small sample. The central claim, however, is causal and confirmatory. The weakest link is the inference from descriptive means to 'confirming our original hypothesis.' The manuscript's own Section 2.5 says inferential tests were omitted because the sample is too small, yet Section 5 uses language of confirmation. The contradictory order descriptions in §2.2 and §2.4 mean we cannot tell whether order was actually counterbalanced; this is not a mere typo because the entire internal-validity defense rests on counterbalancing. A concrete recovery of the session order and an order-aware reanalysis would settle whether the effect is robust. The reader's weakest_assumption points to the same causal-attribution gap, including order, demand characteristics, and confounded manipulation; I agree. Since this is fixable with additional analysis and transparent reporting, the CONDITIONAL verdict stands; no verdict change is needed.","tokens_in":10877,"tokens_out":3843,"duration_ms":36358,"concrete_test":"Recover per-participant session logs (or Google Forms timestamps) to determine the true condition order for each of the 12 participants. Then run a paired Wilcoxon signed-rank test on PANAS positive affect and Godspeed likeability, stratifying or covarying by order. If the personality-condition advantage is confined to participants who experienced the personality condition second, or disappears in an order-adjusted analysis, the Section 5 causal claim fails; if it survives in both order groups, the concern is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the move from Figure 3/Table 1 to Section 5's statement that users 'did feel more comfortable... more engaged and open... thus confirming our original hypothesis.' For that inference, the mean differences in PANAS and Godspeed must be attributable to the personality manipulation. The manuscript gives no inferential statistics (Section 2.5 explicitly declines them) and the order counterbalancing is described inconsistently: Section 2.2 says the first group did the no-personality condition first, while Section 2.4 says the first group started with the personality condition. If the actual order was not balanced, the higher means in the personality condition could reflect practice, fatigue, or contrast effects rather than personality. Moreover, the within-subject questionnaire explicitly asks participants to identify which robot was more social/engaging (Appendix C, Part 3), creating demand characteristics; and the two conditions differed not only in jokes but in tone, responsiveness, and possibly the content of LLM replies, so the manipulation is a bundle. The direction of the effect is plausible and consistent with prior humor-in-HRI literature, but the conclusion as worded is not established by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subject pilot study comparing a personality-driven NAO robot (with humor and informal tone) against a task-oriented control robot in a medical questionnaire interaction. Twelve students interacted with both versions, and the authors measured emotional state with an abbreviated PANAS and robot perception with Godspeed subscales, supplemented by qualitative responses. The central claim is that the personality condition improved user experience, and the paper concludes that the original hypothesis is confirmed. The manuscript includes full experimental instructions, questionnaires, LLM prompts, and an ethical self-check, and it explicitly labels the study as a pilot with a small sample.","tokens_in":11013,"tokens_out":4042,"duration_ms":40433,"significance":"If the reported effect is real, the result is a modest but useful pilot contribution to humor and personality in human-robot interaction for healthcare questionnaires. The study's strengths are its use of standardized instruments (PANAS, Godspeed), a transparent appendix with questionnaires and prompts, explicit ethical consent, and a reproducible open-source LLM pipeline. The qualitative findings on perceived playfulness versus formality are plausible and consistent with prior HRI work. However, the confirmatory conclusion is not supported by the evidence presented: the study reports only descriptive statistics from twelve participants, the counterbalancing order is described inconsistently, and the manipulation is a bundle of tone, responsiveness, and content differences. These issues are load-bearing for the paper's central causal claim.","major_comments":[{"comment":"The conclusion states that the results are 'confirming our original hypothesis,' but §2.5 explicitly says the sample is too small for meaningful inferential tests, and no inferential statistics, effect sizes, or confidence intervals are reported anywhere. The quantitative basis in Table 1 and Figures 3–4 is a set of mean differences from N=12, which cannot support a confirmatory causal claim without accounting for within-subject variability, multiple comparisons, and order. Please either report appropriate inferential statistics (or at least effect sizes and confidence intervals) and reclassify the study as exploratory, or substantially temper the wording of the conclusion.","section":"§5 versus §2.5"},{"comment":"The description of condition order is internally contradictory. §2.2 states that the first group interacted with the robot without personality first and the other half interacted with the personality condition first. §2.4 states that the first group started with the personality condition followed by the control condition, and only 'after the first six participants' did the second group begin in the control condition. These cannot both be true. Because order effects (practice, fatigue, and contrast) are a real threat in within-subject designs, the actual order must be stated precisely and, ideally, checked in the data for order-by-condition interactions.","section":"§2.2 and §2.4"},{"comment":"The two conditions differ along multiple dimensions at once: the personality condition includes jokes and witty follow-ups, informal tone, and possibly different LLM-generated content, while the control condition uses formal tone and no additional commentary. Consequently, the observed differences in PANAS positive affect and Godspeed ratings cannot be attributed specifically to 'personality' or 'humor'; they could reflect any component of this bundle, including response relevance, verbosity, or perceived attentiveness. Please either decompose the manipulation or reframe the research question as comparing two holistic interaction styles, and adjust the causal language in §5 accordingly.","section":"§2.2–§2.4, manipulation confound"},{"comment":"The manipulation check question 'Did you perceive Robot A as more task-oriented and Robot B as more social and engaging?' is leading because it states the expected distinction and is combined with preference questions. Under these conditions, the qualitative finding that all participants detected a difference is unsurprising and does not independently validate the manipulation. A more neutral manipulation check, such as coding free descriptions, or indirect behavioral measures would strengthen the claim that participants perceived the intended manipulation without being cued.","section":"Appendix C, Part 3"},{"comment":"The limitations section correctly acknowledges the small, homogeneous sample, the participants' prior familiarity with NAO robots, the short interaction duration, and occasional nonsensical LLM outputs. These limitations, combined with the post-hoc self-report nature of the qualitative data, further undermine the strength of the confirmatory statement in §5. The qualitative quotes are informative as illustrations, but they should not be used as primary evidence for a causal effect of personality on user experience.","section":"§4.1 and §3.2"}],"minor_comments":[{"comment":"The text attributes the claim about reducing perceived waiting time to 'D. Kang et al.', but reference [18] is Pelikan and Hofstetter; please correct the citation or add the intended reference.","section":"§2.3, references"},{"comment":"There are several typos and formatting issues: 'Thereareexistingoccurrences' in the introduction, 'T able 1' and '0 .669' / '1 .138' in Table 1, 'likability' versus 'likeability', 'ideologues' in §4.1 (likely 'idiosyncrasies'), and 'non-nonsensical' in §4.1 (likely 'nonsensical'). These should be corrected.","section":"Throughout"},{"comment":"The robot name is written inconsistently as both 'NAO' and 'Nao'; please standardize to one form.","section":"Throughout"},{"comment":"Figures 3 and 4 are referenced in the text, but the full text does not show the actual plots; please ensure the figures are included with axis labels and legends, and add explicit captions.","section":"Figures 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"This manuscript has the character of a student course project report, and several presentation issues suggest it is not yet at the polish level of a peer-reviewed venue. The central scientific problem, however, is the gap between the reported descriptive pilot data and the confirmatory conclusion, along with the contradictory counterbalancing description. These are fixable: the authors could reclassify the study as exploratory, report effect sizes and confidence intervals, resolve the order ambiguity, and soften the causal language. Given that the authors acknowledged many limitations themselves, the paper may be suitable after a thorough revision if the venue publishes pilot studies."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this one up front: it is a twelve-participant course pilot, clearly labeled as such, but the Section 5 conclusion says the hypothesis is 'confirmed' while Section 2.5 explicitly says the sample is too small for inferential tests. That gap is the whole ballgame.\n\nThe system description is the genuinely useful part. The authors give the full Llama 3 + Whisper + SIC pipeline on NAO, the actual prompts (Appendix D), the health questionnaire (Appendix B), the full user-experience survey (Appendix C), and the audio-threshold calibration method. That level of transparency is better than many published HRI papers, and it makes the study reproducible in principle. The qualitative results are also fairly presented: some participants found the jokes awkward or inappropriate for a medical setting, and the paper acknowledges the cultural and context sensitivity of humor.\n\nThe soft spots are concentrated in the inference from data to claim. First, no inferential statistics anywhere; the means in Figures 3–4 and Table 1 are descriptive only. Second, the counterbalancing description contradicts itself: Section 2.2 says the first group did the no-personality condition first, while Section 2.4 says the first group started with the personality condition. That is not a minor typo; it directly affects whether the positive mean differences can be attributed to the manipulation or to order. Third, the manipulation is a bundle: the personality condition differs in jokes, tone, responsiveness, and possibly the content of LLM replies, so you cannot isolate 'personality.' Fourth, Appendix C Part 3 asks participants whether they perceived Robot A as task-oriented and Robot B as social/engaging, which invites demand effects.\n\nThat said, the authors are honest about the pilot status and list real limitations (sample diversity, short interaction, model misrecognitions). The direction of the effect is consistent with existing humor-in-HRI literature. I do not see fabrication or self-citation games; the references are standard. The paper just reaches for a stronger conclusion than the evidence supports.\n\nWho is this for? Someone building an LLM-driven social robot and wanting a concrete, transparent example of a within-subject pilot using standardized questionnaires. It would also be a useful case study for an HRI methods seminar about overclaiming from small samples.\n\nMy recommendation: do not send it to a top HRI journal as-is. But it deserves referee time for a short-paper or workshop track if the authors (1) fix the order description, (2) add basic paired tests or at least effect sizes and confidence intervals, and (3) rewrite the conclusion to match the evidence. If they do that, it becomes a legitimate, albeit minor, contribution.","headline":"Honest course-pilot with a solid systems appendix; the conclusion overreaches the descriptive data and the order counterbalancing is contradictory, but the paper would be a fair short-paper revision.","tokens_in":11640,"tokens_out":3096,"would_cite":false,"duration_ms":31068,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a humorous personality in a medical questionnaire robot raises users' positive affect and their perception of the robot's anthropomorphism, likeability, and intelligence, based on a 12-person within-subject pilot.","keywords":["social robots","human-robot interaction","robot personality","humor","user experience","large language models","medical questionnaire","within-subject design"],"falsifier":"Run a between-subjects version with at least 30 participants per condition, a strictly randomized order, and pre-registered inferential statistics, plus a control that delivers the same jokes in a monotone, task-only style. If the flat-joke robot produces the same high likeability and positive-affect scores as the personality robot, or if the personality advantage disappears when order is balanced, the paper's claim is refuted.","tokens_in":10611,"feed_emoji":"🤖","tokens_out":8090,"duration_ms":64157,"temperature":0.7,"pith_summary":"This paper reports a pilot study asking whether a robot that administers medical questions gives people a better experience when it shows a personality, specifically a sense of humor, than when it behaves as a neutral, task-only interviewer. Twelve participants each talked with the same humanoid robot twice, once in a 'personality' condition where it answered with witty, context-aware remarks and an informal tone, and once in a control condition where it read the same questions formally and added nothing. After each interaction, the participants filled in the PANAS mood questionnaire and the Godspeed robot-perception questionnaire. The paper claims that all users noticed the behavioral difference, felt more comfortable and more engaged in the personality condition, and gave it higher mean scores on positive affect, anthropomorphism, likeability, and perceived intelligence, supporting the hypothesis that personality improves user experience.","feed_headline":"A joke-cracking medical robot beats a formal one in user tests","feed_subtitle":"In a 12-person pilot, the witty version of the same robot was rated friendlier, more intelligent, and more engaging than the formal version.","key_machinery":"The central setup is a within-subject comparison of two dialogue policies for the same small humanoid robot, both generated by a locally deployed large language model from different system prompts. The personality prompt instructed a polite, pleasant tone with jokes tied to the user's answers; the control prompt instructed a formal, direct, monotone style with no additional commentary. The robot's spoken output is produced by a text-to-speech module, and the user's spoken answers are transcribed by a speech-to-text model, with the robot signaling that it is listening by changing its eye color. The outcome machinery is two standardized questionnaires, the PANAS mood scale and the Godspeed robot-perception scale, filled out after each condition.","core_discovery":"On its own terms, the paper's central discovery is that a visible, humorous personality in a medical questionnaire robot raises self-reported user experience in a short, single-session interaction. The quantitative basis is the gap in mean scores: personality-condition means exceed control means by roughly 1.3 to 1.75 points on the 1-to-5 PANAS positive-affect items, and by similarly visible margins on the Godspeed dimensions of anthropomorphism, likeability, and intelligence. Qualitative responses confirm the manipulation: participants described the personality robot as playful, conversational, funny, and enjoyable, and the control robot as static, boring, and form-like. The authors conclude that users felt more comfortable and more engaged and open during question answering in the personality condition, which they interpret as confirming the hypothesis that contextual humor creates a better experience.","pith_inferences":["Because the two conditions differed in multiple ways at once (tone, jokes, responsiveness to each answer, and perceived engagement), the study cannot isolate humor as the active ingredient; a condition that delivers the same jokes in a flat monotone would separate 'personality' from 'acknowledging the user'.","The reported mean differences of roughly 1.3 to 1.75 points on a 1-to-5 scale are large enough that a moderately powered replication with proper counterbalancing and inferential statistics could confirm or refute the pattern; the descriptive-only analysis leaves the conclusion open.","The paper's internal inconsistency about the order in which the two groups encountered the conditions (Section 2.2 says the first group got the neutral robot first; Section 2.4 says the first group started with the personality condition) means order effects could not have been properly controlled even within the within-subject design.","For healthcare robotics, a testable extension would be an adaptive robot that gauges a patient's reaction to early jokes and dials humor up or down accordingly, potentially capturing the engagement gain while avoiding the reported downside of jokes felt to be inappropriate for a medical setting."],"forward_implications":["A medical questionnaire robot that shows humor and an informal tone produces higher self-reported positive affect and higher perceived anthropomorphism, likeability, and intelligence than the same robot in a neutral, task-only mode.","Users are able to clearly detect the difference between the two interaction styles, and they report feeling more comfortable, relaxed, and open with the humorous robot.","The humorous robot changed user behavior: participants gave shorter, more concise answers to the formal robot and became more expansive with the humorous one.","Humor's effect is not uniformly positive: some participants felt the jokes were inappropriate for a medical setting, and one participant responded negatively to a culturally insensitive joke.","The effect of humor depends on context and user, indicating that robot personality should be tuned to the individual patient rather than applied one-size-fits-all."],"supporting_citations":[{"why":"Supplies the Godspeed questionnaire used to measure anthropomorphism, likeability, and intelligence.","marker":"[9]"},{"why":"Supplies the PANAS mood scale used to measure positive and negative affect after each interaction.","marker":"[15]"},{"why":"The locally run language model that generates the robot's conversational responses in both conditions.","marker":"[2]"},{"why":"The speech-to-text model that transcribes user answers to feed back into the language model.","marker":"[17]"},{"why":"The framework that connects the robot's text-to-speech and dialogue flow with the language model.","marker":"[13]"}],"fun_headline_variants":["A joke-telling medical robot wins over users","Personality makes robots more likable, study finds","Humor helps robot medical interviews score higher","Witty robot outranks formal one in user test","Robot charm boosts user experience in medical Q&A"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the higher mean scores in the personality condition are caused by the robot's personality (its jokes and informal tone) rather than by which condition came first, by participants expecting to like the more social robot, or by the fact that the humorous robot also acknowledged each user's answer, none of which were controlled for in this pilot.","fun_headline_variants_meta":{"raw":{"variants":["A joke-telling medical robot wins over users","Personality makes robots more likable, study finds","Humor helps robot medical interviews score higher","Witty robot outranks formal one in user test","Robot charm boosts user experience in medical Q&A"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000112,"raw_usage":{"total_tokens":1003,"prompt_tokens":829,"completion_tokens":174,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":101}},"tokens_in":445,"tokens_out":174,"duration_ms":2587,"temperature":1.0,"reasoning_tokens":101,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:52:10.107605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a between-subjects version with at least 30 participants per condition, a strictly randomized order, and pre-registered inferential statistics, plus a control that delivers the same jokes in a monotone, task-only style. If the flat-joke robot produces the same high likeability and positive-affect scores as the personality robot, or if the personality advantage disappears when order is balanced, the paper's claim is refuted.","supporting_citations":[{"cited_title":"A., & Tellegen, A","cited_arxiv_id":null,"evidence_quote":"Supplies the PANAS mood scale used to measure positive and negative affect after each interaction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The framework that connects the robot's text-to-speech and dialogue flow with the language model."}],"review_version":1}