{"id":"de3f2bb8-c3f8-4630-92a6-a479d8b6f016","arxiv_id":"2502.01448","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Proactively telling users what a restaurant robot can do produced higher enjoyment, greater willingness to reuse the robot, and more conversational user speech than reactive disclosure or no disclosure.","lead":"This paper tests whether a restaurant robot should tell users its abilities up front, only after errors, or not at all. A study with 113 participants found that proactive announcements increased user enjoyment and the willingness to use the robot again, while also leading to longer, more conversational user utterances.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proactive condition's longer, more instructive opening is confounded with capability content; the conversational-style finding may be dialogue alignment rather than capability communication.","rationale":"The reader identified the same confound as the weakest assumption, and I agree it is the most load-bearing issue. The paper's headline claim is that capability communication makes conversations more natural; the behavioral measures supporting this are all potentially explained by the longer, more directive robot opening in the proactive condition. Unlike the statistical inconsistency in the post-hoc ANOVA (F = 1.48, p < .05 is impossible for these degrees of freedom) or the video-sample mismatch, the confound cannot be corrected by re-analysis of the existing data; it requires a new control condition. Even if the rating effects (enjoyment, willingness) survive reanalysis, the central behavioral contribution is uninterpretable without such a control. Therefore the appropriate verdict remains CONDITIONAL: the study is well-designed enough to be a useful pilot, but the main causal claim about conversational style needs a stronger design. The concrete follow-up experiment proposed above would settle whether the effect is content-specific.","tokens_in":13095,"tokens_out":6815,"duration_ms":61803,"concrete_test":"Run a follow-up experiment with a fourth 'verbose baseline' condition in which the robot's opening utterance matches the proactive condition in word count, syntactic complexity, and politeness but contains no capability information (e.g., a general welcome, description of the restaurant, and small talk). Compare user word counts, one-word utterance rates, and dialogue-act distributions across this condition, the original baseline, and the proactive condition. If the verbose baseline produces user behavior similar to the proactive condition and significantly different from the original baseline, the behavioral effect is attributable to utterance length or alignment rather than capability content. If it matches the original baseline, the capability message is the active ingredient.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proactive condition differs from baseline not only in capability information but also in the length, structure, and instructiveness of the robot's first utterance. The proactive opening tells users the robot will seat them and take their order, warns about possible speech-recognition problems, asks them to repeat information, and advises 'I can help with small questions, but I work best if we stick to the conversation flow.' The baseline opens with a single greeting question. Users are known to align to a partner's verbosity and conversational style; longer, more detailed robot speech can itself elicit longer, more detailed user responses. The paper's exploratory findings—more words per utterance, more unique words, fewer one-word utterances, and more informing/information-request dialogue acts in the proactive condition (Sec. VI.A.b, Table III)—are exactly what one would expect from alignment with a longer opening. The discussion even attributes the first-interaction word-count increase to 'the robot starting the interaction with explanations,' which is the same confound. No condition controls for utterance length or prefatory speech without capability content, so the central claim that capability communication causes a more conversational interaction style is not identified; the effect could be produced by any lengthy, instructive monologue. This is a structural design issue, not a mere statistical artifact, and it directly affects the paper's main novelty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a between-subjects user study (N=120, 113 analyzed) in which participants interacted three times with a social robot in a restaurant setting. Three conditions are compared: a baseline with no capability communication, a reactive strategy that communicates capabilities only after detected problems, and a proactive strategy that communicates capabilities at the start of the interaction. The authors pre-registered hypotheses that the reactive and proactive robots would be preferred over baseline, and that the two capability strategies would differ, using enjoyment, ease of use, performance trust, and willingness to use again as primary measures, plus exploratory behavioral metrics (words per utterance, unique words, one-word utterance rate, dialogue acts). Results show a significant MANOVA for baseline vs. proactive, with post-hoc ANOVAs significant for enjoyment and willingness to use again, while baseline vs. reactive and reactive vs. proactive were not significant. Exploratory analyses show more words and unique words, fewer one-word utterances, and more informing/information-request dialogue acts in the proactive condition. The authors conclude that proactive capability communication is preferred and leads users to adopt a more conversational interaction style.","tokens_in":13330,"tokens_out":4828,"duration_ms":45261,"significance":"If the causal claims hold, this is a useful HRI contribution: it empirically compares two concrete, deployable strategies for capability communication in spoken interaction, uses a pre-registered design with a priori power analysis and standardized questionnaires, and combines subjective ratings with objective interaction logs and video-based checking. The distinction between proactive and reactive transparency is practically relevant, and the exploratory behavioral findings point to an interesting phenomenon. However, the paper's central causal claim about capability communication causing a 'more conversational interaction style' is currently undermined by an alternative explanation (dialogue alignment to longer, more instructive robot utterances) and by inconsistent statistical reporting. The strengths of the design—pre-registration, sample size planning, multiple measures—make these issues addressable, but they are load-bearing for the main conclusions.","major_comments":[{"comment":"The statistical reporting for the willingness-to-use-again ANOVA is internally inconsistent. The text reports F = 1.48, p < .05, but with the presented means (baseline M = 3.48, SD = 0.93; proactive M = 4.00, SD = 1.05) and per-condition N of roughly 38, an F value of 1.48 is far too small to be significant and is incompatible with the reported p < .05. This is not a minor typo, because the significance of willingness to use again is one of the two post-hoc results used to support H2. Please correct the F (and possibly the p) value, or clarify if a different test was conducted, and re-run the post-hoc analyses accordingly.","section":"Sec. VI.A.a (Table II)"},{"comment":"The proactive condition confounds capability communication with a longer, more instructive opening utterance. The dialogue example in Sec. III.A.c shows the proactive robot delivering a multi-sentence preamble about seating, ordering, speech-recognition limitations, and interaction style, while the baseline opens with a single greeting question. The behavioral differences attributed to capability communication—more words per utterance, more unique words, fewer one-word utterances, and more informing/information-request dialogue acts—are exactly the pattern expected from dialogue accommodation to a more verbose speaking partner. The paper itself, in Sec. VII.B, explains the first-interaction word-count increase as 'the robot starting the interaction with explanations,' which is the same confound. Without a control condition that matches utterance length and structural content but omits capability information, the central claim that capability communication causes a more conversational interaction style is not identified. This is a structural design issue that affects the abstract's and conclusion's causal wording; either add a control condition or substantially temper the causal claims to descriptive/correlational ones.","section":"Sec. III.A.c, Sec. VI.A.b, Sec. VII.B"},{"comment":"The expertise-moderation analysis is presented without an inferential test. The text claims that 'in the reactive condition, people with more experience give higher overall ratings' and that the proactive condition shows a 'steep drop' for extremely familiar users, and Sec. VII.A concludes that 'the strategies are suitable for different expertise levels.' However, the figure is based on combined questionnaire scores with no reported cell sizes, statistical test, or interaction term. The text itself notes that there are few data points at the experience rating of 5. A regression or ANOVA with an experience-by-condition interaction, along with cell sizes, is needed to support these claims. As written, these interpretations are speculative and should be flagged as such.","section":"Sec. VI.A.b, Fig. 3, Sec. VII.A"},{"comment":"The dialogue act tagging—which underpins the conversational-style findings—lacks a documented reliability check. The paper states that 'we tagged all utterances using GPT and manual post-processing,' but does not report how the manual post-processing was performed, whether disagreements with the GPT taggings were adjudicated, or whether any inter-rater reliability (e.g., Cohen's kappa) was computed. Since Table III is used to infer that proactive users 'informing' and 'requesting information' more than baseline, the validity of the tags is load-bearing for the exploratory behavioral claim. Please provide details on the tagging procedure and a reliability estimate.","section":"Sec. VI.A.b, Table III"}],"minor_comments":[{"comment":"The limitations section states 'While we looked at 33 randomly selected interactions,' but Sec. VI.B.b reports selecting 30 interactions (10 per condition). Please reconcile this discrepancy.","section":"Sec. VII.C"},{"comment":"Typo: 'Wile we tried to minimize' should be 'While we tried to minimize.'","section":"Sec. VII.C"},{"comment":"The text says the MANOVA compared conditions pairwise but does not state whether the MANOVA itself was corrected for multiple comparisons; the Benjamini-Hochberg correction is mentioned only for the post-hoc ANOVAs. Please clarify the correction procedure for the pairwise MANOVA.","section":"Sec. VI.A.a"},{"comment":"The description 'Each condition used one LLM prompt for seating and another prompt for ordering, leading to a total of 6 prompts' is clear, but it would be helpful to state explicitly that the same LLM (GPT-4) was used across all conditions and whether the prompts were identical in structure except for the capability-communication content, since this relates to the confound discussed above.","section":"Sec. V.C"},{"comment":"The OSF pre-registration link is given as a view-only URL; ensure the link is stable and accessible for reviewers and readers.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid empirical core—pre-registration, power analysis, standardized measures, and a realistic autonomous robot setup—and could become a solid HRI contribution once the statistical reporting is corrected and the causal claims are either supported by a proper control condition or appropriately tempered. The main risk is the confound between capability content and utterance length, which is structural but could be addressed by reframing the central claim or by future control studies. The authors should also be asked to provide a clear correction for the inconsistent F/p values, as these are easily checkable and currently undermine reader trust in the reported results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is a pre-registered, 120-person in-person study comparing three capability communication strategies for a restaurant robot waiter. The direct comparison of proactive versus reactive general capability communication is new relative to prior work on failure explanations, and the authors did the procedural basics right: power analysis, standardized questionnaires, autonomous robot, random video subsample. The subjective effect on enjoyment and willingness to use the robot again in the proactive versus baseline conditions looks real.\n\nThe main problem is the confound in the proactive manipulation. The proactive opening is not just capability information; it is longer, more instructive, and explicitly asks the user to repeat themselves and 'stick to the conversation flow.' The behavioral differences the paper emphasizes—more words per utterance, more unique words, fewer one-word utterances, more informing dialogue acts—are exactly what you'd expect from users mirroring a more talkative, directive interlocutor. The discussion even credits 'the robot starting the interaction with explanations,' which names the same confound. Without a control condition that adds matched-length, capability-free speech, the central claim that capability communication causes a more conversational style is not identified.\n\nThere are also smaller issues. The reported F = 1.48 with p < .05 for willingness is not a plausible combination and should be corrected. The video analysis is described as 30 interactions in one place and 33 in another. The conclusion says users 'preferred' the proactive robot while ease of use and trust were not significant, which overstates the breadth. The reactive condition also suffered from more speech recognition problems, making the reactive-versus-proactive comparison hard to interpret.\n\nThat said, the preregistration, the powered sample, and the honest limitations section make this a genuine empirical contribution. The enjoyment and willingness effects are worth taking seriously even if the behavioral mechanism is unresolved. The paper deserves a serious referee, but it needs major revision: a control condition or at minimum a reanalysis that adjusts for utterance length, corrected statistics, and a conclusion that matches the significant outcomes.","headline":"A carefully run pre-registered HRI study whose central behavioral claim is confounded with the robot's longer opening; the enjoyment effect may survive, but the 'more conversational style' finding needs a control condition.","tokens_in":13833,"tokens_out":2107,"would_cite":false,"duration_ms":20428,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A robot that announces its limits up front gets higher enjoyment ratings and wordier, more conversational replies than one that only apologizes after errors.","keywords":["human-robot interaction","capability communication","spoken dialogue","transparency","proactive communication","social robots","user study","conversation style"],"falsifier":"Add a fourth condition whose opening turn matches the proactive one in length, structure, and friendliness but contains no capability information, for example a preamble about the restaurant or the day's specials. If enjoyment ratings and words per utterance rise just as much in that condition, the claim that capability content drives the effect would be refuted.","tokens_in":12918,"feed_emoji":"🤖","tokens_out":12026,"duration_ms":91519,"temperature":0.7,"pith_summary":"The paper asks what a robot should tell people about its own limitations during a spoken interaction, and what that telling does to the conversation. In a restaurant scenario with 113 analyzed participants, it compares a robot that announces its capabilities before the task begins, a robot that explains them only when it detects a misunderstanding, and a robot that simply apologizes and moves on. The authors' central claim is that users prefer the proactive robot, rating it significantly higher on enjoyment and willingness to use again, and that they respond to the upfront information by adopting a more conversational style: longer utterances, a wider vocabulary, and more informing and information-requesting turns instead of one-word commands. If correct, this means a robot's communication about its own limits is a practical design lever on the quality of the dialogue, not just a transparency nicety.","feed_headline":"Robots that announce their limits get more natural talk","feed_subtitle":"A 113-person study: upfront capability notices raised enjoyment and replaced one-word commands with fuller talk.","key_machinery":"The central object is the proactive capability communication strategy itself: an opening turn in which the robot states its role, warns that it may have trouble understanding the user, and instructs the user to repeat information when that happens, so the capability information is delivered before any task exchange begins. The implementation runs on a dialogue pipeline that combines Google Dialogflow for speech recognition and intent classification, a confidence threshold of 0.45 for triggering repair sequences, GPT-4 prompted to generate the robot's responses, and ARI's text-to-speech. The argument about conversational style rests on the interaction-log machinery — counts of words per utterance, unique words, one-word utterance share, and dialogue acts tagged with a modified ISO DA schema (the function of each utterance, such as informing, requesting, or confirming) — which is how the paper turns user preference into observable behavior.","core_discovery":"The discovery the paper argues for is that proactive capability communication — the robot telling users, in its own speech, what it can and cannot do — changes both the user's evaluation and the user's actual speech. In a between-groups comparison, the robot that opened each interaction by stating what it would do, warning that it might mishear, and asking the user to repeat information in those cases scored significantly higher than the baseline on enjoyment and on willingness to use the robot again (MANOVA, p < .05, with post-hoc ANOVAs significant on those two scales), while the reactive strategy, which explained what the robot understood only when a problem was detected, did not separate from baseline. The behavioral evidence is used to say what 'prefer' means in practice: participants in the proactive condition used more words per utterance across all three interactions, more unique words overall, fewer one-word utterances, and a dialogue-act profile with more informing and information-requesting statements and fewer direct food requests. The authors interpret this as users shifting from command-like speech to a conversational interaction style once they know what the robot can handle.","pith_inferences":["Part of the 'conversational style' effect may be dialogue accommodation that the study's design does not isolate: a longer, more talkative opening turn can invite longer replies even when its content carries no capability information.","A direct test the authors did not run would hold the opening turn's length and warmth fixed while varying only whether it contains capability statements, separating content from verbosity.","Because the reactive condition delivers capability information bundled with a just-experienced failure, its null result could reflect failure valence rather than information timing; matching failure moments across conditions would separate the two.","The wordiness shift would probably reproduce in a text-only chat version of the same scripts, which would show the effect is about expectation-setting in dialogue rather than the robot's physical presence — a claim the paper does not make."],"forward_implications":["Dialogue designers get a concrete, low-cost lever: stating capability information in the opening turn of a spoken interaction can raise enjoyment and repeat-use intention without changing the task itself.","The behavioral shift — more words, more unique words, fewer one-word commands, more informing turns — implies users generalize capability statements into an interaction style, not just a set of instructions.","Repair-only capability communication looks insufficient for the measures studied, in part because the robot cannot reactively explain failures it never detects, such as speech it did not hear at all.","Because the proactive advantage already appears in the first interaction and persists across all three sessions, the strategy may matter most exactly when a user meets an unfamiliar robot for the first time.","The expertise-dependent rating pattern implies that one fixed strategy will not suit all users, and the authors themselves conclude that capability communication should be adapted to the user."],"supporting_citations":[{"why":"Shows that explanation completeness helps users form more accurate mental models, the theoretical basis for why capability information should help.","marker":"[17]"},{"why":"Reports that the form of (un)solicited corrective communication affects user trust after mistakes, motivating the comparison of communication strategies.","marker":"[18]"},{"why":"Argues that transparency increases interaction efficiency, the premise behind hypotheses H1 and H2.","marker":"[19]"},{"why":"Finds that users prefer proactive over reactive failure explanations and rate the proactive system more trustworthy, the basis for hypothesis H3.","marker":"[26]"},{"why":"Supplies G*Power, used for the a priori power analysis that set the minimum sample size of 111.","marker":"[29]"},{"why":"Provides the enjoyment and perceived ease-of-use questionnaires used as dependent measures.","marker":"[30]"},{"why":"Provides the multidimensional trust measure used for the performance trust ratings.","marker":"[31]"},{"why":"Defines the ISO dialogue-act annotation schema, augmented by the authors, used to tag user utterances for the behavioral comparison.","marker":"[32]"}],"fun_headline_variants":["Robot states limits, users speak more naturally","Upfront robot limits make conversations more human","Proactive robot disclosure yields conversational speech","Robot that sets expectations gets fuller dialogue","When robots explain capabilities, talk turns natural"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the proactive robot's higher ratings and wordier replies come from the capability information itself, rather than from the fact that its opening speech turn is longer and more elaborate — the study has no control that adds extra conversational speech without capability content.","fun_headline_variants_meta":{"raw":{"variants":["Robot states limits, users speak more naturally","Upfront robot limits make conversations more human","Proactive robot disclosure yields conversational speech","Robot that sets expectations gets fuller dialogue","When robots explain capabilities, talk turns natural"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000286,"raw_usage":{"total_tokens":1656,"prompt_tokens":893,"completion_tokens":763,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":699}},"tokens_in":509,"tokens_out":763,"duration_ms":8689,"temperature":1.0,"reasoning_tokens":699,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:15:10.339854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Add a fourth condition whose opening turn matches the proactive one in length, structure, and friendliness but contains no capability information, for example a preamble about the restaurant or the day's specials. If enjoyment ratings and words per utterance rise just as much in that condition, the claim that capability content drives the effect would be refuted.","supporting_citations":[{"cited_title":"Too much, too little, or just right? ways explanations impact end users’ mental models,","cited_arxiv_id":null,"evidence_quote":"Shows that explanation completeness helps users form more accurate mental models, the theoretical basis for why capability information should help."},{"cited_title":"Human trust after robot mistakes: Study of the effects of different forms of robot communication,","cited_arxiv_id":null,"evidence_quote":"Reports that the form of (un)solicited corrective communication affects user trust after mistakes, motivating the comparison of communication strategies."},{"cited_title":"Being transparent about transparency: A model for human- robot interaction,","cited_arxiv_id":null,"evidence_quote":"Argues that transparency increases interaction efficiency, the premise behind hypotheses H1 and H2."},{"cited_title":"Reactive or proactive? how robots should explain failures,","cited_arxiv_id":null,"evidence_quote":"Finds that users prefer proactive over reactive failure explanations and rate the proactive system more trustworthy, the basis for hypothesis H3."},{"cited_title":"G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences,","cited_arxiv_id":null,"evidence_quote":"Supplies G*Power, used for the a priori power analysis that set the minimum sample size of 111."},{"cited_title":"Assessing acceptance of assistive social agent technology by older adults: the almere model,","cited_arxiv_id":null,"evidence_quote":"Provides the enjoyment and perceived ease-of-use questionnaires used as dependent measures."},{"cited_title":"Towards an iso standard for dialogue act annotation,","cited_arxiv_id":null,"evidence_quote":"Defines the ISO dialogue-act annotation schema, augmented by the authors, used to tag user utterances for the behavioral comparison."}],"review_version":1}