REVIEW 4 major objections 5 minor 32 references
What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A robot that announces its limits up front gets higher enjoyment ratings and wordier, more conversational replies than one that only apologizes after errors.
desk verdict A carefully run pre-registered HRI study whose central behavioral claim is confounded with the robot's longer opening; the enjoyment effect may survive, but the 'more conversational style' finding needs a control condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the proactive capability communication strategy itself: an opening turn in which the robot states its role, warns that it may have trouble understanding the user, and instructs the user to repeat information when that happens, so the capability information is delivered before any task exchange begins. The implementation runs on a dialogue pipeline that combines Google Dialogflow for speech recognition and intent classification, a confidence threshold of 0.45 for triggering repair sequences, GPT-4 prompted to generate the robot's responses, and ARI's text-to-speech. The argument about conversational style rests on the interaction-log machinery — counts of words per utterance, unique words, one-word utterance share, and dialogue acts tagged with a modified ISO DA schema (the function of each utterance, such as informing, requesting, or confirming) — which is how the paper turns user preference into observable behavior.
What would settle it
Add a fourth condition whose opening turn matches the proactive one in length, structure, and friendliness but contains no capability information, for example a preamble about the restaurant or the day's specials. If enjoyment ratings and words per utterance rise just as much in that condition, the claim that capability content drives the effect would be refuted.
Extended reading notes
Core claim
The discovery the paper argues for is that proactive capability communication — the robot telling users, in its own speech, what it can and cannot do — changes both the user's evaluation and the user's actual speech. In a between-groups comparison, the robot that opened each interaction by stating what it would do, warning that it might mishear, and asking the user to repeat information in those cases scored significantly higher than the baseline on enjoyment and on willingness to use the robot again (MANOVA, p < .05, with post-hoc ANOVAs significant on those two scales), while the reactive strategy, which explained what the robot understood only when a problem was detected, did not separate from baseline. The behavioral evidence is used to say what 'prefer' means in practice: participants in the proactive condition used more words per utterance across all three interactions, more unique words overall, fewer one-word utterances, and a dialogue-act profile with more informing and information-requesting statements and fewer direct food requests. The authors interpret this as users shifting from command-like speech to a conversational interaction style once they know what the robot can handle.
Load-bearing premise
The load-bearing premise is that the proactive robot's higher ratings and wordier replies come from the capability information itself, rather than from the fact that its opening speech turn is longer and more elaborate — the study has no control that adds extra conversational speech without capability content.
Editorial extensions
If this is right
- Dialogue designers get a concrete, low-cost lever: stating capability information in the opening turn of a spoken interaction can raise enjoyment and repeat-use intention without changing the task itself.
- The behavioral shift — more words, more unique words, fewer one-word commands, more informing turns — implies users generalize capability statements into an interaction style, not just a set of instructions.
- Repair-only capability communication looks insufficient for the measures studied, in part because the robot cannot reactively explain failures it never detects, such as speech it did not hear at all.
- Because the proactive advantage already appears in the first interaction and persists across all three sessions, the strategy may matter most exactly when a user meets an unfamiliar robot for the first time.
- The expertise-dependent rating pattern implies that one fixed strategy will not suit all users, and the authors themselves conclude that capability communication should be adapted to the user.
Reading between the lines
- Part of the 'conversational style' effect may be dialogue accommodation that the study's design does not isolate: a longer, more talkative opening turn can invite longer replies even when its content carries no capability information.
- A direct test the authors did not run would hold the opening turn's length and warmth fixed while varying only whether it contains capability statements, separating content from verbosity.
- Because the reactive condition delivers capability information bundled with a just-experienced failure, its null result could reflect failure valence rather than information timing; matching failure moments across conditions would separate the two.
- The wordiness shift would probably reproduce in a text-only chat version of the same scripts, which would show the effect is about expectation-setting in dialogue rather than the robot's physical presence — a claim the paper does not make.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a between-subjects user study (N=120, 113 analyzed) in which participants interacted three times with a social robot in a restaurant setting. Three conditions are compared: a baseline with no capability communication, a reactive strategy that communicates capabilities only after detected problems, and a proactive strategy that communicates capabilities at the start of the interaction. The authors pre-registered hypotheses that the reactive and proactive robots would be preferred over baseline, and that the two capability strategies would differ, using enjoyment, ease of use, performance trust, and willingness to use again as primary measures, plus exploratory behavioral metrics (words per utterance, unique words, one-word utterance rate, dialogue acts). Results show a significant MANOVA for baseline vs. proactive, with post-hoc ANOVAs significant for enjoyment and willingness to use again, while baseline vs. reactive and reactive vs. proactive were not significant. Exploratory analyses show more words and unique words, fewer one-word utterances, and more informing/information-request dialogue acts in the proactive condition. The authors conclude that proactive capability communication is preferred and leads users to adopt a more conversational interaction style.
Significance. If the causal claims hold, this is a useful HRI contribution: it empirically compares two concrete, deployable strategies for capability communication in spoken interaction, uses a pre-registered design with a priori power analysis and standardized questionnaires, and combines subjective ratings with objective interaction logs and video-based checking. The distinction between proactive and reactive transparency is practically relevant, and the exploratory behavioral findings point to an interesting phenomenon. However, the paper's central causal claim about capability communication causing a 'more conversational interaction style' is currently undermined by an alternative explanation (dialogue alignment to longer, more instructive robot utterances) and by inconsistent statistical reporting. The strengths of the design—pre-registration, sample size planning, multiple measures—make these issues addressable, but they are load-bearing for the main conclusions.
major comments (4)
- [Sec. VI.A.a (Table II)] The statistical reporting for the willingness-to-use-again ANOVA is internally inconsistent. The text reports F = 1.48, p < .05, but with the presented means (baseline M = 3.48, SD = 0.93; proactive M = 4.00, SD = 1.05) and per-condition N of roughly 38, an F value of 1.48 is far too small to be significant and is incompatible with the reported p < .05. This is not a minor typo, because the significance of willingness to use again is one of the two post-hoc results used to support H2. Please correct the F (and possibly the p) value, or clarify if a different test was conducted, and re-run the post-hoc analyses accordingly.
- [Sec. III.A.c, Sec. VI.A.b, Sec. VII.B] The proactive condition confounds capability communication with a longer, more instructive opening utterance. The dialogue example in Sec. III.A.c shows the proactive robot delivering a multi-sentence preamble about seating, ordering, speech-recognition limitations, and interaction style, while the baseline opens with a single greeting question. The behavioral differences attributed to capability communication—more words per utterance, more unique words, fewer one-word utterances, and more informing/information-request dialogue acts—are exactly the pattern expected from dialogue accommodation to a more verbose speaking partner. The paper itself, in Sec. VII.B, explains the first-interaction word-count increase as 'the robot starting the interaction with explanations,' which is the same confound. Without a control condition that matches utterance length and structural content but omits capability information, the central claim that capability communication causes a more conversational interaction style is not identified. This is a structural design issue that affects the abstract's and conclusion's causal wording; either add a control condition or substantially temper the causal claims to descriptive/correlational ones.
- [Sec. VI.A.b, Fig. 3, Sec. VII.A] The expertise-moderation analysis is presented without an inferential test. The text claims that 'in the reactive condition, people with more experience give higher overall ratings' and that the proactive condition shows a 'steep drop' for extremely familiar users, and Sec. VII.A concludes that 'the strategies are suitable for different expertise levels.' However, the figure is based on combined questionnaire scores with no reported cell sizes, statistical test, or interaction term. The text itself notes that there are few data points at the experience rating of 5. A regression or ANOVA with an experience-by-condition interaction, along with cell sizes, is needed to support these claims. As written, these interpretations are speculative and should be flagged as such.
- [Sec. VI.A.b, Table III] The dialogue act tagging—which underpins the conversational-style findings—lacks a documented reliability check. The paper states that 'we tagged all utterances using GPT and manual post-processing,' but does not report how the manual post-processing was performed, whether disagreements with the GPT taggings were adjudicated, or whether any inter-rater reliability (e.g., Cohen's kappa) was computed. Since Table III is used to infer that proactive users 'informing' and 'requesting information' more than baseline, the validity of the tags is load-bearing for the exploratory behavioral claim. Please provide details on the tagging procedure and a reliability estimate.
minor comments (5)
- [Sec. VII.C] The limitations section states 'While we looked at 33 randomly selected interactions,' but Sec. VI.B.b reports selecting 30 interactions (10 per condition). Please reconcile this discrepancy.
- [Sec. VII.C] Typo: 'Wile we tried to minimize' should be 'While we tried to minimize.'
- [Sec. VI.A.a] The text says the MANOVA compared conditions pairwise but does not state whether the MANOVA itself was corrected for multiple comparisons; the Benjamini-Hochberg correction is mentioned only for the post-hoc ANOVAs. Please clarify the correction procedure for the pairwise MANOVA.
- [Sec. V.C] The description 'Each condition used one LLM prompt for seating and another prompt for ordering, leading to a total of 6 prompts' is clear, but it would be helpful to state explicitly that the same LLM (GPT-4) was used across all conditions and whether the prompts were identical in structure except for the capability-communication content, since this relates to the confound discussed above.
- [General] The OSF pre-registration link is given as a view-only URL; ensure the link is stable and accessible for reviewers and readers.
Circularity Check
No significant circularity: empirical user study with independent measures and pre-registered hypotheses.
full rationale
This is an empirical user study, not a derivation chain. The independent variable (capability communication strategy) is operationalized as differing robot utterances, and the dependent variables are questionnaire ratings and logged user behavior, which are measured independently of the manipulation. No fitted parameter is renamed as a prediction: the only hand-set value (speech-recognition confidence threshold of 0.45) is a design parameter, not fitted to any outcome. The self-citation [4] (Reimann et al., survey on dialogue management) is used only as background to motivate why robots' capabilities vary; it is not load-bearing for the hypotheses, which are grounded in external literature and pre-registered (OSF). The skeptical concern that the proactive condition's longer opening is confounded with capability content is a plausible internal-validity threat, but it is a confound in an experimental comparison, not a circular derivation: the results do not reduce by construction to the inputs. The paper also reports non-significant comparisons (reactive vs. baseline), which is inconsistent with a forced or tautological conclusion. Therefore no specific circular step can be quoted with textual evidence, and the appropriate score is 0.
Assumptions & free parameters
free parameters (1)
- Speech recognition confidence threshold =
0.45
assumptions (4)
- domain assumption Self-report Likert scales are valid measures of enjoyment, ease of use, trust and willingness to use again.
- ad hoc to paper The proactive condition's longer initial robot utterance does not by itself cause users to produce longer utterances.
- domain assumption The randomly sampled subset of video interactions is representative of all conditions.
- domain assumption The conditions are comparable aside from the capability communication strategy.
Cite this review
Pith. "Pith review of What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations." pith.science (2026). https://pith.science/paper/75HW6K6O
@misc{pith2026250201448,
author = {Pith},
title = {Pith review of: What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations},
year = {2026},
howpublished = {\url{https://pith.science/paper/75HW6K6O}},
note = {Machine review of arXiv:2502.01448}
}
read the original abstract
When encountering a robot in the wild, it is not inherently clear to human users what the robot's capabilities are. When encountering misunderstandings or problems in spoken interaction, robots often just apologize and move on, without additional effort to make sure the user understands what happened. We set out to compare the effect of two speech based capability communication strategies (proactive, reactive) to a robot without such a strategy, in regard to the user's rating of and their behavior during the interaction. For this, we conducted an in-person user study with 120 participants who had three speech-based interactions with a social robot in a restaurant setting. Our results suggest that users preferred the robot communicating its capabilities proactively and adjusted their behavior in those interactions, using a more conversational interaction style while also enjoying the interaction more.
Figures
Reference graph
Works this paper leans on
-
[1]
Design methodology for the ux of hri: A field study of a commercial social robot at an airport,
M. Tonkin, J. Vitale, S. Herse, M.-A. Williams, W. Judge, and X. Wang, “Design methodology for the ux of hri: A field study of a commercial social robot at an airport,” in Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction , 2018, pp. 407– 415
work page 2018
-
[2]
Y . Iwamura, M. Shiomi, T. Kanda, H. Ishiguro, and N. Hagita, “Do elderly people prefer a conversational humanoid as a shopping assis- tant partner in supermarkets?” in Proceedings of the 6th international conference on Human-robot interaction , 2011, pp. 449–456
work page 2011
-
[3]
Exploring the educational potential of robotics in schools: A systematic review,
F. B. V . Benitti, “Exploring the educational potential of robotics in schools: A systematic review,” Computers & Education , vol. 58, no. 3, pp. 978–988, 2012
work page 2012
-
[4]
A survey on dialogue management in human-robot interaction,
M. M. Reimann, F. A. Kunneman, C. Oertel, and K. V . Hindriks, “A survey on dialogue management in human-robot interaction,” ACM Transactions on Human-Robot Interaction , 2024
work page 2024
-
[5]
Explainable agents and robots: Results from a systematic literature review,
S. Anjomshoae, A. Najjar, D. Calvaresi, and K. Fr ¨amling, “Explainable agents and robots: Results from a systematic literature review,” in 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019), Montreal, Canada, May 13–17, 2019 . International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 1078–1088
work page 2019
-
[6]
Spoken language interaction with robots: Recommendations for future research,
M. Marge, C. Espy-Wilson, N. G. Ward, A. Alwan, Y . Artzi, M. Bansal, G. Blankenship, J. Chai, H. Daum ´e III, D. Dey et al. , “Spoken language interaction with robots: Recommendations for future research,” Computer Speech & Language , vol. 71, p. 101255, 2022
work page 2022
-
[7]
Expectation management in child-robot interaction,
M. Ligthart, O. B. Henkemans, K. Hindriks, and M. A. Neerincx, “Expectation management in child-robot interaction,” in 2017 26th IEEE international symposium on robot and human interactive communication (RO-MAN). IEEE, 2017, pp. 916–921
work page 2017
-
[8]
Dialogue models for socially intelligent robots,
K. Jokinen, “Dialogue models for socially intelligent robots,” in Social Robotics: 10th International Conference, ICSR 2018, Qingdao, China, November 28-30, 2018, Proceedings 10 . Springer, 2018, pp. 127–138
work page 2018
Show all 32 references
-
[9]
Not all robots are evaluated equally: The impact of morphological features on robots’ as- sessment through capability attributions,
L. Kunold, N. Bock, and A. Rosenthal-von der P ¨utten, “Not all robots are evaluated equally: The impact of morphological features on robots’ as- sessment through capability attributions,” ACM Transactions on Human- Robot Interaction, vol. 12, no. 1, pp. 1–31, 2023
2023
-
[10]
Fresh start: Encouraging politeness in wakeword-driven human-robot interaction,
R. Wen, A. Hanson, Z. Han, and T. Williams, “Fresh start: Encouraging politeness in wakeword-driven human-robot interaction,” in Proceedings of the 2023 ACM/IEEE International Conference on Human-Robot Interaction, 2023, pp. 112–121
2023
-
[11]
From talking and listening robots to intelligent commu- nicative machines,
R. K. Moore, “From talking and listening robots to intelligent commu- nicative machines,” Robots that talk and listen , pp. 317–335, 2015
2015
-
[12]
Tulli, F
S. Tulli, F. Correia, S. Mascarenhas, S. Gomes, F. S. Melo, and A. Paiva, Effects of Agents’ Transparency on Teamwork , ser. Lecture Notes in Computer Science. Cham: Springer International Publishing, 2019, vol. 11763, p. 22–37. [Online]. Available: http: //link.springer.com/1...
2019 doi
-
[13]
A literature survey of how to convey transparency in co-located human–robot interaction,
S. Y . Sch ¨ott, R. M. Amin, and A. Butz, “A literature survey of how to convey transparency in co-located human–robot interaction,”Multimodal Technologies and Interaction, vol. 7, no. 3, p. 25, Feb. 2023
2023
-
[14]
Investigating transparency methods in a robot word- learning system and their effects on human teaching behaviors,
M. Hirschmanner, S. Gross, S. Zafari, B. Krenn, F. Neubarth, and M. Vincze, “Investigating transparency methods in a robot word- learning system and their effects on human teaching behaviors,” in 2021 30th IEEE International Conference on Robot & Human Interactive Communicatio...
2021
-
[15]
Transparency in hri: Trust and decision making in the face of robot errors,
B. Nesset, D. A. Robb, J. Lopes, and H. Hastie, “Transparency in hri: Trust and decision making in the face of robot errors,” in Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction. Boulder CO USA: ACM, Mar. 2021, p. 313–317. [Online]. Available:...
2021
-
[16]
When do people want an explanation from a robot?
L. Wachowiak, A. Fenn, H. Kamran, A. Coles, O. Celiktutan, and G. Canal, “When do people want an explanation from a robot?” in Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction . Boulder CO USA: ACM, Mar. 2024, p. 752–761. [Online]. Available...
2024 doi
-
[17]
Too much, too little, or just right? ways explanations impact end users’ mental models,
T. Kulesza, S. Stumpf, M. Burnett, S. Yang, I. Kwan, and W.-K. Wong, “Too much, too little, or just right? ways explanations impact end users’ mental models,” in 2013 IEEE Symposium on Visual Languages and Human Centric Computing . San Jose, CA, USA: IEEE, Sep. 2013, p. 3–10. ...
2013
-
[18]
Human trust after robot mistakes: Study of the effects of different forms of robot communication,
S. Ye, G. Neville, M. Schrum, M. Gombolay, S. Chernova, and A. Howard, “Human trust after robot mistakes: Study of the effects of different forms of robot communication,” in 2019 28th IEEE Inter- national Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, ...
2019
-
[19]
Being transparent about transparency: A model for human- robot interaction,
J. B. Lyons, “Being transparent about transparency: A model for human- robot interaction,” in 2013 AAAI spring symposium series , 2013
2013
-
[20]
Agent transparency: A review of current theory and evidence,
A. Bhaskara, M. Skinner, and S. Loft, “Agent transparency: A review of current theory and evidence,” IEEE Transactions on Human-Machine Systems, vol. 50, no. 3, p. 215–224, Jun. 2020
2020
-
[21]
When transparent does not mean explainable,
K. Fischer, “When transparent does not mean explainable,” in Explain- able Robotic Systems—Workshop in conjunction with HRI 2018 , 2018
2018
-
[22]
Measures for explainable ai: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-ai performance,
R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman, “Measures for explainable ai: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-ai performance,” Frontiers in Computer Science, vol. 5, p. 1096257, 2023
2023
-
[23]
Why did i fail? a causal-based method to find explanations for robot failures,
M. Diehl and K. Ramirez-Amaro, “Why did i fail? a causal-based method to find explanations for robot failures,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 8925–8932, 2022
2022
-
[24]
Explainable ai for robot failures: Generating explanations that improve user assistance in fault recovery,
D. Das, S. Banerjee, and S. Chernova, “Explainable ai for robot failures: Generating explanations that improve user assistance in fault recovery,” in Proceedings of the 2021 ACM/IEEE international conference on human-robot interaction, 2021, pp. 351–360
2021
-
[25]
User study exploring the role of explanation of failures by robots in human robot collaboration tasks,
P. Khanna, E. Yadollahi, M. Bj ¨orkman, I. Leite, and C. Smith, “User study exploring the role of explanation of failures by robots in human robot collaboration tasks,” arXiv preprint arXiv:2303.16010 , 2023
2023 arXiv
-
[26]
Reactive or proactive? how robots should explain failures,
G. LeMasurier, A. Gautam, Z. Han, J. W. Crandall, and H. A. Yanco, “Reactive or proactive? how robots should explain failures,” in Proceed- ings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, 2024, pp. 413–422
2024
-
[27]
Flailing, hailing, prevailing: Perceptions of multi-robot failure recovery strate- gies,
S. Reig, E. J. Carter, T. Fong, J. Forlizzi, and A. Steinfeld, “Flailing, hailing, prevailing: Perceptions of multi-robot failure recovery strate- gies,” in Proceedings of the 2021 ACM/IEEE International Conference on Human-Robot Interaction , 2021, pp. 158–167
2021
-
[28]
What is proactive human- robot interaction?-a review of a progressive field and its definitions,
M. K. van Den broek and T. B. Moeslund, “What is proactive human- robot interaction?-a review of a progressive field and its definitions,” ACM Transactions on Human-Robot Interaction , vol. 13, no. 4, pp. 1– 30, 2024
2024
-
[29]
G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences,
F. Faul, E. Erdfelder, A.-G. Lang, and A. Buchner, “G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences,” Behavior research methods , vol. 39, no. 2, pp. 175–191, 2007
2007
-
[30]
Assessing acceptance of assistive social agent technology by older adults: the almere model,
M. Heerink, B. Kr ¨ose, V . Evers, and B. Wielinga, “Assessing acceptance of assistive social agent technology by older adults: the almere model,” 2010
2010
-
[31]
A multidimensional conception and measure of human-robot trust,
B. F. Malle and D. Ullman, “A multidimensional conception and measure of human-robot trust,” in Trust in human-robot interaction . Elsevier, 2021, pp. 3–25
2021
-
[32]
Towards an iso standard for dialogue act annotation,
H. Bunt, J. Alexandersson, J. Carletta, J.-W. Choe, A. C. Fang, K. Lee, V . Petukhova, A. Popescu-Belis, L. Romary, C. Soria et al. , “Towards an iso standard for dialogue act annotation,” in Seventh conference on International Language Resources and Evaluation (LREC’10) , 201...
2010
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.