Pith. sign in

REVIEW 4 major objections 5 minor 33 references

"Who Should I Believe?": User Interpretation and Decision-Making When a Family Healthcare Robot Contradicts Human Memory

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that when a home healthcare robot contradicts a user's memory, the robot's transparency changes who gets blamed for the contradiction but not whether the user follows the robot's advice, with most users overtrusting the…

desk verdict Two results: a solid overtrust finding, and a transparency-interpretation claim that currently rests on untested percentages. read the letter →

arxiv 2506.21322 v1 pith:YU655XAA submitted 2025-06-26 cs.HC cs.RO

classification cs.HCcs.RO
keywords human-robotinteractionovertrustrobottransparencysociabilityhealthcareroboticsattributionofblamemulti-useraccesscontrolmedicationadherence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what happens when a home healthcare robot gives advice that contradicts what the user remembers. In a 2×2 online experiment with 176 participants watching scripted videos, the authors find that the robot's transparency changes how people interpret the contradiction: with low transparency the most common explanation is that the user misremembered, whereas with high transparency the most common explanation is that another family member or outside party changed the robot's information. Neither transparency nor sociability changed the bottom-line decision, however: about 68% of participants chose the robot's suggested medication time over their own memory, and many did so even while suspecting a system malfunction or tampering. The authors argue this is evidence of overtrust that designers of home healthcare robots should take seriously.

What carries the argument

The central object is a scripted, first-person video interaction with a Furhat robot—a social robot platform with a projected face—acting as a family healthcare assistant. The discrepancy is fixed: the robot says it is 4 PM and time to take medication, while the user recalls 5 PM. Transparency is operationalized as whether the robot, when challenged, offers detailed explanations and system records (high) or simply restates the information (low); sociability is operationalized as a bundle of greeting, self-introduction, memory recall, empathy, and gestures. The argument is carried by comparing coded open-ended attributions (user-related, robot-related, modified-by-others, clock changes, etc.) across conditions, and by relating those attributions to a forced-choice decision item.

What would settle it

A replication that equates the robot's utterance length and informativeness between conditions while varying only whether the explanation is framed as a verifiable system record—and includes a manipulation check for perceived transparency—would test whether attribution differences are really driven by transparency. If the attribution gap disappears when verbosity is controlled, the paper's central interpretive claim would be seriously weakened.

Watch

Extended reading notes

Core claim

The paper's central claim is that, in a family healthcare context, a robot's transparency level shifts users' causal attribution of an information discrepancy without shifting their compliance. When the robot simply restates its 4 PM medication reminder without explanation, half the participants assume the user did not remember correctly; when the robot shows detailed explanations and system records, only about a fifth blame user memory, and a third instead suspect that someone else—a partner, child, or other household member—modified the robot's information. Across all four conditions, a binomial test shows participants chose the robot's time significantly more often than their own memory (68.1%), and the authors interpret this as overtrust, noting that even among participants who suspected robot malfunction or third-party interference, large minorities still said they would follow the robot's advice.

Load-bearing premise

The transparency and sociability conditions bundle several behaviors at once—extra explanations and system records for high transparency, and greeting, self-introduction, memory recall, empathy, and gestures for high sociability—with no manipulation check and no control for utterance length, so the observed differences in attribution could be caused by how much the robot said rather than by transparency as a defined construct.

Editorial extensions

If this is right

  • If users follow a robot's medication advice even when they suspect system faults, adding transparency alone is unlikely to calibrate trust in home healthcare robots.
  • Designers of multi-user home robots need access-control mechanisms that prevent one household member's changes from silently overriding another's health settings.
  • Users' tendency to attribute discrepancies to their own memory errors in low-transparency conditions could delay detection of actual robot faults or tampering.
  • The finding that a third of participants who suspected third-party interference still took the robot's advice suggests that security warnings alone may not change behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not test whether explanation style could be used to steer blame; a follow-up could vary only the framing of the robot's explanation and measure attribution shifts.
  • The study leaves open whether live interaction would change overtrust; with a physically present robot, users might verify records rather than passively accept a video scenario.
  • A further untested consequence is that high transparency might act as a low-cost security cue, since it made users suspect external modification; this could be tested by measuring users' checking behavior after a discrepancy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a 2×2 between-subjects online study (N=176) in which participants watched videos of a Furhat robot that contradicts a fictional user's memory about medication time. It examines how transparency and sociability affect decision-making, interpretation of the discrepancy, and perceived trust. The central findings are that participants tended to follow the robot's recommendation (supported by a binomial test, p < .001), while the transparency manipulation is claimed to shift users' attribution of the discrepancy from user-related causes to external causes on the basis of descriptive percentages only. The paper argues for overtrust in healthcare robots and for the importance of access-control mechanisms in multi-user home environments.

Significance. The overtrust finding is a clean, confirmatory result with appropriate inferential support and speaks to a timely and safety-relevant HRI problem. If the transparency–attribution link were statistically established, it would be a novel and practically valuable contribution to the design of transparent home healthcare robots and to discussions of access control. However, as reported, the paper's headline RQ2 claim is not supported by any statistical inference and is potentially confounded by bundled manipulation components. The study design and scenario are otherwise appropriate and the paper is clearly written, but the main novel claim needs substantial reinforcement before the findings can be considered reliable.

major comments (4)
  1. [Sec. V.B / Table VI and Sec. VI.A] The central RQ2 claim that transparency influenced interpretations is supported only by descriptive percentages in Table VI; no chi-square, Fisher exact test, logistic regression, or confidence intervals are reported. The discussion in Sec. VI.A nevertheless states that 'a significantly larger proportion' of high-transparency participants attributed the discrepancy to external factors, but this significance is never established anywhere in the paper. Please add a proper inferential test on the contingency table (e.g., a 2×2 Fisher exact test contrasting user-related vs. all other categories by transparency level, or a multinomial/ordinal regression), with effect sizes and exact p-values, or revise the abstract and discussion to avoid the causal and significance wording.
  2. [Sec. IV.C] The transparency manipulation bundles multiple components: detailed explanations, system records, and longer utterances, while the low-transparency robot simply restates the information. No manipulation check is reported, so any observed attribution difference could be driven by the quantity or length of information rather than by transparency as a defined construct. Please include a perceived-transparency manipulation check or an experimental design that controls for utterance length/information volume, and discuss this confound explicitly in the limitations.
  3. [Sec. V.B] The qualitative coding that produces the eight categories in Table V and the percentages in Table VI is not described in terms of coding procedure: the number of coders, training, and inter-coder reliability (e.g., Cohen's kappa) are not reported. Because RQ2 hinges on these categories, the reliability of the coding should be documented, or the results should be treated as exploratory rather than confirmatory.
  4. [Sec. V.C / Table VII] The overtrust discussion in RQ3 is based on descriptive follow rates (87.30% vs. 45.71% vs. 34.29%) without an inferential test; a chi-square or Fisher exact test across attribution categories would support the claim that interpretations are associated with decisions. Additionally, the binomial test in Sec. V.A excludes 10 'No'/'Other' responses (113 vs. 53 out of 176); please justify this exclusion or use the full denominator in the test.
minor comments (5)
  1. [Sec. IV.E] The trust questionnaire is referred to as MDMT in the text but as 'MDTM' in the measure description; please correct the abbreviation.
  2. [Sec. VI.A] The phrase 'trust peception' is a typo and should read 'trust perception'.
  3. [Sec. III] The phrasing of RQ2, 'how do ... affect participants assess and interpret the discrepancy', is ungrammatical; please revise to 'affect participants' assessment and interpretation'.
  4. [Sec. V.B] The sentence 'the most frequently responses indicated that the robot is modified by someone else' should be revised to 'the most frequent response indicated that the robot was modified by someone else'.
  5. [Sec. V.A] The binomial test result is reported only as 'p < .001'; please also report the exact statistic (e.g., n and observed k) to improve transparency and reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports direct empirical observations and does not derive its conclusions from its inputs.

full rationale

This paper is an empirical between-subjects study, not a derivation. It manipulates robot transparency and sociability, collects participant decisions, qualitative interpretations, and trust ratings, and reports statistical tests on those observations. There is no fitted parameter that is later renamed as a prediction, no equation that reduces to an assumption, and no load-bearing uniqueness theorem or ansatz imported from the authors' prior work. The two self-citations (e.g., [5] and [12]) are contextual references to prior work on domestic abuse and trustworthy interaction design; they do not carry the paper's empirical claims. The main potential concerns—that the RQ2 transparency effect is reported without an inferential test and that the transparency manipulation bundles multiple behaviors without a manipulation check—are validity and statistical-inference issues, not circularity. The conclusions are not defined by construction from the inputs; they are direct reports of participant responses. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters are fitted because this is an empirical study, not a model. The central claims rest on construct validity of the video manipulations, generalizability of hypothetical decisions, and reliability of qualitative coding.

assumptions (3)
  • domain assumption Participants' self-reported choices in a video-based scenario predict how they would decide when interacting with a real healthcare robot.
    The decision measure is a hypothetical multiple-choice answer after watching a Furhat video, not an actual medication action. The authors acknowledge in Section VII that real-world responses may differ.
  • domain assumption The low/high transparency and low/high sociability manipulations vary only those intended constructs.
    High transparency adds detailed explanations and system records; high sociability bundles greeting, self-introduction, memory recall, empathy, and gestures (Table II). No manipulation check is reported, so observed effects could be due to information quantity or verbosity rather than the named dimension.
  • domain assumption The eight interpretation categories derived from open-ended responses are mutually exclusive and reliable.
    No inter-rater reliability or coding procedure is reported for the qualitative analysis in Section V.B, so the category counts in Tables V and VI could be unstable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "Who Should I Believe?": User Interpretation and Decision-Making When a Family Healthcare Robot Contradicts Human Memory." pith.science (2026). https://pith.science/paper/YU655XAA

@misc{pith2026250621322,
  author       = {Pith},
  title        = {Pith review of: "Who Should I Believe?": User Interpretation and Decision-Making When a Family Healthcare Robot Contradicts Human Memory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YU655XAA}},
  note         = {Machine review of arXiv:2506.21322}
}
read the original abstract

Advancements in robotic capabilities for providing physical assistance, psychological support, and daily health management are making the deployment of intelligent healthcare robots in home environments increasingly feasible in the near future. However, challenges arise when the information provided by these robots contradicts users' memory, raising concerns about user trust and decision-making. This paper presents a study that examines how varying a robot's level of transparency and sociability influences user interpretation, decision-making and perceived trust when faced with conflicting information from a robot. In a 2 x 2 between-subjects online study, 176 participants watched videos of a Furhat robot acting as a family healthcare assistant and suggesting a fictional user to take medication at a different time from that remembered by the user. Results indicate that robot transparency influenced users' interpretation of information discrepancies: with a low transparency robot, the most frequent assumption was that the user had not correctly remembered the time, while with the high transparency robot, participants were more likely to attribute the discrepancy to external factors, such as a partner or another household member modifying the robot's information. Additionally, participants exhibited a tendency toward overtrust, often prioritizing the robot's recommendations over the user's memory, even when suspecting system malfunctions or third-party interference. These findings highlight the impact of transparency mechanisms in robotic systems, the complexity and importance associated with system access control for multi-user robots deployed in home environments, and the potential risks of users' over reliance on robots in sensitive domains such as healthcare.

Figures

Figures reproduced from arXiv: 2506.21322 by the authors.

Figure 1
Figure 1. The Furhat robot employed in the online study. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Different levels of robot transparency condition type. To ensure the validity of the test, only response categories with frequencies greater than five were included in the analysis, limiting it to decisions based on the robot’s suggestion and the user’s memory. A total of 166 responses were analyzed. The results indicated that the association between user’s decision and the level of transparency or sociability was n… view at source ↗
Figure 3
Figure 3. Decision-making under different conditions. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Decision-making and trust level under different conditions [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 32 canonical work pages

  1. [1]

    Robots in healthcare: a scoping review,

    A. A. Morgan, J. Abdi, M. A. Syed, G. E. Kohen, P. Barlow, and M. P. Vizcaychipi, “Robots in healthcare: a scoping review,” Current robotics reports, vol. 3, no. 4, pp. 271–280, 2022

  2. [2]

    Deploying a robotic positive psychology coach to improve college students’ psychological well-being,

    S. Jeong, L. Aymerich-Franch, K. Arias, S. Alghowinem, A. Lapedriza, R. Picard, H. W. Park, and C. Breazeal, “Deploying a robotic positive psychology coach to improve college students’ psychological well-being,” User Modeling and User-Adapted Interaction, vol. 33, no. 2, pp. 571–615, 2023

  3. [3]

    Older adults’ experiences and percep- tions of living with bomy, an assistive dailycare robot: a qualitative study,

    N. Gasteiger, H. S. Ahn, C. Fok, J. Lim, C. Lee, B. A. MacDonald, G. H. Kim, and E. Broadbent, “Older adults’ experiences and percep- tions of living with bomy, an assistive dailycare robot: a qualitative study,” Assistive Technology, vol. 34, no. 4, pp. 487–497, 2022

  4. [4]

    Strengers and J

    Y . Strengers and J. Kennedy, The smart wife: Why Siri, Alexa, and other smart home devices need a feminist reboot . Mit Press, 2020

  5. [5]

    Anticipating the use of robots in domestic abuse: A typology of robot facilitated abuse to support risk assessment and mitigation in human-robot interaction,

    K. Winkle and N. Mulvihill, “Anticipating the use of robots in domestic abuse: A typology of robot facilitated abuse to support risk assessment and mitigation in human-robot interaction,” in Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, 2024, pp. 781–790

  6. [6]

    Overtrust of robots in emergency evacuation scenarios,

    P. Robinette, W. Li, R. Allen, A. M. Howard, and A. R. Wagner, “Overtrust of robots in emergency evacuation scenarios,” in 2016 11th ACM/IEEE international conference on human-robot interaction (HRI). IEEE, 2016, pp. 101–108

  7. [7]

    Would you trust a (faulty) robot? effects of error, task type and personality on human-robot cooperation and trust,

    M. Salem, G. Lakatos, F. Amirabdollahian, and K. Dautenhahn, “Would you trust a (faulty) robot? effects of error, task type and personality on human-robot cooperation and trust,” in Proceedings of the tenth annual ACM/IEEE international conference on human-robot interaction, 2015, pp. 141–148

  8. [8]

    Transparency in hri: Trust and decision making in the face of robot errors,

    B. Nesset, D. A. Robb, J. Lopes, and H. Hastie, “Transparency in hri: Trust and decision making in the face of robot errors,” in Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, 2021, pp. 313–317

Show all 33 references
  1. [9]

    Selecting the right robot: Influence of user attitude, robot sociability and embodiment on user preferences,

    M. Ligthart and K. P. Truong, “Selecting the right robot: Influence of user attitude, robot sociability and embodiment on user preferences,” in 2015 24th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IEEE, 2015, pp. 682–687

  2. [10]

    A meta-analysis of factors affecting trust in human-robot interaction,

    P. A. Hancock, D. R. Billings, K. E. Schaefer, J. Y . Chen, E. J. De Visser, and R. Parasuraman, “A meta-analysis of factors affecting trust in human-robot interaction,” Human factors, vol. 53, no. 5, pp. 517–527, 2011

  3. [11]

    Social interaction moderates human-robot trust-reliance relationship and improves stress coping,

    M. Lohani, C. Stokes, M. McCoy, C. A. Bailey, and S. E. Rivers, “Social interaction moderates human-robot trust-reliance relationship and improves stress coping,” in 2016 11th ACM/IEEE international conference on human-robot interaction (HRI) . IEEE, 2016, pp. 471– 472

  4. [12]

    A case study in designing trustworthy interactions: implications for socially assistive robotics,

    M. Zhong, M. Fraile, G. Castellano, and K. Winkle, “A case study in designing trustworthy interactions: implications for socially assistive robotics,” Frontiers in Computer Science , vol. 5, p. 1152532, 2023

  5. [13]

    Robots and transparency: The multiple dimensions of transparency in the context of robot technologies,

    H. Felzmann, E. Fosch-Villaronga, C. Lutz, and A. Tamo-Larrieux, “Robots and transparency: The multiple dimensions of transparency in the context of robot technologies,” IEEE Robotics & Automation Magazine, vol. 26, no. 2, pp. 71–78, 2019

  6. [14]

    De- velopment of the sociability of non-anthropomorphic robot home companions,

    J. Saez-Pons, H. Lehmann, D. S. Syrdal, and K. Dautenhahn, “De- velopment of the sociability of non-anthropomorphic robot home companions,” in 4th International Conference on Development and Learning and on Epigenetic Robotics . IEEE, 2014, pp. 111–116

  7. [15]

    System transparency in shared autonomy: A mini review,

    V . Alonso and P. De La Puente, “System transparency in shared autonomy: A mini review,” Frontiers in neurorobotics, vol. 12, p. 83, 2018

  8. [16]

    Integrating transparency, trust, and acceptance: The intelligent systems technology acceptance model (is- tam),

    E. S. V orm and D. J. Combs, “Integrating transparency, trust, and acceptance: The intelligent systems technology acceptance model (is- tam),” International Journal of Human–Computer Interaction, vol. 38, no. 18-20, pp. 1828–1845, 2022

  9. [17]

    Who should i blame? effects of autonomy and transparency on attributions in human-robot interaction,

    T. Kim and P. Hinds, “Who should i blame? effects of autonomy and transparency on attributions in human-robot interaction,” in ROMAN 2006-The 15th IEEE international symposium on robot and human interactive communication. IEEE, 2006, pp. 80–85

  10. [18]

    A survey of expectations about the role of robots in robot-assisted therapy for children with asd: ethical acceptability, trust, sociability, appearance, and attachment,

    M. Coeckelbergh, C. Pop, R. Simut, A. Peca, S. Pintea, D. David, and B. Vanderborght, “A survey of expectations about the role of robots in robot-assisted therapy for children with asd: ethical acceptability, trust, sociability, appearance, and attachment,” Science and enginee...

  11. [19]

    Factors for personalization and localization to optimize human–robot interaction: A literature review,

    N. Gasteiger, M. Hellou, and H. S. Ahn, “Factors for personalization and localization to optimize human–robot interaction: A literature review,” International Journal of Social Robotics , vol. 15, no. 4, pp. 689–701, 2023

  12. [20]

    The influence of empathy in human–robot relations,

    I. Leite, A. Pereira, S. Mascarenhas, C. Martinho, R. Prada, and A. Paiva, “The influence of empathy in human–robot relations,” International journal of human-computer studies , vol. 71, no. 3, pp. 250–260, 2013

  13. [21]

    Adaptive emotional expression in robot-child interaction,

    M. Tielman, M. Neerincx, J.-J. Meyer, and R. Looije, “Adaptive emotional expression in robot-child interaction,” in Proceedings of the 2014 ACM/IEEE international conference on Human-robot interac- tion, 2014, pp. 407–414

  14. [22]

    Security issues and challenges in healthcare automated devices,

    A. Jangid, P. K. Dubey, and B. Chandavarkar, “Security issues and challenges in healthcare automated devices,” in 2020 International Conference on COMmunication Systems & NETworkS (COMSNETS) . IEEE, 2020, pp. 19–23

  15. [23]

    The elephant in the room: cybersecurity in health- care,

    A. J. Cartwright, “The elephant in the room: cybersecurity in health- care,” Journal of Clinical Monitoring and Computing , vol. 37, no. 5, pp. 1123–1132, 2023

  16. [24]

    The greeting machine: an abstract robotic object for opening encounters,

    L. Anderson-Bashan, B. Megidish, H. Erel, I. Wald, G. Hoffman, O. Zuckerman, and A. Grishko, “The greeting machine: an abstract robotic object for opening encounters,” in 2018 27th IEEE Interna- tional Symposium on Robot and Human Interactive Communication (RO-MAN). IEEE, 2018...

  17. [25]

    Towards episodic memory- based long-term affective interaction with a human-like robot,

    Z. Kasap and N. Magnenat-Thalmann, “Towards episodic memory- based long-term affective interaction with a human-like robot,” in 19th International Symposium in Robot and Human Interactive Communi- cation. IEEE, 2010, pp. 452–457

  18. [26]

    A multidimensional conception and measure of human-robot trust,

    B. F. Malle and D. Ullman, “A multidimensional conception and measure of human-robot trust,” in Trust in human-robot interaction . Elsevier, 2021, pp. 3–25

  19. [27]

    Measuring gains and losses in human- robot trust: Evidence for differentiable components of trust,

    D. Ullman and B. F. Malle, “Measuring gains and losses in human- robot trust: Evidence for differentiable components of trust,” in 2019 14th ACM/IEEE international Conference on human-robot interaction (HRI). IEEE, 2019, pp. 618–619

  20. [28]

    Be more transparent and users will like you: A robot privacy and user experience design experiment,

    J. Vitale, M. Tonkin, S. Herse, S. Ojha, J. Clark, M.-A. Williams, X. Wang, and W. Judge, “Be more transparent and users will like you: A robot privacy and user experience design experiment,” in Proceedings of the 2018 ACM/IEEE international conference on human-robot interacti...

  21. [29]

    The development of overtrust: An empirical simulation and psychological analysis in the context of human–robot interaction,

    D. Ullrich, A. Butz, and S. Diefenbach, “The development of overtrust: An empirical simulation and psychological analysis in the context of human–robot interaction,” Frontiers in Robotics and AI , vol. 8, p. 554578, 2021

  22. [30]

    An explanation is not an excuse: Trust calibration in an age of transparent robots,

    A. R. Wagner and P. Robinette, “An explanation is not an excuse: Trust calibration in an age of transparent robots,” in Trust in human-robot interaction. Elsevier, 2021, pp. 197–208

  23. [31]

    An experimental security analysis of an industrial robot controller,

    D. Quarta, M. Pogliani, M. Polino, F. Maggi, A. M. Zanchettin, and S. Zanero, “An experimental security analysis of an industrial robot controller,” in 2017 IEEE Symposium on Security and Privacy (SP) . IEEE, 2017, pp. 268–286

  24. [32]

    Security of indus- trial robots: Vulnerabilities, attacks, and mitigations,

    H. Pu, L. He, P. Cheng, M. Sun, and J. Chen, “Security of indus- trial robots: Vulnerabilities, attacks, and mitigations,” IEEE Network, vol. 37, no. 1, pp. 111–117, 2022

  25. [33]

    Ethics guidelines for trustworthy ai,

    E. Commission, “Ethics guidelines for trustworthy ai,” 2019, accessed: 2025-03-25. [Online]. Available: https://digital-strategy.ec.europa.eu/ en/library/ethics-guidelines-trustworthy-ai

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.