{"id":"56b75544-ba33-4746-a5e8-14217c3389cf","arxiv_id":"2506.21322","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 176-participant online video study, home healthcare robots that explained themselves shifted blame from the user's memory to third-party tampering, yet most participants still followed the robot's medication advice.","lead":"This study shows that when a home healthcare robot gives medication advice that conflicts with a person's memory, how clearly the robot explains itself changes who people blame for the mismatch. It also finds that people usually follow the robot's advice even when they suspect the robot may be wrong or tampered with.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The transparency effect on interpretation (RQ2) is the paper's novel claim, but it rests on descriptive percentages with no inferential test, so the central claim is currently unsupported.","rationale":"The paper is a competently reported vignette study, and the overtrust result is genuinely supported by a binomial test, so I do not see grounds for full rejection. However, the headline claim about transparency influencing interpretation is the novel contribution, and Section V.B provides no inferential statistic linking transparency to the interpretation categories. The Discussion's use of 'significantly larger proportion' is unsupported by any reported test. The reader's weakest assumption correctly identifies the bundled manipulation and missing manipulation check, which is a real construct-validity issue; my concern is adjacent but more fundamental: even before asking what caused the difference, we have no statistical evidence that the difference is real. A Fisher-Freeman-Halton or Monte Carlo chi-square test on Table VI would settle this directly. The verdict should remain CONDITIONAL because the problem is addressable and the overtrust finding stands, but the transparency claim needs either a significant inferential test or a substantial softening of the wording.","tokens_in":10734,"tokens_out":4410,"duration_ms":57878,"concrete_test":"Run a Fisher-Freeman-Halton exact test (or Monte Carlo chi-square) on the 2x8 table of transparency level (LT vs HT) by the eight interpretation categories in Table VI, using the raw counts behind the percentages and including 'Other'. Report the p-value and Cramér's V. If p>=.05, the abstract's RQ2 transparency claim is indistinguishable from chance; if p<.05, the statistical existence of the effect is established, though the construct-validity confound would still need a separate manipulation-check study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V.B reports only percentages (Table VI) for how participants interpreted the discrepancy; no chi-square, Fisher, or regression is run for RQ2. Section VI.A nonetheless asserts that 'a significantly larger proportion' in the high-transparency condition attributed the discrepancy to external factors. That significance is not established anywhere, and with only 176 single-response codings spread over eight categories, the LT vs HT differences (User-related 50.0% vs 21.59%; Modified-by-others 4.55% vs 35.23%) could easily arise by chance. The overtrust finding is different: it is supported by a binomial test (p<.001) and a chi-square with a reported p-value, so this attack is targeted at the transparency-interpretation link, not the whole paper. In addition, the transparency manipulation bundles detailed explanations with system records and longer speech (Sec. IV.C), with no manipulation check; even if a test were significant, the causal attribution to transparency would remain confounded. This is why the core contribution, as stated in the abstract, is not yet secure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a 2×2 between-subjects online study (N=176) in which participants watched videos of a Furhat robot that contradicts a fictional user's memory about medication time. It examines how transparency and sociability affect decision-making, interpretation of the discrepancy, and perceived trust. The central findings are that participants tended to follow the robot's recommendation (supported by a binomial test, p < .001), while the transparency manipulation is claimed to shift users' attribution of the discrepancy from user-related causes to external causes on the basis of descriptive percentages only. The paper argues for overtrust in healthcare robots and for the importance of access-control mechanisms in multi-user home environments.","tokens_in":10882,"tokens_out":4584,"duration_ms":50985,"significance":"The overtrust finding is a clean, confirmatory result with appropriate inferential support and speaks to a timely and safety-relevant HRI problem. If the transparency–attribution link were statistically established, it would be a novel and practically valuable contribution to the design of transparent home healthcare robots and to discussions of access control. However, as reported, the paper's headline RQ2 claim is not supported by any statistical inference and is potentially confounded by bundled manipulation components. The study design and scenario are otherwise appropriate and the paper is clearly written, but the main novel claim needs substantial reinforcement before the findings can be considered reliable.","major_comments":[{"comment":"The central RQ2 claim that transparency influenced interpretations is supported only by descriptive percentages in Table VI; no chi-square, Fisher exact test, logistic regression, or confidence intervals are reported. The discussion in Sec. VI.A nevertheless states that 'a significantly larger proportion' of high-transparency participants attributed the discrepancy to external factors, but this significance is never established anywhere in the paper. Please add a proper inferential test on the contingency table (e.g., a 2×2 Fisher exact test contrasting user-related vs. all other categories by transparency level, or a multinomial/ordinal regression), with effect sizes and exact p-values, or revise the abstract and discussion to avoid the causal and significance wording.","section":"Sec. V.B / Table VI and Sec. VI.A"},{"comment":"The transparency manipulation bundles multiple components: detailed explanations, system records, and longer utterances, while the low-transparency robot simply restates the information. No manipulation check is reported, so any observed attribution difference could be driven by the quantity or length of information rather than by transparency as a defined construct. Please include a perceived-transparency manipulation check or an experimental design that controls for utterance length/information volume, and discuss this confound explicitly in the limitations.","section":"Sec. IV.C"},{"comment":"The qualitative coding that produces the eight categories in Table V and the percentages in Table VI is not described in terms of coding procedure: the number of coders, training, and inter-coder reliability (e.g., Cohen's kappa) are not reported. Because RQ2 hinges on these categories, the reliability of the coding should be documented, or the results should be treated as exploratory rather than confirmatory.","section":"Sec. V.B"},{"comment":"The overtrust discussion in RQ3 is based on descriptive follow rates (87.30% vs. 45.71% vs. 34.29%) without an inferential test; a chi-square or Fisher exact test across attribution categories would support the claim that interpretations are associated with decisions. Additionally, the binomial test in Sec. V.A excludes 10 'No'/'Other' responses (113 vs. 53 out of 176); please justify this exclusion or use the full denominator in the test.","section":"Sec. V.C / Table VII"}],"minor_comments":[{"comment":"The trust questionnaire is referred to as MDMT in the text but as 'MDTM' in the measure description; please correct the abbreviation.","section":"Sec. IV.E"},{"comment":"The phrase 'trust peception' is a typo and should read 'trust perception'.","section":"Sec. VI.A"},{"comment":"The phrasing of RQ2, 'how do ... affect participants assess and interpret the discrepancy', is ungrammatical; please revise to 'affect participants' assessment and interpretation'.","section":"Sec. III"},{"comment":"The sentence 'the most frequently responses indicated that the robot is modified by someone else' should be revised to 'the most frequent response indicated that the robot was modified by someone else'.","section":"Sec. V.B"},{"comment":"The binomial test result is reported only as 'p < .001'; please also report the exact statistic (e.g., n and observed k) to improve transparency and reproducibility.","section":"Sec. V.A"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of cs.HC and addresses a relevant and timely HRI research question. The missing inferential statistics for the main RQ2 claim and the confounded transparency manipulation are fixable through additional analysis or revised framing; the overtrust result itself is on solid ground. I recommend major revision rather than rejection, provided the authors either provide the missing statistical tests and manipulation checks or soften the central claims to match the descriptive evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper has one solid finding and one unsupported headline. The solid finding: across 176 participants, a significant binomial test (p < .001) shows people follow the robot's medication timing over the user's memory, and this overtrust persists even among participants who suspect system malfunction or third-party interference (RQ3, Table VII). That is a real, useful data point for HRI and for anyone designing home healthcare robots.\n\nThe unsupported headline is the transparency effect on interpretation (RQ2). The paper reports only percentages in Table VI and the discussion claims a 'significantly larger proportion' in the high-transparency condition attributed the discrepancy to external factors. No chi-square, Fisher, or regression is run. With 176 codings spread over eight categories, the LT vs HT differences (50% vs 21.59% for user-related; 4.55% vs 35.23% for modified-by-others) could plausibly be chance, especially with multiple categories. This is a load-bearing flaw because RQ2 is the novel contribution, as the abstract leads with it. The reader's stress-test note is correct.\n\nA second soft spot: the transparency manipulation bundles detailed explanations, system records, and longer speech; sociability bundles greeting, self-introduction, memory recall, empathy, and gestures. No manipulation check is reported, so even with an inferential test the causal attribution to 'transparency' or 'sociability' as single constructs is confounded with verbosity and information quantity. The qualitative coding also lacks reliability evidence (no kappa or agreement).\n\nCredit where due: the study is honestly designed, the methods are transparent, the limitations section acknowledges the video-based paradigm and single platform, and the authors do not overfit. The overtrust result is independent of the interpretations and stands on its own.\n\nWho is this for? HRI researchers and designers of home healthcare robots. The overtrust finding is worth a serious referee; the transparency claim needs revision before it can be cited. I'd recommend sending it to peer review with major revision: require inferential statistics for RQ2, manipulation checks, coding reliability, and softer wording in the discussion.\n\nBest.","headline":"Two results: a solid overtrust finding, and a transparency-interpretation claim that currently rests on untested percentages.","tokens_in":11400,"tokens_out":2277,"would_cite":false,"duration_ms":26850,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that when a home healthcare robot contradicts a user's memory, the robot's transparency changes who gets blamed for the contradiction but not whether the user follows the robot's advice, with most users overtrusting the…","keywords":["human-robot interaction","overtrust","robot transparency","robot sociability","healthcare robotics","attribution of blame","multi-user access control","medication adherence"],"falsifier":"A replication that equates the robot's utterance length and informativeness between conditions while varying only whether the explanation is framed as a verifiable system record—and includes a manipulation check for perceived transparency—would test whether attribution differences are really driven by transparency. If the attribution gap disappears when verbosity is controlled, the paper's central interpretive claim would be seriously weakened.","tokens_in":10534,"feed_emoji":"🤖","tokens_out":7366,"duration_ms":64332,"temperature":0.7,"pith_summary":"This paper asks what happens when a home healthcare robot gives advice that contradicts what the user remembers. In a 2×2 online experiment with 176 participants watching scripted videos, the authors find that the robot's transparency changes how people interpret the contradiction: with low transparency the most common explanation is that the user misremembered, whereas with high transparency the most common explanation is that another family member or outside party changed the robot's information. Neither transparency nor sociability changed the bottom-line decision, however: about 68% of participants chose the robot's suggested medication time over their own memory, and many did so even while suspecting a system malfunction or tampering. The authors argue this is evidence of overtrust that designers of home healthcare robots should take seriously.","feed_headline":"Transparency changes blame, not decisions, for home healthcare robots","feed_subtitle":"Transparency shifted whom users blamed, but not their choice to take the robot's medication time.","key_machinery":"The central object is a scripted, first-person video interaction with a Furhat robot—a social robot platform with a projected face—acting as a family healthcare assistant. The discrepancy is fixed: the robot says it is 4 PM and time to take medication, while the user recalls 5 PM. Transparency is operationalized as whether the robot, when challenged, offers detailed explanations and system records (high) or simply restates the information (low); sociability is operationalized as a bundle of greeting, self-introduction, memory recall, empathy, and gestures. The argument is carried by comparing coded open-ended attributions (user-related, robot-related, modified-by-others, clock changes, etc.) across conditions, and by relating those attributions to a forced-choice decision item.","core_discovery":"The paper's central claim is that, in a family healthcare context, a robot's transparency level shifts users' causal attribution of an information discrepancy without shifting their compliance. When the robot simply restates its 4 PM medication reminder without explanation, half the participants assume the user did not remember correctly; when the robot shows detailed explanations and system records, only about a fifth blame user memory, and a third instead suspect that someone else—a partner, child, or other household member—modified the robot's information. Across all four conditions, a binomial test shows participants chose the robot's time significantly more often than their own memory (68.1%), and the authors interpret this as overtrust, noting that even among participants who suspected robot malfunction or third-party interference, large minorities still said they would follow the robot's advice.","pith_inferences":["The authors do not test whether explanation style could be used to steer blame; a follow-up could vary only the framing of the robot's explanation and measure attribution shifts.","The study leaves open whether live interaction would change overtrust; with a physically present robot, users might verify records rather than passively accept a video scenario.","A further untested consequence is that high transparency might act as a low-cost security cue, since it made users suspect external modification; this could be tested by measuring users' checking behavior after a discrepancy."],"forward_implications":["If users follow a robot's medication advice even when they suspect system faults, adding transparency alone is unlikely to calibrate trust in home healthcare robots.","Designers of multi-user home robots need access-control mechanisms that prevent one household member's changes from silently overriding another's health settings.","Users' tendency to attribute discrepancies to their own memory errors in low-transparency conditions could delay detection of actual robot faults or tampering.","The finding that a third of participants who suspected third-party interference still took the robot's advice suggests that security warnings alone may not change behavior."],"supporting_citations":[{"why":"Supplies the concept and prior evidence of overtrust in robots that the paper's decision findings are framed against.","marker":"[6]"},{"why":"Directly motivates RQ2, providing earlier evidence that autonomy and transparency affect blame attributions in human-robot interaction.","marker":"[17]"},{"why":"Establishes the domestic-abuse and malicious-modification risk that the 'information modified by others' interpretation draws on.","marker":"[5]"},{"why":"Provides the Multi-Dimensional Measure of Trust (MDMT) used to measure trust dimensions in RQ4.","marker":"[26]"},{"why":"Shows that transparency affects trust and decision making in the face of robot errors, the prior work whose scope the paper extends to memory-contradiction scenarios.","marker":"[8]"},{"why":"Supports the overtrust-development interpretation offered for why transparency did not calibrate decisions.","marker":"[29]"}],"fun_headline_variants":["Transparency shifts blame, not overtrust, in robot med advice","Robot transparency changes blame, but users still follow it","For home robots, transparency redirects blame but not compliance","Users blame memory less when robot explains, yet still comply","Transparency alters who's blamed, not whether to trust robot"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The transparency and sociability conditions bundle several behaviors at once—extra explanations and system records for high transparency, and greeting, self-introduction, memory recall, empathy, and gestures for high sociability—with no manipulation check and no control for utterance length, so the observed differences in attribution could be caused by how much the robot said rather than by transparency as a defined construct.","fun_headline_variants_meta":{"raw":{"variants":["Transparency shifts blame, not overtrust, in robot med advice","Robot transparency changes blame, but users still follow it","For home robots, transparency redirects blame but not compliance","Users blame memory less when robot explains, yet still comply","Transparency alters who's blamed, not whether to trust robot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1498,"prompt_tokens":977,"completion_tokens":521,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":438}},"tokens_in":593,"tokens_out":521,"duration_ms":5683,"temperature":1.0,"reasoning_tokens":438,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:26:45.126112+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that equates the robot's utterance length and informativeness between conditions while varying only whether the explanation is framed as a verifiable system record—and includes a manipulation check for perceived transparency—would test whether attribution differences are really driven by transparency. If the attribution gap disappears when verbosity is controlled, the paper's central interpretive claim would be seriously weakened.","supporting_citations":[{"cited_title":"Overtrust of robots in emergency evacuation scenarios,","cited_arxiv_id":null,"evidence_quote":"Supplies the concept and prior evidence of overtrust in robots that the paper's decision findings are framed against."},{"cited_title":"Who should i blame? effects of autonomy and transparency on attributions in human-robot interaction,","cited_arxiv_id":null,"evidence_quote":"Directly motivates RQ2, providing earlier evidence that autonomy and transparency affect blame attributions in human-robot interaction."},{"cited_title":"Anticipating the use of robots in domestic abuse: A typology of robot facilitated abuse to support risk assessment and mitigation in human-robot interaction,","cited_arxiv_id":null,"evidence_quote":"Establishes the domestic-abuse and malicious-modification risk that the 'information modified by others' interpretation draws on."},{"cited_title":"A multidimensional conception and measure of human-robot trust,","cited_arxiv_id":null,"evidence_quote":"Provides the Multi-Dimensional Measure of Trust (MDMT) used to measure trust dimensions in RQ4."},{"cited_title":"Transparency in hri: Trust and decision making in the face of robot errors,","cited_arxiv_id":null,"evidence_quote":"Shows that transparency affects trust and decision making in the face of robot errors, the prior work whose scope the paper extends to memory-contradiction scenarios."},{"cited_title":"The development of overtrust: An empirical simulation and psychological analysis in the context of human–robot interaction,","cited_arxiv_id":null,"evidence_quote":"Supports the overtrust-development interpretation offered for why transparency did not calibrate decisions."}],"review_version":1}