{"id":"3e94fd21-0181-4242-885e-efbec17dd5a4","arxiv_id":"2509.01450","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"People in human-robot assembly tasks hesitate to ask human helpers for support, but report they would ask robots sooner and prefer proactive, non-intrusive robotic assistance.","lead":"In a lab assembly task with a robot, 20 people delayed asking a remote human helper for help, with many showing clear signs of stress while hesitating. Most said they would feel more comfortable and ask sooner if the assistant were a non-judgmental robot, suggesting concrete design features for future assistive agents.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim about robot-assistance preferences rests on hypothetical interviews, not on any deployed robot condition.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: stated willingness to ask a robot for help does not predict behavior with a real robot. The study is a well-described preliminary qualitative/quantitative exploration, and the Limitations section explicitly acknowledges the lack of robot assistance, which is appropriate. However, the Conclusions and abstract overreach by presenting robot-assistance preference as a finding rather than a hypothesis. The internal inconsistencies in Table I further undermine the quantitative foundation but are addressable. Since the paper already has a CONDITIONAL verdict, my stress-test does not change that assessment; it reinforces the need for a real-robot condition before the design guidelines are treated as validated. No ad hominem or theatrical language is needed: the issue is scope of evidence, not researcher intent.","tokens_in":10955,"tokens_out":2921,"duration_ms":33026,"concrete_test":"Run a between-subjects follow-up using the same 32-piece assembly task with three conditions: (1) remote human on-demand assistant, (2) robot on-demand assistant, and (3) proactive robot assistant with indicator-light/hint behaviors, all with the same measurements (time from need onset to request, completion rate, stress coding with inter-rater reliability, and post-task interviews). If the robot conditions do not significantly reduce help-seeking delay or improve completion rate relative to human assistance, the central claim that users prefer and benefit from proactive robotic assistance is not supported. Additionally, recompute Table I statistics from raw behavioral logs and verify that reported means, SDs, and 95% CIs are mutually consistent; correct any transcription errors before reinterpreting the reluctance data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Conclusions state users 'prefer proactive, non-judgmental robotic assistance to enhance performance in HRC tasks,' but the study deployed only a remote human assistant. All robot-related results come from post-task interviews about imagined scenarios (Section V, Figs. 4-5); no condition tested a real robot assistant, proactive assistance, or unsolicited help. The quantitative reluctance (2.1-min delay, 28% time in help-need states) was measured with human assistance only, and the paper itself concedes in Section VII that the study 'is limited to human assistance.' Self-reported willingness to ask a robot for help is a weak proxy for actual behavior, especially because participants never encountered a robot's failure modes, response delays, or interaction costs. Therefore the central claim is an unsupported extrapolation from imagined-preference data to a behavioral conclusion. This concern is compounded by internally inconsistent statistics in Table I (e.g., Level 3 time: M=0.9, SD=4.8, but CI=[0.3,1.6]; Total time: M=7.5, SD=2.5, CI=[4.4,10.2]), which makes even the human-assistance reluctance magnitude uncertain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a single-condition user study (N=20) in which participants assembled a 3D puzzle alongside a UR5 robot and could request help from a remote human assistant by saying 'NEED HELP.' The task was intentionally made difficult by hiding a required piece and presenting grayscale instructions. Video annotations were used to classify time spent in four help-need levels, yielding a mean delay of 2.1 minutes between apparent need onset and help request and an average of 28.1% of task time spent in help-need states. Semi-structured interviews explored participants' feelings about human assistance and their attitudes toward hypothetical robot assistance, including on-demand and proactive forms. From these data, the paper proposes design guidelines for assistive robots and concludes that users prefer proactive, non-judgmental robotic assistance.","tokens_in":11163,"tokens_out":6527,"duration_ms":73345,"significance":"The study addresses an important and understudied issue in HRC: the reluctance of human workers to ask for help. The direct behavioral measurement of delay-to-ask (2.1 minutes) is a potentially useful empirical contribution, and the qualitative material gives a rich picture of the social and emotional barriers to help-seeking. However, the paper's central conclusion about robot-assistance preferences is based on imagined scenarios rather than interaction with a real robot, and the reported statistics contain internal inconsistencies. The contribution is best characterized as a preliminary, qualitative/descriptive study. Its value depends on correcting the quantitative reporting and re-scoping the claims from established behavior to stated attitudes.","major_comments":[{"comment":"The conclusion that users 'prefer proactive, non-judgmental robotic assistance' is not supported by the data. Section III.B describes a single-condition study in which assistance was provided exclusively by a remote human and only after the participant verbally requested it; no robot-assistance condition and no proactive/unsolicited-assistance condition were run. All robot-related results in Section V and Figs. 4-5 come from post-task interviews about hypothetical scenarios, as Section VII concedes ('it is limited to human assistance'). Self-reported willingness to ask a robot for help is not evidence of actual behavior with a real robot. Please re-scope the abstract, Section V, Section VI, and Section VIII to 'self-reported attitudes toward imagined robot assistance' or add a real robot-assistance condition.","section":"Section VIII / Abstract / Section V-A"},{"comment":"The reported standard deviations and 95% confidence intervals are mutually inconsistent. For Level 3 time, M=0.9, SD=4.8, CI=[0.3,1.6] is impossible for n=20: the CI half-width of 0.65 implies SD approximately 1.4, not 4.8. For Total time, M=7.5, SD=2.5, CI=[4.4,10.2] implies SD approximately 6.2, not 2.5. Similarly, '12.2% (SD=0.01) CI 95%[8.7, 15.6]' in Section IV-A is implausible, as are several other rows. Since the paper's central quantitative findings (2.1-min delay, 28% time in need) rest on these values, the numerical magnitudes and uncertainties are not trustworthy as reported. In addition, the video coding of help-need levels is central to the delay measurement, yet no inter-rater reliability statistic is reported for the two investigators who performed the labeling. Please correct the table and report reliability.","section":"Table I / Section IV-A"},{"comment":"The abstract and several passages speak of analyzing the 'impact' of on-demand versus unsolicited assistance and of human versus artificial assistants, but the study has no control or comparison condition. Participants always had access to a remote human on-demand; there was no no-assistance condition and no proactive/unsolicited condition. Therefore statements about 'impact on task performance' and the design guidelines in Section VI are not established by the data. The paper should describe the findings as descriptive/preliminary or introduce the missing comparison conditions. The deliberate creation of help-need situations (hidden piece, grayscale instructions) also deserves a more prominent caveat: the measured delay and frustration occur in an artificially induced need context, and the representativeness of that context for real HRC tasks is an assumption rather than a demonstrated","section":"Section III.B / Section IV / Section VI"}],"minor_comments":[{"comment":"The numbers are not fully consistent: the text says 'twenty participants feeling comfortable seeking help from a robot' and 'seventeen participants reported no such issues with robot assistance.' Clarify whether these are different subquestions and report the exact wording used in the interview.","section":"Section V-A"},{"comment":"The description of the coding levels says 'Levels 1 through 3 indicate increasing assistance needs,' but the first level is called 'Flow' in Table I. Clarify the naming (Flow vs. Level 1) so the table and definitions align.","section":"Section III-D"},{"comment":"The sentence 'Notably, 17 users were stuck at some point ... for around one minute, and 9 of them for more than 2 minutes' would benefit from defining the threshold for 'stuck' and from reporting the durations as ranges rather than approximate summaries.","section":"Section IV-A"},{"comment":"The figures in Section IV-B present the affinity diagram results with green/red categories; if the paper is read in grayscale, the distinction should be indicated by shape or hatching in addition to color.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a single-condition, qualitative/descriptive HRC study. The direct behavioral delay measurement is a useful contribution, but Table I contains statistical inconsistencies that undermine the precision of the headline numbers, and the robot-preference conclusions go beyond what the study actually tested. I recommend major revision rather than rejection because the human-help-seeking reluctance finding is directly observed and the limitation is acknowledged in Section VII. If the statistical issues cannot be corrected or the claims cannot be re-scoped to stated preferences, the paper would need to be reconsidered. The fit with an HRI/HRC venue is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has one solid empirical nugget and one load-bearing overreach. The nugget is the measured help-seeking delay in a physical human-robot assembly task — a mean of 2.1 minutes between when a participant first signals needing help and when they actually ask. They also spent 28% of task time in assistance-needing states. That is the kind of behavioral ground truth HRI designers need. The overreach is the conclusion that users \"prefer proactive, non-judgmental robotic assistance,\" because no robot assistant was ever deployed. All robot-related results come from post-task interviews about imagined scenarios. The authors admit this in Section VII (\"it is limited to human assistance\"), but the abstract and Conclusions don't carry that caveat.\n\nWhat is actually new: a direct, video-coded measure of help-need levels in a realistic cobot assembly task, with a hidden piece and deliberately ambiguous instructions to create genuine stuck moments. The qualitative analysis of why people hesitate — fear of bothering, being judged — is coherent and consistent with the organizational psychology literature. The design guidelines in Section VI (signal availability, customizable assistance levels, non-judgmental neutrality) are reasonable hypotheses, clearly derived from participant statements rather than invented.\n\nThe soft spots are real but fixable. Table I has internally inconsistent statistics: Level 3 time is reported as M=0.9, SD=4.8, CI=[0.3,1.6], and the total time CI is too narrow; these numbers cannot all be true with N=20. The \"70% of users displayed clear frustration and stress signs\" claim needs a rubric and inter-rater reliability data. The abstract promises analysis of \"on-demand versus unsolicited help,\" but the study never manipulates that; it only discusses it hypothetically. And the participant counts shift between sections (six vs. eight vs. twenty for robot comfort), which is confusing even if the metrics differ.\n\nNone of this kills the core finding, which is the measured delay. But the paper as currently written could mislead a skimmer into thinking a robot condition existed.\n\nWho is this for? HRI researchers who want a behaviorally grounded number to design against, and the workplace help-seeking community as a case study in a human-robot context. Send it to peer review. A careful revision — fix the table, add a coding rubric, and reword the conclusions to separate measured behavior from expressed preferences — would make this a solid contribution.","headline":"Worth a referee: a real behavioral measurement of help-seeking delay in HRC, wrapped in an overreaching robot-preference narrative.","tokens_in":11661,"tokens_out":2541,"would_cite":true,"duration_ms":29389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users in a human-robot assembly task delayed asking a remote helper for an average of 2.1 minutes after recognizing a problem, and interviews indicate they would ask a non-judgmental robot sooner.","keywords":["human-robot collaboration","help-seeking","assistance","proactive robot assistance","user study","design guidelines","non-judgmental assistance","assembly task"],"falsifier":"Run the same assembly task with a real proactive robot assistant implementing the paper's design guidelines and measure the delay from help need to request; if the delay does not drop below 2.1 minutes, or if users report the robot feels intrusive or untrustworthy, the central claim fails.","tokens_in":10813,"feed_emoji":"🤖","tokens_out":4223,"duration_ms":46470,"temperature":0.7,"pith_summary":"This paper argues that people working alongside robots in assembly tasks are reluctant to ask a remote human helper for assistance, even when they are stuck. Using a 20-person study in which a cobot and a human assembled a 3D puzzle, the authors measured an average 2.1-minute gap between the first signs of needing help and the verbal request, with about 28% of task time spent in states of unresolved need. Interviews revealed that most participants avoided human help out of fear of bothering or being judged, while most said they would ask a robot sooner because a robot is non-judgmental and has no feelings. The paper draws design guidelines for assistive robots: proactive but non-intrusive, neutral in tone, customizable in help level, and privacy-transparent. The conclusion is that robotic assistance, built this way, could reduce the stress and performance loss caused by help-seeking reluctance in human-robot collaboration.","feed_headline":"HRC users wait 2.1 minutes before asking for help","feed_subtitle":"Fear of judgment delays requests; users say a non-judgmental robot would get asked sooner.","key_machinery":"The instrument that carries the argument is a four-level behavioral coding scheme applied to video recordings: Flow, then Level 1 (social cues of uncertainty), Level 2 (visible attempts to resolve the problem), and Level 3 (verbal help request). This scheme turns the invisible moment of 'needing help' into a time-stamped event, letting the authors compute the delay between need onset and request. The second mechanism is the semi-structured interview and affinity-diagram analysis, which connects those delays to participants' feelings about human versus robot assistance, and produces the design guidelines.","core_discovery":"On the paper's own terms, the central discovery is that reluctance to ask for help is measurable and consequential in human-robot collaboration, and that users' stated preferences point toward proactive, non-judgmental robot assistance as a remedy. In a controlled HRC assembly task with intentional ambiguities, participants showed signs of needing help for a mean of 7.5 minutes per task and took a mean of 2.1 minutes after recognizing a problem before saying 'NEED HELP'; 70% displayed visible frustration or stress during this period. Interview data then supplied the mechanism: 12 of 20 participants reported barriers to asking humans, citing fear of bothering or being judged, whereas 20 of 20","pith_inferences":["If the 2.1-minute delay reflects a general human cost of help-seeking, then any assistive agent—human or artificial—should be evaluated on how much it shortens that delay, not only on task success.","The same non-judgmental preference may generalize beyond factories to remote work, education, and health coaching, where asking a human feels socially risky; this is an untested extrapolation.","A real robot implementation could reveal backlash: users who imagine asking sooner may, in practice, find proactive prompts annoying or intrusive, especially experts; the paper's own data on prompt-assistance discomfort hints at this.","The hidden-piece and grayscale-instruction task is a reusable elicitation method for help-seeking studies, since it reliably creates stuck states without artificially forcing requests."],"forward_implications":["Designers of assistive robots should treat the time between need and request as a primary performance metric, not just task completion time.","A robot that offers help without stopping its own work, signaled by a light, may reduce users' feeling that they are bothering it and encourage earlier requests.","Robot assistants should default to hints and spatial guidance rather than full solutions, because participants wanted to retain autonomy.","Privacy transparency—no recording, no performance tracking, a visible active-state indicator—may be a precondition for trust in proactive assistance.","Because users preferred asking a robot sooner, proactive assistance may be accepted even if on-demand remains preferred for human helpers."],"supporting_citations":[{"why":"Supplies prior evidence that employees remain silent rather than communicate problems upward, grounding the reluctance finding.","marker":"[7]"},{"why":"Links performance anxiety to threat interference and working-memory impairment, supporting the claim that delayed help harms performance.","marker":"[3]"},{"why":"Documents underestimation of the discomfort of help-seeking, supporting the interpretation of the measured delays.","marker":"[12]"},{"why":"Supports the claim that AI can provide non-judgmental on-demand support that reduces stigma associated with seeking human help.","marker":"[17]"},{"why":"Shows that proactive robot suggestion improves task efficiency, serving as the baseline for proactive assistance design.","marker":"[18]"},{"why":"Shows predictive assistance reduces cognitive load, supporting the proactive-assistance direction the paper recommends.","marker":"[19]"},{"why":"Defines the flow state used as the baseline level in the video-coding scheme.","marker":"[34]"},{"why":"Supplies the affinity-diagram methodology used to organize interview comments into themes and design guidelines.","marker":"[36]"}],"fun_headline_variants":["2.1-min stall: why users fear asking robots for help","Fear of judgment delays robot help requests by 2.1 min","Users wait 2.1 min to ask robots, cite fear of judgment","70% show frustration before asking robot for help","All 20 users want judgment-free robot help in HRC"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that participants' interview-stated willingness to ask a robot for help predicts what they would actually do with a real assistive robot, since the study only provided a remote human assistant.","fun_headline_variants_meta":{"raw":{"variants":["2.1-min stall: why users fear asking robots for help","Fear of judgment delays robot help requests by 2.1 min","Users wait 2.1 min to ask robots, cite fear of judgment","70% show frustration before asking robot for help","All 20 users want judgment-free robot help in HRC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001893,"raw_usage":{"total_tokens":7230,"prompt_tokens":689,"completion_tokens":6541,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":6467}},"tokens_in":433,"tokens_out":6541,"duration_ms":40740,"temperature":1.0,"reasoning_tokens":6467,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:30:03.047355+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same assembly task with a real proactive robot assistant implementing the paper's design guidelines and measure the delay from help need to request; if the delay does not drop below 2.1 minutes, or if users report the robot feels intrusive or untrustworthy, the central claim fails.","supporting_citations":[{"cited_title":"An exploratory study of employee silence: Issues that employees don’t communicate upward and why,","cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence that employees remain silent rather than communicate problems upward, grounding the reluctance finding."},{"cited_title":"I’m going to fail! acute cognitive performance anxiety increases threat-interference and impairs wm performance,","cited_arxiv_id":null,"evidence_quote":"Links performance anxiety to threat interference and working-memory impairment, supporting the claim that delayed help harms performance."},{"cited_title":"“why didn’t you just ask?","cited_arxiv_id":null,"evidence_quote":"Documents underestimation of the discomfort of help-seeking, supporting the interpretation of the measured delays."},{"cited_title":"Beyond text: ChatGPT as an emotional resilience support tool for gen z–a sequential ex- planatory design exploration,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that AI can provide non-judgmental on-demand support that reduces stigma associated with seeking human help."},{"cited_title":"Efficient human-robot collaboration: when should a robot take initiative?","cited_arxiv_id":null,"evidence_quote":"Shows that proactive robot suggestion improves task efficiency, serving as the baseline for proactive assistance design."},{"cited_title":"Czikszentmihalyi, Flow: The psychology of optimal experience","cited_arxiv_id":null,"evidence_quote":"Defines the flow state used as the baseline level in the video-coding scheme."},{"cited_title":"Using affinity diagrams to evaluate interactive prototypes,","cited_arxiv_id":null,"evidence_quote":"Supplies the affinity-diagram methodology used to organize interview comments into themes and design guidelines."}],"review_version":1}