{"id":"bb8eec60-b787-412c-8aff-3d53d3bf1e40","arxiv_id":"2501.02127","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"People vary in how they pull objects from a robot, and 48 of 68 participants changed their pulling or handedness behavior across repeated handovers.","lead":"A 68-person user study found three distinct ways people take objects from a robot: a quick pull, a hesitant pull, and holding the object while waiting for the robot to let go. Most participants changed their style across repeated handovers, suggesting fixed grip-release rules may need to be personalized.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The conclusion that humans adapt to the handover itself is confounded by the fixed ordering of objects/failures and the explanation-level manipulation in Section 2.1; a reanalysis with mixed-effects controls is needed before claiming adaptive strategies are necessary.","rationale":"The reader's weakest-assumption analysis correctly identifies the fixed ordering and explanation-level manipulation as the main confound. The paper's headline contribution is descriptive—documenting that human pull behavior and handedness vary across repeated handovers—and that part is plausible from the data as presented. However, the conclusion goes further, asserting that the observed changes show 'humans adapt to the robot based on the interaction' and that robots must therefore modify handover strategies online. This causal claim is load-bearing for the paper's practical recommendation. The experiment as described cannot separate adaptation to the handover itself from reactions to the systematically varied explanation levels or from the fixed object/failure sequence. The proposed mixed-effects reanalysis would directly test whether trial position still predicts behavior after controlling for the alternative causes. Since the reader already conditioned the verdict on further validation, my stress-test does not shift the verdict; it sharpens the specific test needed. The paper deserves credit for a clearly presented taxonomy and for releasing the observation that behavior varies within a session, but the necessity of adaptive control should be treated as a hypothesis, not an established conclusion, until the confound is addressed.","tokens_in":3727,"tokens_out":2232,"duration_ms":24325,"concrete_test":"Reanalyze the logged handover trials from [6,7] with a generalized linear mixed model predicting behavior category (PF/PS/HNP and handedness) from trial index (1–8) while including explanation level, object type, and failure type as fixed effects and participant as a random intercept. If the trial-index effect is no longer significant or is substantially reduced when explanation level and order covariates are included, the observed ΔB changes are confounded and the adaptation claim fails. As a secondary check, count ΔB>0 transitions only within adjacent trials that share the same explanation level; if most changes occur at explanation-level boundaries, the explanation manipulation is the likely driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 states every participant underwent the same order of objects and robotic failures and was exposed to different explanation levels at each round. The ΔB measure in Section 4 collapses across the 8 robot-to-human handovers, so a reported 'behavior change' could be a response to the explanation manipulation, to a particular object or failure type, or to practice effects, rather than adaptation to repeated handovers with the robot. The paper reports no statistical test that change rates exceed chance, nor any control for trial order. Because all 68 participants share the same sequence, there is no between-participant counterbalancing that could separate these causes. The central recommendation—that robots must plan and adapt handover strategies online because 'humans adapt to the robot based on the interaction'—depends on attributing the observed changes to repeated handover experience. That attribution is not identifiable from the current analysis. The descriptive taxonomy (PF/PS/HNP, handedness) may stand, but the causal claim of human adaptation and the necessity conclusion are not supported by the reported data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 68 participants performing a collaborative shelf-filling task with a robot, focusing on robot-to-human handovers. The robot used a pull-force threshold (3 N) for grip release with a timed automatic release after 10 seconds. The authors classify pulling behavior into three categories (Pull Fine, Pull Slow, Hold No Pull), plus a handedness dimension, and introduce a ΔB metric to quantify within-session behavior changes. They report that 48 of 68 participants showed some change, with 27 showing a large change, and conclude that robots must plan and adapt handover strategies online because humans adapt to the robot during interaction.","tokens_in":3954,"tokens_out":3657,"duration_ms":35533,"significance":"The paper provides a descriptive taxonomy of human handover-taking behaviors in a relatively large study (N=68), with concrete, replicable grip-release parameters (3 N threshold, 10 s timeout). If the behavioral variability is real, the finding that a substantial fraction of users do not conform to the expected pull-to-release pattern is useful for designing more robust handover controllers. However, the central interpretive claim — that the observed changes represent adaptation to the handover itself, and that adaptivity is therefore necessary — is not supported by the reported analysis due to uncontrolled confounds. The descriptive contribution could stand, but the paper's conclusions overreach the evidence.","major_comments":[{"comment":"The fixed ordering of objects, robotic failures, and round-varying explanation levels confounds the ΔB measure. All 68 participants experienced the same sequence of eight robot-to-human handovers, yet the ΔB values in Table 2 collapse across these handovers without controlling for trial order, object type, failure type, or the explanation-level manipulation. Consequently, the reported behavior changes could reflect responses to the explanation condition, to specific objects or failures, or practice effects, rather than adaptation to repeated handovers with the robot. The claim in Section 5 that 'humans adapt to the robot based on the interaction' and the resulting necessity of online adaptation require an analysis that can separate these causes; no such analysis is provided.","section":"§2.1, §4, §5"},{"comment":"The definitions of Pull Fine, Pull Slow, and Hold No Pull are not operationalized from the recorded force sensor data. For example, PS is described as 'little to no pull as they figure out how the robot releases the object,' and HNP relies on whether a 'sufficient pull' is applied. No thresholds for force magnitude or duration beyond the 3-second cutoff are given, and no inter-rater reliability is reported for the classification. Because the entire ΔB analysis depends on these subjective categories, the quantitative results in Table 2 lack a demonstrated basis in the measured data.","section":"§3.1"},{"comment":"The ΔB metric assigns arbitrary magnitudes to behavior changes with no validation. A change in pull category is always weighted more heavily than a change in handedness, and the numeric values (1, 2, 3) are presented as if they reflect meaningful degrees of change. The claim that 'a high number of people showing a large (27) and moderate (17) change' is quantitatively meaningful presupposes that these magnitudes are calibrated to interactional significance. No evidence or sensitivity analysis is offered to support this.","section":"§4, Table 1"},{"comment":"The conclusion that the occurrence of different pull-based behaviors 'make it necessary for a robot to plan and adapt its handover strategies' is not supported by the study design. The paper measures behavior but does not measure the cost of a fixed strategy (e.g., task success, completion time, user trust, or satisfaction), nor does it compare adaptive and non-adaptive strategies. At most, the data support the weaker claim that human pulling behavior is variable and changes within a session, suggesting that adaptive strategies are a promising direction for future investigation.","section":"§5"}],"minor_comments":[{"comment":"The text states that handovers were 'needed for 8 objects in each experiment per participant,' but Section 2.1 describes 4 rounds of 4 objects. Please clarify the relationship between the 16 object trials and the 8 handovers, and state explicitly that each participant performed exactly eight robot-to-human handovers.","section":"§2.2"},{"comment":"The row 'Sum 20 48' is cryptic. It would be clearer to provide a separate row or sentence stating 'No change (ΔB=0): 20 participants; change (ΔB>0): 48 participants.'","section":"Table 2"},{"comment":"The caption reads 'A snippet of a robot-to-handover in the study'; this should be 'robot-to-human handover.'","section":"Figure 1"},{"comment":"The definition of PF sets a 3-second limit after 'first human contact' but does not specify how first contact was determined (manual coding vs. force sensor event). A brief operational definition would aid reproducibility.","section":"§3.1.1"},{"comment":"The phrase 'a popular technique in the literature [2–5]' would benefit from naming the technique as pull-force thresholding, which is already done in Section 2.2. No action required beyond consistency.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short conference-style report of an observational dataset. The descriptive taxonomy and the reported counts could be a useful contribution to HRI practice, but the design (fixed order, no counterbalancing, no statistical control) makes the causal 'adaptation' claim unidentifiable. A major revision that either (a) restricts conclusions to descriptive observations of variability and explicitly acknowledges confounds, or (b) reanalyzes the data with appropriate mixed-effects models and inter-rater reliability, would be necessary before publication. Editors may also wish to consider whether the scope of the current analysis is sufficient for the journal's standards, given that the conference version is already published (HAI '23)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core of this paper is the three-way pull behavior taxonomy (PF/PS/HNP), the handedness dimension, and the delta B change metric applied to 68 participants. That is a genuine descriptive addition to the HRI handover literature, and the within-session variation they report (48/68 with delta B > 0) is worth knowing. The paper is honest as a short empirical report, and the categories are clear enough for other researchers to try coding their own handover videos.\n\nWhat holds it back is not the taxonomy but the leap from observation to mechanism. The conclusion says humans adapt to the robot based on the interaction, but every participant went through the same object order, the same failure order, and a changing explanation level per round. The observed changes could be practice effects, responses to particular objects or failures, or reactions to the explanation manipulation. The paper reports no statistical test, no confidence intervals, no inter-rater reliability on the behavioral categories, and no control for trial order. The PS category in particular relies on 'little to no pull' judgments with no coding protocol. Given that the authors list 'studying the factors leading to the changes' as future work, the current data cannot identify the cause.\n\nI would not claim the conclusion is false; the descriptive pattern is suggestive and the recommendation for adaptive handover controllers is reasonable. But the necessity conclusion—'make it necessary for a robot to plan and adapt its handover strategies'—is stronger than the evidence. A mixed-effects reanalysis of the trial-level data, with trial number, object, failure type, and explanation level as predictors, would go a long way. If that is not possible, the paper should be reframed as observational and hypothesis-generating.\n\nThis deserves a serious referee for a workshop or late-breaking track, and with revisions it could be an acceptable short paper. It is not a desk-reject: there is a real dataset and a usable taxonomy. The citation pattern is appropriate—the grip-release technique and the experimental design are the authors' own prior work, and the categories do not depend on unverified external claims. I'd bring it to a reading group focused on handover control or adaptive HRI, but I would not cite it as evidence that humans adapt to the robot itself without a confound-free follow-up.","headline":"A useful descriptive taxonomy of handover-taking behavior, but the 'humans adapt' conclusion is confounded by fixed task ordering—keep the categories, temper the causal claim.","tokens_in":4487,"tokens_out":1951,"would_cite":true,"duration_ms":19024,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Humans change how they take objects from a robot across repeated handovers, and the paper argues robots must adapt their grip-release strategy accordingly.","keywords":["human-robot handover","pull force behavior","grip release strategy","user study","behavior adaptation","handedness","human-robot interaction"],"falsifier":"Run the same 68-participant protocol with object order and failure types counterbalanced across participants and explanation level held constant; if the ΔB distribution (48 changers, 27 large) shrinks toward zero, the behavior changes are artifacts of order or explanation rather than evidence that handover-taking behavior adapts.","tokens_in":3555,"feed_emoji":"🤖","tokens_out":6227,"duration_ms":53146,"temperature":0.7,"pith_summary":"This paper tries to establish that people do not take objects from a robot in one fixed way, and that their taking behavior changes within a single session of repeated handovers. In a study of 68 novices completing a collaborative shelf-filling task, 48 participants changed their behavior at least once, and 27 switched between pulling the object out and simply holding it until the robot released the object automatically. The authors categorize behavior along two dimensions, pull force and handedness, and introduce a ΔB measure to quantify the amount of change. If the result holds, robots that release objects using a fixed pull-force threshold will mismatch a large fraction of users, and grip-release strategies should adapt online to the current user.","feed_headline":"48 of 68 users changed handover behavior with a robot","feed_subtitle":"In a 68-person study, 48 altered their pulling or handedness behavior across repeated handovers.","key_machinery":"The load-bearing mechanism is the ΔB behavior-change score, built from a two-dimensional coding scheme: pull-force behavior (Pull Fine, Pull Slow, Hold-no-pull with Verbal Commands) and handedness (one hand versus two hands). The score assigns 3 to a large change between pulling and holding without pulling, 2 to a change between Pull Fine and Pull Slow, 1 to a change in handedness alone, and 0 to no change. This coding is used to summarize handover behavior and to count how many participants changed across repeated robot-to-human handovers. The robot side of the setup is a pull-force thresholding grip-release strategy set at 3 N, with a 10-second timed automatic release, which defines the two behaviors that the ΔB categories distinguish.","core_discovery":"The central claim, stated on the paper's own terms, is that human handover-taking behavior is diverse and changes with repeated interaction: only 20 of 68 participants kept the same behavior across all handovers, while 48 showed a change and 27 showed a large change between pulling the object out and holding it without pulling until the timed automatic release. The paper classifies pull behavior as Pull Fine (object removed within 3 seconds of contact), Pull Slow (more than 3 seconds while the person tests how the robot releases), and Hold-no-pull with Verbal Commands (no sufficient pull; the robot releases only by its 10-second timeout). Handedness adds a second dimension, one hand versus two hands. From these categories the paper builds the ΔB scale, prioritizing changes in pulling over handedness, and reports that 17 participants showed a moderate change, 4 a small change, and 27 a large change. The stated conclusion is that the occurrence of different pull-force behaviors makes it necessary for a robot to plan and adapt its handover strategies rather than rely on a fixed threshold.","pith_inferences":["Because every participant saw the same object and failure order and a different explanation level each round, part of the reported ΔB may reflect order or explanation effects rather than adaptation to the handover itself; a counterbalanced replication would separate these.","The 10-second automatic release may itself teach users to wait, so the large shift toward hold-no-pull could partly be learned behavior; varying the timeout would isolate this mechanism.","The paper's taxonomy could be used as the input to an online classifier that adapts release strategy within the first few handovers, though the paper does not build such a controller."],"forward_implications":["A fixed grip-release threshold is insufficient: 27 of 68 participants either pulled then stopped or held without pulling, behaviors the 3 N threshold does not directly serve.","Handover strategy should be personalized to the current user, since different participants exhibit different stable behaviors.","Strategy should be adapted online during the same interaction, because nearly half the participants changed behavior across repeated handovers.","Modalities beyond force, such as verbal commands, should be considered, as some participants spoke to the robot expecting it to obey.","The stated future direction is to identify the factors that drive behavior changes and let the robot observe them and adapt in advance."],"supporting_citations":[{"why":"Documents the broader case for adaptivity in human-robot interaction that motivates reading behavior changes.","marker":"[1]"},{"why":"Supplies a representative human-inspired handover controller that grip-release strategies in this study build on.","marker":"[2]"},{"why":"Shows proactive release behavior alternatives, contextualizing the pull-threshold approach the study tests.","marker":"[4]"},{"why":"Introduces the pull force thresholding-based grip-release technique with the 3 N threshold used by the robot.","marker":"[5]"},{"why":"Presents the full experimental design whose robot-to-human handover data this paper analyzes.","marker":"[6]"},{"why":"Provides the workshop-level user-study description behind the 68-participant dataset.","marker":"[7]"},{"why":"Provides the mutual-adaptation framework that implies robots should adjust to changing human behavior.","marker":"[8]"}],"fun_headline_variants":["48 of 68 users changed handover style with a robot","Most people adapt their grip and pull when taking from a robot","Robot handovers: 70% of users change behavior on repeat","Study: humans vary pull force and hand use when taking from robots","Only 20 of 68 users kept same handover behavior every time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every participant encountered the same order of objects and robotic failures and a different explanation level in each round, so the paper assumes the observed behavior changes reflect adaptation to the handover itself rather than the fixed order or the explanation manipulation.","fun_headline_variants_meta":{"raw":{"variants":["48 of 68 users changed handover style with a robot","Most people adapt their grip and pull when taking from a robot","Robot handovers: 70% of users change behavior on repeat","Study: humans vary pull force and hand use when taking from robots","Only 20 of 68 users kept same handover behavior every time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000558,"raw_usage":{"total_tokens":2588,"prompt_tokens":812,"completion_tokens":1776,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":1686}},"tokens_in":428,"tokens_out":1776,"duration_ms":12522,"temperature":1.0,"reasoning_tokens":1686,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:14:07.185255+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 68-participant protocol with object order and failure types counterbalanced across participants and explanation level held constant; if the ΔB distribution (48 changers, 27 large) shrinks toward zero, the behavior changes are artifacts of order or explanation rather than evidence that handover-taking behavior adapts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies a representative human-inspired handover controller that grip-release strategies in this study build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows proactive release behavior alternatives, contextualizing the pull-threshold approach the study tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the pull force thresholding-based grip-release technique with the 3 N threshold used by the robot."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents the full experimental design whose robot-to-human handover data this paper analyzes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the mutual-adaptation framework that implies robots should adjust to changing human behavior."}],"review_version":1}