{"id":"1a03064e-43b0-43e4-a127-c8e6c7183e37","arxiv_id":"2501.03515","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A user study shows that robot competency and motion legibility shift when and whether people correct a robot, contradicting two common assumptions in learning-from-corrections.","lead":"This paper reports a 60-person user study on how a robot's apparent competence and the readability of its motions change when people step in to correct it. The authors find that people correct highly competent robots sooner and miss more failures of incompetent robots, which challenges assumptions used in learning from corrections.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's headline RQ1 p-values rely on trial-level ANOVAs that pool 1,944 corrections from only 60 participants (10 per condition), so within-participant clustering may inflate significance; the authors acknowledge this, but the abstract reports the p-values unqualified.","rationale":"The reader identified the same fundamental weakness: the RQ1 trial-level ANOVAs assume independence of 1,944 corrections from 60 participants. I agree this is the most load-bearing concern because the paper's novelty and abstract rest on those p-values. My read is partial rather than full agreement because the same pseudoreplication also affects the RQ3 correlation comparisons in Table I/II, which the reader's weakest-assumption statement did not mention; those correlations pool corrections within condition and compare them with Fisher z-tests that assume independent observations. The concern is about inference, not about the study design or internal consistency: the authors transparently flag the issue in Sec. VI.D, and the direction of effects may well survive clustering. However, as reported, the strength of the central RQ1 claims is not yet established. RQ2's per-participant analyses (F(1,54)) are not subject to this critique, and the overall contribution is a useful empirical study with transparent limitations. A conditional verdict is therefore appropriate; the paper should either provide clustered reanalyses or restrict its headline claims to per-participant results. I would not reject the paper, and I would not change the reader's conditional verdict.","tokens_in":93,"tokens_out":3977,"duration_ms":101685,"concrete_test":"Re-run the RQ1 analyses on the existing trial-level data with participant as a random intercept (and if the design supports it, random slopes for competency and legibility) in a linear mixed-effects model, or equivalently run a two-way ANOVA on per-participant means (60 observations). Record whether the competency-by-legibility interaction and the Tukey contrasts for legible and predictable motions remain below α=0.05, and report the adjusted p-values and intraclass correlations. If any headline contrast loses significance, the abstract's unqualified p-values should be replaced by per-participant or clustered analyses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that people correct a high-competency robot at smaller task-objective divergence (p=0.0015 for legible, p=0.0055 for predictable) and earlier in the trajectory—rests on two-way ANOVAs with degrees of freedom like (2, 1944). Those 1,944 observations are not independent: each of the 60 participants contributed up to 64 trials, and only the first correction per trial was used. Responses within a participant are likely correlated (individual thresholds, attention, calibration to the robot's error rate), so the effective sample size is far below 1,944 and the reported p-values are anti-conservative. Sec. VI.D discloses this, but the abstract presents the p-values as established. The same nesting affects the RQ3 Spearman correlations and Fisher z-tests, where the n's in Table I are corrections pooled across 10 participants per condition. RQ2's per-participant confusion-matrix rates (F(1,54)) avoid this issue and are more robust, but they do not support the RQ1 timing/divergence claims.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subject user study (N=60; 2 competency levels × 3 legibility levels, 10 participants per condition) in which participants supervised and kinesthetically corrected a robot performing 64 pick-and-place trials. The authors measure three families of outcomes: timing of corrections (task-objective divergence, time until correction, proportion of trajectory untraveled), prediction accuracy (missed and unnecessary correction rates, false omission and false correction rates), and the precision–effort tradeoff in corrections (Spearman correlations between a precision metric and integrated torque). They report that people correct a highly competent robot at smaller task-objective divergence and earlier in the trajectory than an incompetent robot when motions are legible or predictable; that people withhold necessary corrections from incompetent robots and give unnecessary corrections to competent robots; and that physical effort is positively correlated with correction precision overall, but significantly less so for an incompetent robot with legible motions than for an incompetent robot with predictable motions. The paper interprets these findings as evidence against three common assumptions in Learning from Corrections and offers design and learning-algorithm recommendations.","tokens_in":18092,"tokens_out":5538,"duration_ms":52419,"significance":"If the headline effects are real, the paper makes a valuable contribution: it provides one of the first empirical demonstrations that a robot's apparent competency biases the timing and necessity of human correction feedback, which has direct implications for LfC algorithms that treat corrections as noisy labels generated from a stable threshold. The RQ2 per-participant confusion-matrix analyses are methodologically sound (F(1,54), with one summary rate per participant), and the empirical test of the precision–effort tradeoff addresses an assumption that is widely used but rarely validated. The concrete recommendations for interaction designers and learning researchers are appropriately grounded in the results. The main limitation, acknowledged in Sec. VI.D, is that the RQ1 and RQ3 analyses treat trial-level data as independent despite nesting within participants, which undermines the headline p-values until reanalysis is performed. The strength of the RQ2 evidence, together with the fixability of the statistical issue, makes the paper suitable for major revision rather than rejection.","major_comments":[{"comment":"The RQ1 two-way ANOVAs treat each correction as an independent observation, with error degrees of freedom around 1944, but the corrections are nested within only 60 participants (up to 64 trials per participant). Within-participant responses are likely correlated due to individual divergence thresholds, attention levels, and calibration to the robot's error rate, so the effective sample size is far smaller than the analysis assumes. This makes the reported p-values (e.g., p=0.0015 for legible high vs. low competency in task-objective divergence; p<0.0001 for time until correction and proportion of trajectory untraveled) anti-conservative. The manuscript itself acknowledges this in Sec. VI.D. Because the abstract presents these p-values without qualification, the central RQ1 claims are load-bearing. Please reanalyze RQ1 using linear mixed-effects models with participant as a random intercept (or participant-level aggregated means) and report whether the direction and significance of the high- vs. low-competency comparisons within the legible and predictable conditions survive.","section":"V-A (Figs. 5–7)"},{"comment":"The RQ3 Spearman correlations are computed on corrections pooled across participants within each condition (n ranges from 153 to 482 per cell), and the pairwise Fisher z-tests in Table II treat these observations as independent. Precision and effort values from the same participant are likely correlated, which would affect both the correlation estimates and their standard errors. The only significant pairwise difference (legible low vs. predictable low, p=0.0075) is based on n=468 and n=461 corrections from just 10 participants per condition, so this result is particularly vulnerable to clustering. Please reanalyze using cluster-robust methods (e.g., bootstrap resampling by participant) or by computing per-participant correlations and testing the difference at the participant level, and report the updated pairwise comparisons.","section":"V-C and Table I/II"}],"minor_comments":[{"comment":"The phrase \"we present an between-subject user study\" should be \"a between-subject user study.\"","section":"Abstract and Sec. I"},{"comment":"In the time-until-correction results, the sentence \"people correct a competent robot earlier in high-competency conditions\" is unclear; it should read \"people correct a highly competent robot earlier in high-competency conditions than in low-competency conditions.\"","section":"Sec. V-A"},{"comment":"In the proportion-of-trajectory results, the text contains \"in in low-competency conditions\" (duplicate preposition) in the legible comparison; please fix this typo.","section":"Sec. V-A"},{"comment":"The sentence beginning \"This which strongly supports the inverse of H1B\" is missing a word or should be rephrased, e.g., \"This result strongly supports the inverse of H1B.\"","section":"Sec. VI.A"},{"comment":"The precision metric normalizes EEF position error and rotation error \"by their mean and standard deviation,\" but it is not stated whether the mean and SD are computed across all corrections, per participant, or per condition; please specify the normalization sample for reproducibility.","section":"Appendix VIII-A"},{"comment":"The paper reports raw p-values throughout but states that the Benjamini–Hochberg procedure was applied; please clarify whether the reported p-values are adjusted, and if they are raw, state which comparisons remained significant after adjustment.","section":"Sec. V-A and V-C"},{"comment":"For the significant correlation difference (legible low vs. predictable low, p=0.0075), please also report an effect size or confidence interval for the difference between correlations, not only the p-value.","section":"Sec. V-C and Table II"}],"recommendation":"major_revision","confidential_remarks":"The authors are transparent about the independence assumption in Sec. VI.D, but the abstract reports the RQ1 p-values unconditionally, and those values depend on an invalid error structure. Since the RQ2 per-participant analyses are robust and the RQ1/RQ3 issues are fixable through reanalysis, I recommend major revision rather than rejection. The final decision should hinge on whether the reanalyzed RQ1 and RQ3 effects survive participant-level clustering; if they do not, the manuscript's contribution would be substantially weakened, as the timing/divergence claims are central to the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on arXiv:2501.03515. The paper is worth engaging with: it's the first user study I know of that varies both robot competency and motion legibility in a supervision-and-correction paradigm, and it directly probes three assumptions that LfC papers quietly rely on. The empirical setup is serious—a Kinova arm with admittance control, 64 pick-and-place trials per participant, and sensible measures (KLD of inferred goal, timing, confusion-matrix rates). The RQ2 findings are the most robust: people miss more necessary corrections when supervising an incompetent robot, and give more unnecessary corrections for a highly competent robot. Those come from per-participant ANOVAs (F(1,54)) and survive as a clean competency main effect.\n\nThe soft spot is exactly where the stress test points. The RQ1 analysis pools 1,944 first corrections across 60 participants and runs ANOVA with df ~1944. Those observations aren't independent, so the p=0.0015 and p=0.0055 values are anti-conservative. The authors disclose this in Section VI.D, but the abstract states them without qualification. The same issue affects the RQ3 correlation comparisons in Table I, where n's are pooled trials per condition, though the one significant contrast (p=0.0075) is between two conditions with n≈460-480, so it's a bit less fragile. The RQ2 analyses avoid the problem entirely and should be the paper's centerpiece.\n\nI'd push the authors to re-run the trial-level analyses with mixed-effects models (participant as random effect) or to demote the RQ1 claims to per-participant aggregates (e.g., each participant's mean divergence at correction). I suspect the direction of the effects will survive, because the pattern is consistent across three correlated measures (divergence, time, proportion untraveled), but the precise p-values will change and some may lose significance. That's not a fatal flaw; it's a fixable statistical issue in an otherwise well-designed study.\n\nThe paper also does a good job with alternatives: it proposes a \"high expectations\" explanation for the inverse trust effects, and checks whether feedback accuracy changes as people calibrate to the robot. The citation pattern seems fair; the self-citations to the authors' prior legibility work are relevant.\n\nBottom line: worthwhile contribution for HRI and interactive learning researchers. It deserves a serious referee, and the right outcome is likely 'revise and resubmit' with the re-analysis as the main request. I'd cite the RQ2 results with a caveat.","headline":"Useful first test of how competency and legibility bias correction feedback, but the headline RQ1 p-values overstate confidence by treating repeated corrections as independent.","tokens_in":18630,"tokens_out":3240,"would_cite":true,"duration_ms":28890,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A robot's apparent competence systematically biases human correction feedback, violating assumptions of learning-from-corrections algorithms.","keywords":["learning from corrections","robot competency","motion legibility","human-robot interaction","kinesthetic teaching","correction feedback","user study","task objective divergence"],"falsifier":"Re-analyze the RQ1 data with mixed-effects models including participant as a random intercept and random slopes for the conditions; if the competency-by-legibility interactions on divergence, time-to-correction, or trajectory-untraveled cease to be significant, the claim that apparent competence shifts correction timing fails.","tokens_in":17680,"feed_emoji":"🤖","tokens_out":7788,"duration_ms":66325,"temperature":0.7,"pith_summary":"The paper tries to establish that two features of a robot's behavior—its apparent competency and the legibility of its motions—systematically change how people supervise and correct it, violating the assumptions implicit in most learning-from-corrections (LfC) algorithms. In a 60-participant study of people correcting a robot arm during pick-and-place tasks, the authors find that people correct a highly competent robot earlier and at smaller task-objective divergences than an incompetent robot when motions are legible or predictable, and that they miss more necessary corrections for incompetent robots while giving more unnecessary corrections to competent ones. They also confirm a positive correlation between physical effort and correction precision overall, but show the correlation weakens significantly for an incompetent robot with legible motions. If correct, these results mean correction data cannot be read as clean, objective labels: the robot's own track record shapes the signal the human produces.","feed_headline":"People correct skilled robots sooner, clumsy robots later","feed_subtitle":"A 60-person study shows correction feedback is biased by the robot's apparent competence, complicating robot learning.","key_machinery":"The engine of the study is a between-subject 2×3 design: 60 participants each supervised a Kinova Gen3 robot arm through 64 pick-and-place trials, with competency set by the robot's intended success rate (25% vs 75%) and legibility by the style of the executed trajectories (predictable, legible, or illegible, the first generated as efficient RRT* paths and the others by optimizing a legibility score). The load-bearing measures are task-objective divergence (the Kullback-Leibler divergence between the goal distribution inferred from the robot's partial motion and the true goal), time until the first correction, the fraction of the intended trajectory left untraveled at correction, missed and unnecessary correction rates framed as a confusion matrix, and the Spearman correlation between correction precision and physical effort. These measures translate raw physical corrections into quantities that can be directly compared against the three LfC assumptions the paper targets.","core_discovery":"The central claim is that the robot's displayed competence shifts both when people intervene and how accurate their intervention labels are, in the opposite direction of what a simple trust story predicts. People supervising a highly competent robot corrected it earlier in its trajectory and at significantly smaller task-objective divergence than people supervising an incompetent robot, for both legible (p=0.0015 for divergence) and predictable (p=0.0055) motions; the same pattern appeared for time until correction and proportion of trajectory untraveled. Missed necessary corrections were far more common in low-competency conditions (11.3% vs 2.8%, p<0.0001), while unnecessary corrections were more common in high-competency conditions (9.8% vs 2.0%, p=0.0171). The authors interpret this as people holding competent robots to a higher standard and giving incompetent robots the benefit of the doubt. The precision–effort tradeoff held, but the correlation was significantly weaker for an incompetent robot with legible motions than for the same robot with predictable motions (p=0.0075).","pith_inferences":["If this effect generalizes, a robot learning from corrections will need an estimate of how competent the human believes it is, because identical corrections carry different task information depending on that belief.","The trust-based hypotheses predict tolerant correction of competent robots; the data show the opposite, pointing to expectation-based strictness. A modeling consequence is that correction thresholds should be functions of expected competence, not just task divergence.","The design confounds the robot's reputation with the robot's actual error distribution, so a cleaner test of the reputation effect would hold the robot's behavior fixed while merely varying the competence label or the participant's prior information.","Because the effort–precision correlation collapsed only for the incompetent-plus-legible condition, a testable extension is to vary legibility continuously and measure whether the correlation drop tracks perceived goal clarity."],"forward_implications":["Algorithms that learn from corrections should stop treating the absence of a correction as an endorsement, especially for low-competency robots, where people systematically miss necessary corrections.","For high-competency robots, a higher rate of unnecessary corrections means algorithms should down-weight corrections as evidence of task-constraint violations.","The precision–effort tradeoff cannot be assumed uniformly; in the incompetent-plus-legible condition, effort is a much weaker guide to correction quality.","Interaction designers can steer feedback quality: competent robots should minimize deviating behavior, while incompetent robots can explore with less risk of triggering misleading corrections.","Robot learning evaluations should record or control the robot's apparent competency and motion legibility, since both change the meaning of the feedback signal."],"supporting_citations":[{"why":"Foundational LfC methods that assume people correct only under significant divergence and trade off precision for effort, providing the assumptions this study tests.","marker":"[16]–[21]"},{"why":"Learning-from-corrections methods that explicitly assume a person's decision to intervene reflects significant task divergence; [23] incorporates legibility into LfC.","marker":"[22], [23]"},{"why":"Defines legible versus predictable robot motion and provides the legibility optimization used to generate the study's three motion conditions.","marker":"[61]"},{"why":"Demonstrates that a robot's task accuracy changes human trust and perceived intelligence, motivating the competency manipulation.","marker":"[68]"},{"why":"Shows an incompetent robot lowers trust in both the robot and the human's own teaching ability, motivating the trust-based hypotheses.","marker":"[69]"},{"why":"Provides the method for estimating the distribution of plausible goals from partial trajectories, the basis of the task-objective divergence measure.","marker":"[89]"},{"why":"Benjamini-Hochberg procedure applied to control false discovery rate across multiple comparisons.","marker":"[93]"}],"fun_headline_variants":["Competent robots get corrected earlier, incompetent ones get a pass","Robot competence skews when humans step in to correct","Skilled robots face stricter correction, clumsy ones forgiven","Correction timing and accuracy depend on robot's shown skill","Human feedback bias: competent robots corrected sooner, clumsy later"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The timing and accuracy results treat each of the roughly 1,944 corrections as an independent observation, even though they come from just 60 participants, so unmodeled within-person correlations could drive the significant interaction effects.","fun_headline_variants_meta":{"raw":{"variants":["Competent robots get corrected earlier, incompetent ones get a pass","Robot competence skews when humans step in to correct","Skilled robots face stricter correction, clumsy ones forgiven","Correction timing and accuracy depend on robot's shown skill","Human feedback bias: competent robots corrected sooner, clumsy later"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3148,"prompt_tokens":1068,"completion_tokens":2080,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":2000}},"tokens_in":684,"tokens_out":2080,"duration_ms":14541,"temperature":1.0,"reasoning_tokens":2000,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:51:43.237387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the RQ1 data with mixed-effects models including participant as a random intercept and random slopes for the conditions; if the competency-by-legibility interactions on divergence, time-to-correction, or trajectory-untraveled cease to be significant, the claim that apparent competence shifts correction timing fails.","supporting_citations":[{"cited_title":"Legibility and predictabil- ity of robot motion,","cited_arxiv_id":null,"evidence_quote":"Defines legible versus predictable robot motion and provides the legibility optimization used to generate the study's three motion conditions."},{"cited_title":"Helping robots learn: a human-robot master-apprentice model using demonstrations via virtual reality teleoperation,","cited_arxiv_id":null,"evidence_quote":"Demonstrates that a robot's task accuracy changes human trust and perceived intelligence, motivating the competency manipulation."},{"cited_title":"The effects of a robot’s performance on human teachers for learning from demonstration tasks,","cited_arxiv_id":null,"evidence_quote":"Shows an incompetent robot lowers trust in both the robot and the human's own teaching ability, motivating the trust-based hypotheses."}],"review_version":1}