Pith. sign in

REVIEW 2 major objections 7 minor 1 cited by

Effects of Robot Competency and Motion Legibility on Human Correction Feedback

T0 review · 2 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A robot's apparent competence systematically biases human correction feedback, violating assumptions of learning-from-corrections algorithms.

desk verdict Useful first test of how competency and legibility bias correction feedback, but the headline RQ1 p-values overstate confidence by treating repeated corrections as independent. read the letter →

arxiv 2501.03515 v1 pith:QPNYWP2H submitted 2025-01-07 cs.RO

classification cs.RO
keywords learningfromcorrectionsrobotcompetencymotionlegibilityhuman-robotinteractionkinestheticteachingcorrectionfeedbackuserstudytaskobjectivedivergence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that two features of a robot's behavior—its apparent competency and the legibility of its motions—systematically change how people supervise and correct it, violating the assumptions implicit in most learning-from-corrections (LfC) algorithms. In a 60-participant study of people correcting a robot arm during pick-and-place tasks, the authors find that people correct a highly competent robot earlier and at smaller task-objective divergences than an incompetent robot when motions are legible or predictable, and that they miss more necessary corrections for incompetent robots while giving more unnecessary corrections to competent ones. They also confirm a positive correlation between physical effort and correction precision overall, but show the correlation weakens significantly for an incompetent robot with legible motions. If correct, these results mean correction data cannot be read as clean, objective labels: the robot's own track record shapes the signal the human produces.

What carries the argument

The engine of the study is a between-subject 2×3 design: 60 participants each supervised a Kinova Gen3 robot arm through 64 pick-and-place trials, with competency set by the robot's intended success rate (25% vs 75%) and legibility by the style of the executed trajectories (predictable, legible, or illegible, the first generated as efficient RRT* paths and the others by optimizing a legibility score). The load-bearing measures are task-objective divergence (the Kullback-Leibler divergence between the goal distribution inferred from the robot's partial motion and the true goal), time until the first correction, the fraction of the intended trajectory left untraveled at correction, missed and unnecessary correction rates framed as a confusion matrix, and the Spearman correlation between correction precision and physical effort. These measures translate raw physical corrections into quantities that can be directly compared against the three LfC assumptions the paper targets.

What would settle it

Re-analyze the RQ1 data with mixed-effects models including participant as a random intercept and random slopes for the conditions; if the competency-by-legibility interactions on divergence, time-to-correction, or trajectory-untraveled cease to be significant, the claim that apparent competence shifts correction timing fails.

Watch

Extended reading notes

Core claim

The central claim is that the robot's displayed competence shifts both when people intervene and how accurate their intervention labels are, in the opposite direction of what a simple trust story predicts. People supervising a highly competent robot corrected it earlier in its trajectory and at significantly smaller task-objective divergence than people supervising an incompetent robot, for both legible (p=0.0015 for divergence) and predictable (p=0.0055) motions; the same pattern appeared for time until correction and proportion of trajectory untraveled. Missed necessary corrections were far more common in low-competency conditions (11.3% vs 2.8%, p<0.0001), while unnecessary corrections were more common in high-competency conditions (9.8% vs 2.0%, p=0.0171). The authors interpret this as people holding competent robots to a higher standard and giving incompetent robots the benefit of the doubt. The precision–effort tradeoff held, but the correlation was significantly weaker for an incompetent robot with legible motions than for the same robot with predictable motions (p=0.0075).

Load-bearing premise

The timing and accuracy results treat each of the roughly 1,944 corrections as an independent observation, even though they come from just 60 participants, so unmodeled within-person correlations could drive the significant interaction effects.

Editorial extensions

If this is right

  • Algorithms that learn from corrections should stop treating the absence of a correction as an endorsement, especially for low-competency robots, where people systematically miss necessary corrections.
  • For high-competency robots, a higher rate of unnecessary corrections means algorithms should down-weight corrections as evidence of task-constraint violations.
  • The precision–effort tradeoff cannot be assumed uniformly; in the incompetent-plus-legible condition, effort is a much weaker guide to correction quality.
  • Interaction designers can steer feedback quality: competent robots should minimize deviating behavior, while incompetent robots can explore with less risk of triggering misleading corrections.
  • Robot learning evaluations should record or control the robot's apparent competency and motion legibility, since both change the meaning of the feedback signal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this effect generalizes, a robot learning from corrections will need an estimate of how competent the human believes it is, because identical corrections carry different task information depending on that belief.
  • The trust-based hypotheses predict tolerant correction of competent robots; the data show the opposite, pointing to expectation-based strictness. A modeling consequence is that correction thresholds should be functions of expected competence, not just task divergence.
  • The design confounds the robot's reputation with the robot's actual error distribution, so a cleaner test of the reputation effect would hold the robot's behavior fixed while merely varying the competence label or the participant's prior information.
  • Because the effort–precision correlation collapsed only for the incompetent-plus-legible condition, a testable extension is to vary legibility continuously and measure whether the correlation drop tracks perceived goal clarity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper reports a between-subject user study (N=60; 2 competency levels × 3 legibility levels, 10 participants per condition) in which participants supervised and kinesthetically corrected a robot performing 64 pick-and-place trials. The authors measure three families of outcomes: timing of corrections (task-objective divergence, time until correction, proportion of trajectory untraveled), prediction accuracy (missed and unnecessary correction rates, false omission and false correction rates), and the precision–effort tradeoff in corrections (Spearman correlations between a precision metric and integrated torque). They report that people correct a highly competent robot at smaller task-objective divergence and earlier in the trajectory than an incompetent robot when motions are legible or predictable; that people withhold necessary corrections from incompetent robots and give unnecessary corrections to competent robots; and that physical effort is positively correlated with correction precision overall, but significantly less so for an incompetent robot with legible motions than for an incompetent robot with predictable motions. The paper interprets these findings as evidence against three common assumptions in Learning from Corrections and offers design and learning-algorithm recommendations.

Significance. If the headline effects are real, the paper makes a valuable contribution: it provides one of the first empirical demonstrations that a robot's apparent competency biases the timing and necessity of human correction feedback, which has direct implications for LfC algorithms that treat corrections as noisy labels generated from a stable threshold. The RQ2 per-participant confusion-matrix analyses are methodologically sound (F(1,54), with one summary rate per participant), and the empirical test of the precision–effort tradeoff addresses an assumption that is widely used but rarely validated. The concrete recommendations for interaction designers and learning researchers are appropriately grounded in the results. The main limitation, acknowledged in Sec. VI.D, is that the RQ1 and RQ3 analyses treat trial-level data as independent despite nesting within participants, which undermines the headline p-values until reanalysis is performed. The strength of the RQ2 evidence, together with the fixability of the statistical issue, makes the paper suitable for major revision rather than rejection.

major comments (2)
  1. [V-A (Figs. 5–7)] The RQ1 two-way ANOVAs treat each correction as an independent observation, with error degrees of freedom around 1944, but the corrections are nested within only 60 participants (up to 64 trials per participant). Within-participant responses are likely correlated due to individual divergence thresholds, attention levels, and calibration to the robot's error rate, so the effective sample size is far smaller than the analysis assumes. This makes the reported p-values (e.g., p=0.0015 for legible high vs. low competency in task-objective divergence; p<0.0001 for time until correction and proportion of trajectory untraveled) anti-conservative. The manuscript itself acknowledges this in Sec. VI.D. Because the abstract presents these p-values without qualification, the central RQ1 claims are load-bearing. Please reanalyze RQ1 using linear mixed-effects models with participant as a random intercept (or participant-level aggregated means) and report whether the direction and significance of the high- vs. low-competency comparisons within the legible and predictable conditions survive.
  2. [V-C and Table I/II] The RQ3 Spearman correlations are computed on corrections pooled across participants within each condition (n ranges from 153 to 482 per cell), and the pairwise Fisher z-tests in Table II treat these observations as independent. Precision and effort values from the same participant are likely correlated, which would affect both the correlation estimates and their standard errors. The only significant pairwise difference (legible low vs. predictable low, p=0.0075) is based on n=468 and n=461 corrections from just 10 participants per condition, so this result is particularly vulnerable to clustering. Please reanalyze using cluster-robust methods (e.g., bootstrap resampling by participant) or by computing per-participant correlations and testing the difference at the participant level, and report the updated pairwise comparisons.
minor comments (7)
  1. [Abstract and Sec. I] The phrase "we present an between-subject user study" should be "a between-subject user study."
  2. [Sec. V-A] In the time-until-correction results, the sentence "people correct a competent robot earlier in high-competency conditions" is unclear; it should read "people correct a highly competent robot earlier in high-competency conditions than in low-competency conditions."
  3. [Sec. V-A] In the proportion-of-trajectory results, the text contains "in in low-competency conditions" (duplicate preposition) in the legible comparison; please fix this typo.
  4. [Sec. VI.A] The sentence beginning "This which strongly supports the inverse of H1B" is missing a word or should be rephrased, e.g., "This result strongly supports the inverse of H1B."
  5. [Appendix VIII-A] The precision metric normalizes EEF position error and rotation error "by their mean and standard deviation," but it is not stated whether the mean and SD are computed across all corrections, per participant, or per condition; please specify the normalization sample for reproducibility.
  6. [Sec. V-A and V-C] The paper reports raw p-values throughout but states that the Benjamini–Hochberg procedure was applied; please clarify whether the reported p-values are adjusted, and if they are raw, state which comparisons remained significant after adjustment.
  7. [Sec. V-C and Table II] For the significant correlation difference (legible low vs. predictable low, p=0.0075), please also report an effect size or confidence interval for the difference between correlations, not only the p-value.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the results are empirical comparisons between manipulated conditions and measured behaviors, not outputs forced by the measures or by self-citations; the disclosed non-independence caveat is a statistical-validity concern, not circularity.

full rationale

This paper is an empirical user study, so its derivation chain is the mapping from manipulated independent variables (robot competency, motion legibility) to measured dependent variables (correction timing, correction rates, precision-effort correlation). I checked each load-bearing link. The task-objective-divergence, time-until-correction, and untraveled-proportion measures are defined independently of the experimental conditions and are not algebraically equal to the hypotheses they support; for example, KLD at correction time is computed from the robot's motion distribution and the correct goal, not from the participant's condition label. The RQ2 confusion-matrix rates are conditional proportions, and the competency main effects are not forced by the 25%/75% success-rate manipulation: missed-correction rate divides by intended failures and unnecessary-correction rate divides by intended successes, so the base rates do not by themselves determine the direction of the reported differences. RQ3 is a measured Spearman correlation, not a fitted parameter renamed as a finding. The only self-citations ([23], [30]) appear in related-work descriptions of LfC assumptions and legibility-aware LfC; the present statistical results do not rest on them. The paper's own limitation in Sec. VI.D, noting that trial-level ANOVAs may inflate degrees of freedom due to within-participant correlation, is a real statistical caveat about p-values, but it is not circularity: it concerns independence of observations, not equivalence between inputs and outputs. No step reduces by construction to its own input, so no circularity steps are listed.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

This is an empirical user study with no new theoretical entities. The central claims rest on participant behavior measures and statistical assumptions. The main free parameter is the researcher-chosen precision weighting. The key questionable axiom is trial-level independence, disclosed by the authors.

free parameters (1)
  • Precision metric feature weights = 0.5 position, 0.5 rotation
    Precision is computed as an equally weighted linear combination of normalized EEF position and rotation errors to the correct goal (Appendix VIII-A). The weights are chosen by the authors, not estimated from data or theory.
assumptions (4)
  • domain assumption Each trial from each participant is treated as independent in RQ1 and RQ3 analyses.
    The paper assumes trial-level independence to use 1,944 corrections in RQ1 analyses. This is explicitly flagged in Sec VI.D as potentially inflating degrees of freedom due to within-participant correlation.
  • domain assumption Participant corrections reflect the measured constructs (task objective divergence, precision, physical effort) as operationalized.
    The study assumes the chosen metrics (KLD, precision linear combination, torque L2 norm) capture the intended psychological constructs.
  • standard math ANOVA is robust to non-normal data.
    Invoked in Sec V to justify ANOVA on non-normal measures, citing Blanca et al. 2017.
  • domain assumption Competency is operationalized as a 25% vs 75% intended success rate.
    The manipulation of competency is defined by the proportion of intended failures, which may not match a user's subjective sense of competence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Effects of Robot Competency and Motion Legibility on Human Correction Feedback." pith.science (2026). https://pith.science/paper/QPNYWP2H

@misc{pith2026250103515,
  author       = {Pith},
  title        = {Pith review of: Effects of Robot Competency and Motion Legibility on Human Correction Feedback},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QPNYWP2H}},
  note         = {Machine review of arXiv:2501.03515}
}
abstract

As robot deployments become more commonplace, people are likely to take on the role of supervising robots (i.e., correcting their mistakes) rather than directly teaching them. Prior works on Learning from Corrections (LfC) have relied on three key assumptions to interpret human feedback: (1) people correct the robot only when there is significant task objective divergence; (2) people can accurately predict if a correction is necessary; and (3) people trade off precision and physical effort when giving corrections. In this work, we study how two key factors (robot competency and motion legibility) affect how people provide correction feedback and their implications on these existing assumptions. We conduct a user study ($N=60$) under an LfC setting where participants supervise and correct a robot performing pick-and-place tasks. We find that people are more sensitive to suboptimal behavior by a highly competent robot compared to an incompetent robot when the motions are legible ($p=0.0015$) and predictable ($p=0.0055$). In addition, people also tend to withhold necessary corrections ($p < 0.0001$) when supervising an incompetent robot and are more prone to offering unnecessary ones ($p = 0.0171$) when supervising a highly competent robot. We also find that physical effort positively correlates with correction precision, providing empirical evidence to support this common assumption. We also find that this correlation is significantly weaker for an incompetent robot with legible motions than an incompetent robot with predictable motions ($p = 0.0075$). Our findings offer insights for accounting for competency and legibility when designing robot interaction behaviors and learning task objectives from corrections.

Figures

Figures reproduced from arXiv: 2501.03515 by the authors.

Figure 1
Figure 1. The robot begins moving along the yellow trajectory [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. A comparison between predictable (solid lines) and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. This figure depicts the series of pick-and-place tasks performed by the robot. The tasks involved manipulating various [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: In an intended-success trial, the robot will attempt to [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Task Objective Divergence: The KLD between the estimated robot’s task belief distribution and the actual, correct motion goal [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 10
Figure 10. Figure 10: False Omission Rate: The pro￾portion of uncorrected trials that were intended failures. between these two variables. Overall, the trend in each con￾dition exhibited a weak positive correlation (Table I in the Appendix). After adopting a Fisher z-transformation on the …
Figure 11
Figure 11. Figure 11: False Correction Rate: The proportion of corrected [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A hierarchical teaching algorithm selects complementary environments and feedback modalities to learn reward functions that generalize across unseen MDPs, proving that single-environment teaching leaves structural rew...

Reference graph

Works this paper leans on

98 extracted references · 61 canonical work pages · cited by 1 Pith paper

  1. [1]

    Robot learning from demonstration,

    C. G. Atkeson and S. Schaal, “Robot learning from demonstration,” in ICML, vol. 97, 1997, pp. 12–20

  2. [2]

    Robot learning from demonstration: a task- level planning approach,

    S. Ekvall and D. Kragic, “Robot learning from demonstration: a task- level planning approach,” International Journal of Advanced Robotic Systems, vol. 5, no. 3, p. 33, 2008

  3. [3]

    Learning from demonstration,

    S. Schaal, “Learning from demonstration,” Advances in neural informa- tion processing systems , vol. 9, 1996

  4. [4]

    A robot learning from demonstra- tion framework to perform force-based manipulation tasks,

    L. Rozo, P. Jim ´enez, and C. Torras, “A robot learning from demonstra- tion framework to perform force-based manipulation tasks,” Intelligent service robotics, vol. 6, no. 1, pp. 33–51, 2013

  5. [5]

    Gaussian-process-based robot learning from demonstration,

    M. Arduengo, A. Colom ´e, J. Lobo-Prat, L. Sentis, and C. Torras, “Gaussian-process-based robot learning from demonstration,” Journal of Ambient Intelligence and Humanized Computing , pp. 1–14, 2023

  6. [6]

    Preference-based policy learning,

    R. Akrour, M. Schoenauer, and M. Sebag, “Preference-based policy learning,” in Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greece, September 5-9, 2011. Proceedings, Part I 11 . Springer, 2011, pp. 12–27

  7. [7]

    April: Active preference learning-based reinforcement learning,

    ——, “April: Active preference learning-based reinforcement learning,” in Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II 23 . Springer, 2012, pp. 116–131

  8. [8]

    Preference learning on the execution of collaborative human-robot tasks,

    T. Munzer, M. Toussaint, and M. Lopes, “Preference learning on the execution of collaborative human-robot tasks,” in 2017 IEEE Interna- tional Conference on Robotics and Automation (ICRA) . IEEE, 2017, pp. 879–885

Show all 98 references
  1. [9]

    Improving user specifications for robot behavior through active preference learning: Framework and evaluation,

    N. Wilde, A. Blidaru, S. L. Smith, and D. Kuli ´c, “Improving user specifications for robot behavior through active preference learning: Framework and evaluation,” The International Journal of Robotics Research, vol. 39, no. 6, pp. 651–667, 2020

  2. [10]

    Skill preferences: Learning to extract and execute robotic skills from human feedback,

    X. Wang, K. Lee, K. Hakhamaneshi, P. Abbeel, and M. Laskin, “Skill preferences: Learning to extract and execute robotic skills from human feedback,” in Conference on Robot Learning . PMLR, 2022, pp. 1259– 1268

  3. [11]

    Few-shot preference learning for human- in-the-loop rl,

    D. J. Hejna III and D. Sadigh, “Few-shot preference learning for human- in-the-loop rl,” in Conference on Robot Learning . PMLR, 2023, pp. 2014–2025

  4. [12]

    Policy shaping with su- pervisory attention driven exploration,

    T. K. Faulkner, E. S. Short, and A. L. Thomaz, “Policy shaping with su- pervisory attention driven exploration,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 842–847

  5. [13]

    Modeling supervisor safe sets for improving collaboration in human- robot teams,

    D. L. McPherson, D. R. Scobee, J. Menke, A. Y . Yang, and S. S. Sastry, “Modeling supervisor safe sets for improving collaboration in human- robot teams,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 861–868

  6. [14]

    Human supervision of robotic site surveys,

    D. Schreckenghost, T. Fong, and T. Milam, “Human supervision of robotic site surveys,” in AIP Conference Proceedings , vol. 969, no. 1. American Institute of Physics, 2008, pp. 776–783

  7. [15]

    Providing the basis for human-robot-interaction: A multi- modal attention system for a mobile robot,

    S. Lang, M. Kleinehagenbrock, S. Hohenner, J. Fritsch, G. A. Fink, and G. Sagerer, “Providing the basis for human-robot-interaction: A multi- modal attention system for a mobile robot,” in Proceedings of the 5th international conference on Multimodal interfaces , 2003, pp. 28–35

  8. [16]

    Learning robot objectives from physical human interaction,

    A. Bajcsy, D. P. Losey, M. K. O’malley, and A. D. Dragan, “Learning robot objectives from physical human interaction,” in Conference on Robot Learning. PMLR, 2017, pp. 217–226

  9. [17]

    Learning from physical human corrections, one feature at a time,

    A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan, “Learning from physical human corrections, one feature at a time,” in Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction, 2018, pp. 141–149

  10. [18]

    Including uncertainty when learning from human corrections,

    D. P. Losey and M. K. O’Malley, “Including uncertainty when learning from human corrections,” in Conference on Robot Learning . PMLR, 2018, pp. 123–132

  11. [19]

    Less is more: Rethinking probabilistic models of human behavior,

    A. Bobu, D. R. Scobee, J. F. Fisac, S. S. Sastry, and A. D. Dragan, “Less is more: Rethinking probabilistic models of human behavior,” in Proceedings of the 2020 acm/ieee international conference on human- robot interaction, 2020, pp. 429–437

  12. [20]

    Quantifying hypothesis space misspecification in learning from human– robot demonstrations and physical corrections,

    A. Bobu, A. Bajcsy, J. F. Fisac, S. Deglurkar, and A. D. Dragan, “Quantifying hypothesis space misspecification in learning from human– robot demonstrations and physical corrections,” IEEE Transactions on Robotics, vol. 36, no. 3, pp. 835–854, 2020

  13. [21]

    Physical interaction as communication: Learning robot objectives online from human corrections,

    D. P. Losey, A. Bajcsy, M. K. O’Malley, and A. D. Dragan, “Physical interaction as communication: Learning robot objectives online from human corrections,” The International Journal of Robotics Research , vol. 41, no. 1, pp. 20–44, 2022

  14. [22]

    Correct me if i am wrong: Interactive learning for robotic manipula- tion,

    E. Chisari, T. Welschehold, J. Boedecker, W. Burgard, and A. Valada, “Correct me if i am wrong: Interactive learning for robotic manipula- tion,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 3695– 3702, 2022

  15. [23]

    Legibility-aware learning from corrections,

    A. Wang and T. Fitzgerald, “Legibility-aware learning from corrections,” 2024

  16. [24]

    Modeling and learning con- straints for creative tool use,

    T. Fitzgerald, A. Goel, and A. Thomaz, “Modeling and learning con- straints for creative tool use,” Frontiers in Robotics and AI , vol. 8, p. 674292, 2021

  17. [25]

    Human-guided trajectory adaptation for tool transfer,

    T. Fitzgerald, E. Short, A. Goel, and A. Thomaz, “Human-guided trajectory adaptation for tool transfer,” in Proceedings of the 18th Inter- national Conference on Autonomous Agents and MultiAgent Systems, ser. AAMAS ’19. Richland, SC: International Foundation for Autonomous Age...

  18. [26]

    Robot motion adap- tation through user intervention and reinforcement learning,

    A. Jevti ´c, A. Colom ´e, G. Alenya, and C. Torras, “Robot motion adap- tation through user intervention and reinforcement learning,” Pattern Recognition Letters, vol. 105, pp. 67–75, 2018

  19. [27]

    Learning from interventions,

    J. Spencer, S. Choudhury, M. Barnes, M. Schmittle, M. Chiang, P. Ra- madge, and S. Srinivasa, “Learning from interventions,” in Robotics: Science and Systems (RSS) , 2020

  20. [28]

    Expert intervention learning: An online framework for robot learning from explicit and implicit human feedback,

    ——, “Expert intervention learning: An online framework for robot learning from explicit and implicit human feedback,” Autonomous Robots, pp. 1–15, 2022

  21. [29]

    Interactive robot learning from verbal correction,

    H. Liu, A. Chen, Y . Zhu, A. Swaminathan, A. Kolobov, and C.- A. Cheng, “Interactive robot learning from verbal correction,” arXiv preprint arXiv:2310.17555, 2023

  22. [30]

    Toward measuring the effect of robot competency on human kinesthetic feedback in long-term task learning,

    S. Wang, B. Scassellati, and T. Fitzgerald, “Toward measuring the effect of robot competency on human kinesthetic feedback in long-term task learning,” 4th Workshop on Lifelong Learning and Personalization in Long-Term Human-Robot Interaction (LEAP-HRI) , 2024

  23. [31]

    Human supervisory control of robot systems,

    T. Sheridan, “Human supervisory control of robot systems,” in Proceed- ings. 1986 IEEE International Conference on Robotics and Automation , vol. 3. IEEE, 1986, pp. 808–812

  24. [32]

    A study on dense and sparse (visual) rewards in robot policy learning,

    A. Mohtasib, G. Neumann, and H. Cuay ´ahuitl, “A study on dense and sparse (visual) rewards in robot policy learning,” inTowards Autonomous Robotic Systems: 22nd Annual Conference, TAROS 2021, Lincoln, UK, September 8–10, 2021, Proceedings 22 . Springer, 2021, pp. 3–13

  25. [33]

    From real-time attention assessment to “with-me-ness

    S. Lemaignan, F. Garcia, A. Jacq, and P. Dillenbourg, “From real-time attention assessment to “with-me-ness” in human-robot interaction,” in 2016 11th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 2016, pp. 157–164

  26. [34]

    Robot navigation in crowds by graph convolutional networks with attention learned from human gaze,

    Y . Chen, C. Liu, B. E. Shi, and M. Liu, “Robot navigation in crowds by graph convolutional networks with attention learned from human gaze,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 2754–2761, 2020

  27. [35]

    Imitation learning with in- consistent demonstrations through uncertainty-based data manipulation,

    P. Valletta, R. P ´erez-Dattari, and J. Kober, “Imitation learning with in- consistent demonstrations through uncertainty-based data manipulation,” in 2021 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2021, pp. 3655–3661

  28. [36]

    Robot learning from human demonstrations with inconsistent contexts,

    Z. Qian, M. You, H. Zhou, X. Xu, and B. He, “Robot learning from human demonstrations with inconsistent contexts,” Robotics and Autonomous Systems, vol. 166, p. 104466, 2023

  29. [37]

    Interactively shaping agents via human reinforcement: The tamer framework,

    W. B. Knox and P. Stone, “Interactively shaping agents via human reinforcement: The tamer framework,” in Proceedings of the fifth in- ternational conference on Knowledge capture , 2009, pp. 9–16

  30. [38]

    Active reward learning

    C. Daniel, M. Viering, J. Metz, O. Kroemer, and J. Peters, “Active reward learning.” in Robotics: Science and systems , vol. 98, 2014

  31. [39]

    Chernova and A

    S. Chernova and A. L. Thomaz, Robot learning from human teachers . Morgan & Claypool Publishers, 2014

  32. [40]

    The empathic framework for task learning from implicit human feedback,

    Y . Cui, Q. Zhang, B. Knox, A. Allievi, P. Stone, and S. Niekum, “The empathic framework for task learning from implicit human feedback,” in Conference on Robot Learning . PMLR, 2021, pp. 604–626

  33. [41]

    Learning something from nothing: Leveraging implicit human feedback strategies,

    R. Loftin, B. Peng, J. MacGlashan, M. L. Littman, M. E. Taylor, J. Huang, and D. L. Roberts, “Learning something from nothing: Leveraging implicit human feedback strategies,” in The 23rd IEEE in- ternational symposium on robot and human interactive communication . IEEE, 2014, ...

  34. [42]

    Learning behaviors via human-delivered discrete feedback: mod- eling implicit feedback strategies to speed up learning,

    ——, “Learning behaviors via human-delivered discrete feedback: mod- eling implicit feedback strategies to speed up learning,” Autonomous agents and multi-agent systems , vol. 30, pp. 30–59, 2016

  35. [43]

    Personal robot training via natural-language instructions,

    S. Lauria, G. Bugmann, T. Kyriacou, J. Bos, and E. Klein, “Personal robot training via natural-language instructions,” IEEE Intelligent sys- tems, vol. 16, no. 3, pp. 38–45, 2001

  36. [44]

    Learning to parse natural language commands to a robot control system,

    C. Matuszek, E. Herbst, L. Zettlemoyer, and D. Fox, “Learning to parse natural language commands to a robot control system,” in Experimental robotics: the 13th international symposium on experimental robotics . Springer, 2013, pp. 403–415

  37. [45]

    Learning to interpret natural language commands through human-robot dialog,

    J. Thomason, S. Zhang, R. J. Mooney, and P. Stone, “Learning to interpret natural language commands through human-robot dialog,” in Twenty-Fourth International Joint Conference on Artificial Intelligence , 2015

  38. [46]

    Language-conditioned imitation learning for robot ma- nipulation tasks,

    S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor, “Language-conditioned imitation learning for robot ma- nipulation tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 13 139–13 150, 2020

  39. [47]

    Understanding the relationship between interactions and outcomes in human-in-the-loop machine learning,

    Y . Cui, P. Koppol, H. Admoni, S. Niekum, R. Simmons, A. Steinfeld, and T. Fitzgerald, “Understanding the relationship between interactions and outcomes in human-in-the-loop machine learning,” in International Joint Conference on Artificial Intelligence , 2021

  40. [48]

    Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance,

    A. L. Thomaz, C. Breazeal et al., “Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance,” in Aaai, vol. 6. Boston, MA, 2006, pp. 1000– 1005

  41. [49]

    Teachable robots: Understanding human teaching behavior to build more effective robot learners,

    A. L. Thomaz and C. Breazeal, “Teachable robots: Understanding human teaching behavior to build more effective robot learners,” Artificial Intelligence, vol. 172, no. 6-7, pp. 716–737, 2008

  42. [50]

    Improving robot mo- tor learning with negatively valenced reinforcement signals,

    N. Navarro-Guerrero, R. J. Lowe, and S. Wermter, “Improving robot mo- tor learning with negatively valenced reinforcement signals,” Frontiers in neurorobotics, vol. 11, p. 10, 2017

  43. [51]

    Inquire: Interactive querying for user- aware informative reasoning,

    T. Fitzgerald, P. Koppol, P. Callaghan, R. Q. J. H. Wong, R. Simmons, O. Kroemer, and H. Admoni, “Inquire: Interactive querying for user- aware informative reasoning,” in 6th Annual Conference on Robot Learning, 2022

  44. [52]

    Bayesian inverse reinforcement learn- ing

    D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learn- ing.” in IJCAI, vol. 7, 2007, pp. 2586–2591

  45. [53]

    Joint estimation of expertise and reward preferences from human demonstrations,

    P. Carreno-Medrano, S. L. Smith, and D. Kuli ´c, “Joint estimation of expertise and reward preferences from human demonstrations,” IEEE Transactions on Robotics , vol. 39, no. 1, pp. 681–698, 2022

  46. [54]

    Maximum entropy inverse reinforcement learning

    B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey et al., “Maximum entropy inverse reinforcement learning.” in Aaai, vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438

  47. [55]

    The effects of meaningful and meaningless explanations on trust and perceived system accuracy in intelligent systems,

    M. Nourani, S. Kabir, S. Mohseni, and E. D. Ragan, “The effects of meaningful and meaningless explanations on trust and perceived system accuracy in intelligent systems,” in Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , vol. 7, 2019, pp. 97–105

  48. [56]

    Detection and elimination of systematic labeling bias in code reviewer recommen- dation systems,

    K. A. Tecimer, E. T ¨uz¨un, H. Dibeklioglu, and H. Erdogmus, “Detection and elimination of systematic labeling bias in code reviewer recommen- dation systems,” in Proceedings of the 25th International Conference on Evaluation and Assessment in Software Engineering, 2021, pp. 181–190

  49. [57]

    Robot learning in homes: Improving generalization and reducing dataset bias,

    A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto, “Robot learning in homes: Improving generalization and reducing dataset bias,” Advances in neural information processing systems , vol. 31, 2018

  50. [58]

    Learning with noisy labels revisited: A study using real-world human annotations,

    J. Wei, Z. Zhu, H. Cheng, T. Liu, G. Niu, and Y . Liu, “Learning with noisy labels revisited: A study using real-world human annotations,” arXiv preprint arXiv:2110.12088 , 2021

  51. [59]

    Guided cost learning: Deep inverse optimal control via policy optimization,

    C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning. PMLR, 2016, pp. 49–58

  52. [60]

    Infinite time horizon maximum causal entropy inverse reinforcement learning,

    M. Bloem and N. Bambos, “Infinite time horizon maximum causal entropy inverse reinforcement learning,” in 53rd IEEE conference on decision and control . IEEE, 2014, pp. 4911–4916

  53. [61]

    Legibility and predictabil- ity of robot motion,

    A. D. Dragan, K. C. Lee, and S. S. Srinivasa, “Legibility and predictabil- ity of robot motion,” in 2013 8th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 2013, pp. 301–308

  54. [62]

    Sim-to-real transfer in deep reinforcement learning for robotics: a survey,

    W. Zhao, J. P. Queralta, and T. Westerlund, “Sim-to-real transfer in deep reinforcement learning for robotics: a survey,” in 2020 IEEE symposium series on computational intelligence (SSCI) . IEEE, 2020, pp. 737–744

  55. [63]

    A study on challenges of testing robotic systems,

    A. Afzal, C. Le Goues, M. Hilton, and C. S. Timperley, “A study on challenges of testing robotic systems,” in 2020 IEEE 13th International Conference on Software Testing, Validation and Verification (ICST) . IEEE, 2020, pp. 96–107

  56. [64]

    Tiny robot learning: Challenges and directions for machine learning in resource-constrained robots,

    S. M. Neuman, B. Plancher, B. P. Duisterhof, S. Krishnan, C. Banbury, M. Mazumder, S. Prakash, J. Jabbour, A. Faust, G. C. de Croon et al., “Tiny robot learning: Challenges and directions for machine learning in resource-constrained robots,” in 2022 IEEE 4th International Conf...

  57. [65]

    A meta-analysis of factors affecting trust in human-robot interaction,

    P. A. Hancock, D. R. Billings, K. E. Schaefer, J. Y . Chen, E. J. De Visser, and R. Parasuraman, “A meta-analysis of factors affecting trust in human-robot interaction,” Human factors , vol. 53, no. 5, pp. 517–527, 2011

  58. [66]

    Judging a bot by its cover: An experiment on expectation setting for personal robots,

    S. Paepcke and L. Takayama, “Judging a bot by its cover: An experiment on expectation setting for personal robots,” in 2010 5th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 2010, pp. 45–52

  59. [67]

    Warmth and competence to predict human preference of robot behavior in physical human-robot interaction,

    M. M. Scheunemann, R. H. Cuijpers, and C. Salge, “Warmth and competence to predict human preference of robot behavior in physical human-robot interaction,” in 2020 29th IEEE international conference on robot and human interactive communication (RO-MAN) . IEEE, 2020, pp. 1340–1347

  60. [68]

    Helping robots learn: a human-robot master-apprentice model using demonstrations via virtual reality teleoperation,

    J. DelPreto, J. I. Lipton, L. Sanneman, A. J. Fay, C. Fourie, C. Choi, and D. Rus, “Helping robots learn: a human-robot master-apprentice model using demonstrations via virtual reality teleoperation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IE...

  61. [69]

    The effects of a robot’s performance on human teachers for learning from demonstration tasks,

    E. Hedlund, M. Johnson, and M. Gombolay, “The effects of a robot’s performance on human teachers for learning from demonstration tasks,” in Proceedings of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, 2021, pp. 207–215

  62. [70]

    Generating legible motion,

    A. Dragan and S. Srinivasa, “Generating legible motion,” 2013

  63. [71]

    Facilitating intention prediction for humans by optimizing robot motions,

    F. Stulp, J. Grizou, B. Busch, and M. Lopes, “Facilitating intention prediction for humans by optimizing robot motions,” in 2015 ieee/rsj international conference on intelligent robots and systems (iros). IEEE, 2015, pp. 1249–1255

  64. [72]

    Sadigh, A

    D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia, Active preference- based learning of reward functions , 2017

  65. [73]

    Active uncertainty reduction for human- robot interaction: An implicit dual control approach,

    H. Hu and J. F. Fisac, “Active uncertainty reduction for human- robot interaction: An implicit dual control approach,” in International Workshop on the Algorithmic Foundations of Robotics. Springer, 2022, pp. 385–401

  66. [74]

    Active probing and influencing human behaviors via autonomous agents,

    S. Wang, Y . Lyu, and J. M. Dolan, “Active probing and influencing human behaviors via autonomous agents,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 1514–1521

  67. [75]

    Influence of legibility on perceived safety in a virtual human-robot path crossing task,

    C. Lichtenth ¨aler, T. Lorenzy, and A. Kirsch, “Influence of legibility on perceived safety in a virtual human-robot path crossing task,” in 2012 IEEE RO-MAN: The 21st IEEE International Symposium on Robot and Human Interactive Communication . IEEE, 2012, pp. 676–681

  68. [76]

    Hey robot, which way are you going? nonverbal motion legibility cues for human- robot spatial interaction,

    N. J. Hetherington, E. A. Croft, and H. M. Van der Loos, “Hey robot, which way are you going? nonverbal motion legibility cues for human- robot spatial interaction,” IEEE Robotics and Automation Letters, vol. 6, no. 3, pp. 5010–5015, 2021

  69. [77]

    Effects of robot motion on human-robot collaboration,

    A. D. Dragan, S. Bauman, J. Forlizzi, and S. S. Srinivasa, “Effects of robot motion on human-robot collaboration,” in Proceedings of the tenth annual ACM/IEEE international conference on human-robot interaction, 2015, pp. 51–58

  70. [78]

    A review on robot motion legibility for human- robot interaction,

    S. Kim and W. Park, “A review on robot motion legibility for human- robot interaction,” Korean Society of Ergonomics Conference , pp. 220– 228, 2017

  71. [79]

    Learning legible mo- tion from human–robot interactions,

    B. Busch, J. Grizou, M. Lopes, and F. Stulp, “Learning legible mo- tion from human–robot interactions,” International Journal of Social Robotics, vol. 9, no. 5, pp. 765–779, 2017

  72. [80]

    The Open Motion Planning Library,

    I. A. S ¸ucan, M. Moll, and L. E. Kavraki, “The Open Motion Planning Library,” IEEE Robotics & Automation Magazine , vol. 19, no. 4, pp. 72–82, December 2012, https://ompl.kavrakilab.org

  73. [81]

    Model-free friction observers for flexible joint robots with torque measurements,

    M. J. Kim, F. Beck, C. Ott, and A. Albu-Sch ¨affer, “Model-free friction observers for flexible joint robots with torque measurements,” IEEE Transactions on Robotics , vol. 35, no. 6, pp. 1508–1515, 2019

  74. [82]

    Feel the bite: Robot-assisted inside-mouth bite transfer using robust mouth perception and physical interaction-aware control,

    R. K. Jenamani, D. Stabile, Z. Liu, A. Anwar, K. Dimitropoulou, and T. Bhattacharjee, “Feel the bite: Robot-assisted inside-mouth bite transfer using robust mouth perception and physical interaction-aware control,” in Proceedings of the 2024 ACM/IEEE International Confer- ence...

  75. [83]

    Analytical derivatives of rigid body dynamics algorithms,

    J. Carpentier and N. Mansard, “Analytical derivatives of rigid body dynamics algorithms,” in Robotics: Science and Systems (RSS) , 2018

  76. [84]

    Sus: A quick and dirty usability scale,

    J. Brooke, “Sus: A quick and dirty usability scale,” Usability Evaluation in INdustry/Taylor and Francis , 1996

  77. [85]

    The perception of agency: Scale reduction and construct validity,

    J. G. Trafton, C. R. Frazier, K. Zish, B. J. Bio, and J. M. McCurry, “The perception of agency: Scale reduction and construct validity,” in 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, 2023, pp. 936–942

  78. [86]

    Nasa-task load index (nasa-tlx); 20 years later,

    S. G. Hart, “Nasa-task load index (nasa-tlx); 20 years later,” in Pro- ceedings of the human factors and ergonomics society annual meeting , vol. 50, no. 9. Sage publications Sage CA: Los Angeles, CA, 2006, pp. 904–908

  79. [87]

    Theoretical considerations and development of a ques- tionnaire to measure trust in automation,

    M. K ¨orber, “Theoretical considerations and development of a ques- tionnaire to measure trust in automation,” in Proceedings of the 20th Congress of the International Ergonomics Association (IEA 2018) Vol- ume VI: Transport Ergonomics and Human Factors (TEHF), Aerospace Human...

  80. [88]

    On information and sufficiency,

    S. Kullback and R. A. Leibler, “On information and sufficiency,” The annals of mathematical statistics , vol. 22, no. 1, pp. 79–86, 1951

  81. [89]

    The boltzmann policy distribution: Ac- counting for systematic suboptimality in human models,

    C. Laidlaw and A. Dragan, “The boltzmann policy distribution: Ac- counting for systematic suboptimality in human models,” arXiv preprint arXiv:2204.10759, 2022

  82. [90]

    Selecting and interpreting measures of thematic classi- fication accuracy,

    S. V . Stehman, “Selecting and interpreting measures of thematic classi- fication accuracy,” Remote sensing of Environment , vol. 62, no. 1, pp. 77–89, 1997

  83. [91]

    Statistical methods for research workers,

    R. A. Fisher, “Statistical methods for research workers,” in Break- throughs in statistics: Methodology and distribution . Springer, 1970, pp. 66–70

  84. [92]

    Non-normal data: Is anova still a valid option?

    M. J. Blanca Mena, R. Alarc ´on Postigo, J. Arnau Gras, R. Bono Cabr ´e, and R. Bendayan, “Non-normal data: Is anova still a valid option?” Psicothema, 2017, vol. 29, num. 4, p. 552-557 , 2017

  85. [93]

    Controlling the false discovery rate: a practical and powerful approach to multiple testing,

    Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” Journal of the Royal statistical society: series B (Methodological) , vol. 57, no. 1, pp. 289–300, 1995

  86. [94]

    An analysis of variance test for normality (complete samples),

    S. S. Shapiro and M. B. Wilk, “An analysis of variance test for normality (complete samples),” Biometrika, vol. 52, no. 3-4, pp. 591–611, 1965

  87. [95]

    The proof and measurement of association between two things

    C. Spearman, “The proof and measurement of association between two things.” 1961

  88. [96]

    Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population,

    R. A. Fisher, “Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population,” Biometrika, vol. 10, no. 4, pp. 507–521, 1915

  89. [97]

    014: On the

    R. A. Fisher et al. , “014: On the” probable error” of a coefficient of correlation deduced from a small sample.” 1921

  90. [98]

    Examining the effects of robots’ physical appearance, warmth, and competence in frontline services: The humanness-value-loyalty model,

    D. Belanche, L. V . Casal ´o, J. Schepers, and C. Flavi ´an, “Examining the effects of robots’ physical appearance, warmth, and competence in frontline services: The humanness-value-loyalty model,” Psychology & Marketing, vol. 38, no. 12, pp. 2357–2376, 2021. VIII. A PPENDIX A...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.