{"id":"80e786e6-1c11-48e4-9328-130ceffdc82f","arxiv_id":"2505.10151","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Providing novice teachers with machine-teaching-derived feedback during a scaffolding curriculum improved their reward demonstrations for a robot trained with RLfD, with mixed evidence of transfer to unseen skills.","lead":"Novice teachers who received machine-teaching-derived feedback during a short scaffolded practice session gave notably better reward demonstrations to a robot learner on the trained task, and their reward-assignment quality also improved on an unseen task. The abstract claims the unseen-task robot learning improved by 70%, but the statistical test for that improvement was not significant.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central transfer claim is not supported by the reported inferential statistics: on unseen skill S2, only E_ADE is significant, while the direct robot-learning metrics E_ARMSE (p=0.27) and E_ATR (p=0.16) are not, so the abstract's '70% improvement' overstates the evidence.","rationale":"The reader's verdict is CONDITIONAL, and my read agrees with that assessment. The strongest, most central claim is the transfer result (h2): MT-guidance improves robot learning performance on unseen skills. The paper's own results show that the direct robot-learning metrics for S2 are not statistically significant, while only demonstration-error E_ADE is. The abstract's phrasing that MT-guidance 'causes a 70% improvement in robot learning performance' leans on a non-significant E_ARMSE comparison. That is an overstatement of the evidence, though not evidence of fraud or even of a false effect; it is a mismatch between the reported statistics and the headline. The reader's weakest_assumption focused on Eq. (20), but I see the statistical mismatch as more directly load-bearing because the paper measures robot learning outcomes directly; even if Eq. (20) were tightened, the non-significant E_ARMSE and E_ATR would still undermine the strongest claim. I nevertheless mark agreement as 'partial' because the theoretical link, with its self-cited proof and proportionality-from-bound issue, remains a real secondary concern. A CONDITIONAL verdict is appropriate: the training-skill evidence (h1) is solid, and the transfer idea is plausible, but the transfer claim and abstract need revision, and a more careful statistical analysis should accompany any re-submission. Rejecting outright would be too harsh given the strong within-task results and the reasonable experimental design; accepting as-is would ignore the unsupported headline.","tokens_in":8163,"tokens_out":2807,"duration_ms":28628,"concrete_test":"Re-analyse the S2 transfer data (P2 vs P8) with paired bootstrap or a robust Wilcoxon test on E_ARMSE and E_ATR, reporting 95% confidence intervals and standardized effect sizes rather than only p-values; if the intervals include zero or the effect sizes are small, the abstract's '70% improvement in robot learning performance' should be removed or explicitly qualified as non-significant.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is a transfer effect on robot learning performance for unseen skills. The evidence for that claim rests on the S2 comparison (P2 vs P8). There, only demonstration error E_ADE reaches significance (64% reduction, p=4.73e-9); the two direct measures of robot learning, E_ARMSE (70% reduction, p=0.27) and E_ATR (91% reduction, p=0.16), do not. The abstract's '70% improvement in robot learning performance' is therefore not supported by the reported statistics. The authors attribute the null to 'reward ambiguity,' but that explanation is not itself tested and does not transform non-significant learning-outcome differences into evidence of transfer. The theoretical bridge claimed in Eq. (20), ∥ω−ω̄∥∝∥j−j̄∥, is likewise only an upper-bound argument with a condition-number-dependent constant, and its proof is delegated to self-cited prior work [4], so it cannot carry the load of converting the significant E_ADE result into a claim about learned policies. The central claim would hold only if robot learning metrics improved significantly or if a validated mapping from E_ADE to policy performance were established; neither condition is met in the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a machine-teaching-based scaffolding framework for training novice humans to provide reward demonstrations for reinforcement learning from demonstration (RLfD). In a between-subjects experiment, participants teach a two-link robot arm reaching skills S1 (point reaching) and S2 (line reaching); guided participants receive visual feedback based on the ideal reward values in a five-phase curriculum. The paper reports large within-task improvements in demonstration error and robot learning metrics on S1 and claims transfer to S2 with a 64% reduction in demonstration error and a 70% improvement in robot learning. The abstract and introduction present the 70% transfer improvement on robot learning as a headline result.","tokens_in":8463,"tokens_out":10686,"duration_ms":98393,"significance":"If the within-task results hold, the paper demonstrates that a short scaffolded curriculum with visual reward feedback can substantially improve novices' reward giving in RLfD, and the comparison with supervised LfD is a useful additional data point. The experimental design has real strengths: random assignment, a control group, a power analysis, multiple evaluation metrics, and a separate transfer task. The paper does not provide code or machine-checked proofs, so all claims rest on the reported statistics, which makes the mismatch between the significant E_ADE transfer result and the non-significant robot-learning metrics on S2 particularly damaging for the abstract's central claim. The transfer claim as written is not supported by the reported tests.","major_comments":[{"comment":"The central transfer claim is not supported by the reported inferential statistics. For unseen skill S2, comparing P2 and P8 gives a significant 64% reduction in E_ADE (p=4.73e-9), but the direct measures of robot learning, E_ARMSE (p=0.27) and E_ATR (p=0.16), are not statistically significant. The abstract's statement that MT-guidance 'causes a 70% improvement in robot learning performance on skills not seen by subjects during training' therefore overstates the evidence. The post-hoc attribution of this discrepancy to 'reward ambiguity' is not tested with any quantitative measure and cannot convert non-significant learning-outcome differences into evidence of transfer. Please either report the transfer claim only for demonstration quality (E_ADE), or supply an analysis that directly supports the robot-learning transfer claim.","section":"§IV-C (Transfer of Teaching Ability); Abstract"},{"comment":"Equation (20) asserts ∥ω−ω̄∥2 ∝ ∥j−j̄∥2 and attributes the proof to the sub-multiplicative property of matrix norms, citing prior work [4]. Sub-multiplicativity yields at most an upper bound of the form ∥ω−ω̄∥ ≤ ∥(Ψ⊤(Ψ−γΨ′))−1∥ ∥Ψ⊤∥ ∥j−j̄∥, with a condition-number-dependent constant; it does not establish proportionality, and no lower bound is given. In addition, the reported metric E_ADE is an ℓ1-style sum over individual reward differences, whereas Eq. (20) is stated for an ℓ2 norm. This relation is the theoretical bridge from the significant E_ADE result to claims about learned policies, so it cannot be left as a self-cited 'it can be shown' step; the paper should either prove a precise two-sided bound or weaken the interpretation of E_ADE as a proxy for robot learning.","section":"§III-B, Eq. (20)"},{"comment":"The h1 analysis compares within-group changes only: P1 versus P9 for the target group and P1 versus P9 for the control group. A significant within-group improvement in one group and a non-significant change in the other does not by itself establish a treatment effect. The paper should report a between-group comparison of change scores or a group-by-time interaction test for E_ADE, E_ARMSE, and E_ATR to support the causal claim that guidance, rather than practice or time, produced the improvement.","section":"§IV-B/C (between-subjects analysis)"},{"comment":"The control-group results are used to infer 'the absence of training effects' from p>0.1 (S1) and p>0.05 (S2). Failure to reject the null in a study with ten participants per group is not evidence of absence. Please report effect sizes and confidence intervals for the control-group comparisons and avoid framing non-significant p-values as demonstrating that no effect exists.","section":"§IV-C (Transfer of Teaching Ability)"}],"minor_comments":[{"comment":"The parameter definitions are confusing: Eq. (21) sets R=βI, while Eq. (22) also carries a β factor, and the text says Q=R for S1 yet Eq. (22) defines a different Q for S2. Please spell out the exact matrices used for each skill and report the sensitivity of the results to the chosen β and ε.","section":"§IV-A, Eq. (21)-(22)"},{"comment":"P1 and P9 use the same skill S1 with newly sampled random probes, so h1 is a within-task improvement rather than evidence of generalization to new reward functions; h2 is the true generalization test. The phrase 'generalise this to previously unseen ones' in the abstract should be aligned with this distinction.","section":"§II-B / §IV-B"},{"comment":"The supervised-LfD versus RLfD comparison appears to be descriptive only; no statistical tests or error bars are reported for the threshold crossing times, so the crossover points t=276 and t=69 should be labeled as exploratory.","section":"§IV-C, Fig. 6"},{"comment":"The definition of E_ADE in metric m1 starts the sum at n=0; since demonstrations are indexed n=1,...,N=8, the sum should run from n=1 to N unless a separate term is intended.","section":"§IV-A (metric m1)"},{"comment":"The paper does not state whether ethics approval or informed consent was obtained for the human-subject experiment; this should be reported in the experimental protocol.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The current version is not acceptable because the headline transfer claim is contradicted by the reported p-values. I would encourage the authors to re-frame the central claim around the significant transfer of demonstration quality (E_ADE) and the within-task robot-learning improvements, and to add the missing between-group interaction tests. If the data allow, reporting confidence intervals for the non-significant S2 learning metrics would help readers calibrate the strength of the transfer evidence. The supervised-LfD comparison is interesting but needs statistical support before it can carry weight in the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper does something new and useful: it adapts the authors' earlier MT-guided teacher training from supervised LfD to RLfD with LSPI, adds a scaffolding curriculum, and tests whether the training transfers to an unseen skill. Second, the paper is honest in its results section about the main weakness, but the abstract overstates it: the within-task training effect is solid, while the headline transfer effect is not statistically significant.\n\nWhat the paper does well: the S1 within-task results are strong and consistent. The target group improves E_ADE by 83% (p = 9.78e-14), E_ARMSE by 89% (p = 0.013), and E_ATR by 98% (p = 0.006), while the control group shows no significant change. That is a clean, credible result with a random assignment and a between-subjects design. The S2 transfer test is a reasonable design, and the authors report all metrics, including the non-significant ones. That transparency deserves credit.\n\nThe soft spots are real. The abstract claims a 70% improvement in robot learning performance on unseen skills, but that is the E_ARMSE reduction on S2, and it is not significant (p = 0.27); E_ATR is also not significant (p = 0.16). The only significant S2 result is E_ADE (64% reduction, p = 4.73e-9), which measures demonstration quality, not robot learning performance. The paper attributes the discrepancy to \"reward ambiguity,\" which is plausible but untested. Eq. (20) is also weaker than presented: sub-multiplicativity of matrix norms gives an upper bound with a condition-number-dependent constant, not the claimed proportionality, and the proof is delegated to a self-cited prior work. That matters because the interpretation of reduced E_ADE as directly improving robot policy learning rests on that equation.\n\nMinor issues: no code or data are provided, and the statistics lack effect sizes and confidence intervals. The sample size is small but justified by a power analysis. The supervised-vs-RLfD comparison is a nice extra but not central.\n\nBottom line: this is a useful empirical study for people working on human-in-the-loop LfD or interactive RL. It deserves a serious referee after the abstract and Eq. (20) are fixed. The core training result is worth evaluating; the transfer claim is not yet established.","headline":"Useful extension with solid within-task training results, but the headline transfer claim to unseen skills is not supported by its own statistics.","tokens_in":8992,"tokens_out":1878,"would_cite":false,"duration_ms":18135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-teaching guidance sharpens novice reward demonstrations for reinforcement-learning robots and transfers to skills the teachers never trained on.","keywords":["reinforcement learning from demonstration","machine teaching","learning from demonstration","scaffolding training","reward design","transfer of teaching skill","least-squares policy iteration","human-robot interaction"],"falsifier":"Compute the ratio $\\|\\omega-\\bar{\\omega}\\|_2 / \\|j-\\bar{\\jmath}\\|_2$ across many random state-action samples in the least-squares policy iteration setup; if it varies by orders of magnitude, the proportionality in Eq. (20) fails and reward-error gains would not reliably become robot-learning gains. A companion check: rerun the transfer experiment with an unseen skill that has a unique optimal policy; if trajectory-level gains still fail to reach significance, the reward-ambiguity explanation is not sufficient.","tokens_in":7896,"feed_emoji":"🤖","tokens_out":12961,"duration_ms":112618,"temperature":0.7,"pith_summary":"This paper asks whether machine teaching—the practice of choosing the smallest set of examples that lets a learner reach a target—can train novice humans to give better reward demonstrations to robots that learn by reinforcement learning from demonstration (RLfD). The authors build a scaffolding interface that shows trainees the ideal reward for each state-action pair while they adjust a reward slider, using only eight demonstrations for a simple reaching skill. Their central claim is that this guidance improves the robot's learning on the training skill and transfers: trainees also produce better demonstrations for a second skill they never practiced. On the training skill, the target group's reward error fell 83% and the average root-mean-square trajectory error fell 89%; on the unseen skill, reward error fell 64%, while trajectory-level gains were large but not statistically significant, which the paper attributes to reward ambiguity in that task.","feed_headline":"Eight guided demos sharpen robot teaching and transfer to new skills","feed_subtitle":"Eight guided demonstrations cut reward error by 83% on trained skills and 64% on new ones, improving robot learning.","key_machinery":"The machinery is an interactive scaffolding training interface built on a machine-teaching formulation for RLfD. For a linear action-value model $Q^{\\pi}(x,u)=\\omega^{\\top}\\psi(x,u)$ learned by least-squares policy iteration, the teaching risk is $\\rho=\\|\\omega-\\bar{\\omega}\\|_2$, and with teaching dimension $T$ the teaching budget is exactly $N=T=8$ state-action-reward tuples. The interface shows the current state and action, lets the user assign a reward with a slider, and displays the ideal reward as a reference bar. The training curriculum P3-P7 progressively fixes state, action direction, and action magnitude so novices internalise the reward structure. The argument that this improves robot learning passes through Eq. (20), $\\|\\omega-\\bar{\\omega}\\|_2\\propto\\|j-\\bar{\\jmath}\\|_2$, which connects the measurable reward error to the value-parameter error that determines policy quality.","core_discovery":"The central claim is that a teacher's absolute reward error, $\\|j-\\bar{\\jmath}\\|$, can serve as a training signal for novice humans: with a least-squares policy iteration learner and linear value features, the value-parameter error $\\|\\omega-\\bar{\\omega}\\|$ is taken to be proportional to that reward error, so guiding people toward ideal rewards should improve what the robot learns. The experiment supports this on the trained skill: eight machine-teaching-guided demonstrations cut the target group's reward error by 83% and the learned policy's trajectory error by 89%, with no significant change in the control group. The transfer claim is that this teaching skill generalises: on a line-reaching skill not seen in training, reward error fell by 64%, a statistically significant change, while the larger trajectory improvements (70% lower ARMSE, 91% lower ATR) did not reach significance, which the paper explains by noting that multiple policies give the same cumulative reward for this task.","pith_inferences":["Taken as an inference from the mathematical fact the paper cites, Eq. (20) is an inequality in disguise: the constant linking reward error to value-parameter error depends on the conditioning of the sampled state-action matrix, so in poorly conditioned samples the training benefit may not propagate to the robot's policy.","The reward-ambiguity explanation for the non-significant S2 trajectory gains is directly testable: repeat the transfer test with a skill whose optimal policy is unique; if trajectory metrics become significant, ambiguity is confirmed, and if not, the transfer claim needs a different mechanism.","The scaffolding curriculum separates state distance, action direction, and action magnitude, so the same sequence could be applied to other learners with linear value features, though the machine-teaching target would no longer be closed-form if the learner changes.","A practical consequence the authors leave implicit is that short, targeted reward-teaching drills could substitute for extensive demonstration training in workplaces, since the measured benefit appears in demonstration quality and transfers across reward functions."],"forward_implications":["Eight guided demonstrations on one reaching skill suffice to produce significant gains in a novice's reward-teaching accuracy on that skill.","The teaching improvement transfers to a never-practiced skill at the demonstration level: reward error fell 64% on the line-reaching task.","On the training skill, the robot's trajectory error fell 89% and its true-reward shortfall by 98%, so better reward labels translate into better control.","On the unseen skill, trajectory-level gains (70% lower trajectory error, 91% lower total-reward shortfall) were large but not statistically significant; the paper attributes this to reward ambiguity.","Trained teachers who teach via reinforcement-learning demonstrations produce policies with better long-horizon stability than teachers trained for supervised learning-from-demonstration, which matters when rollouts are long."],"supporting_citations":[{"why":"supplies the earlier MT-training protocol for supervised LfD that this paper adapts to RLfD, and is the cited source for Eq. (20).","marker":"[4]"},{"why":"defines machine teaching as a bi-level optimisation problem from which the paper's teaching-risk formulation is taken.","marker":"[9]"},{"why":"establishes the teaching dimension of linear learners, fixing the teaching budget at N=T=8 demonstrations.","marker":"[12]"},{"why":"provides a robotics-oriented instance of least-squares policy iteration that supports the choice of LSPI as the learner.","marker":"[10]"},{"why":"supplies the LSPI algorithm whose closed-form parameter update appears in Eq. (12) and drives the MT problem.","marker":"[11]"},{"why":"gives the sub-multiplicative matrix-norm property used to derive the proportionality between reward error and value-parameter error.","marker":"[13]"},{"why":"supports the scaffolding/curriculum-learning sequence that structures the training phases P3-P7.","marker":"[18]"}],"fun_headline_variants":["Guided demos boost robot learning 89% and transfer to new tasks","Teaching humans to teach robots: 8 demos, 89% better learning","Eight guided demos: 89% robot learning gain, 70% gain on new skills","Guided teaching improves robot learning 89% and transfers to novel skills"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that bringing a novice's reward judgments closer to the ideal will bring the robot's learned value function closer to the target by a proportional amount; the paper derives this from a matrix-norm inequality, but the inequality only guarantees a bound, so the strength of that link can vary from task to task.","fun_headline_variants_meta":{"raw":{"variants":["Guided demos boost robot learning 89% and transfer to new tasks","Teaching humans to teach robots: 8 demos, 89% better learning","Eight guided demos: 89% robot learning gain, 70% gain on new skills","Guided teaching improves robot learning 89% and transfers to novel skills"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000502,"raw_usage":{"total_tokens":2432,"prompt_tokens":902,"completion_tokens":1530,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1443}},"tokens_in":518,"tokens_out":1530,"duration_ms":11461,"temperature":1.0,"reasoning_tokens":1443,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:15:15.520279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the ratio $\\|\\omega-\\bar{\\omega}\\|_2 / \\|j-\\bar{\\jmath}\\|_2$ across many random state-action samples in the least-squares policy iteration setup; if it varies by orders of magnitude, the proportionality in Eq. (20) fails and reward-error gains would not reliably become robot-learning gains. A companion check: rerun the transfer experiment with an unseen skill that has a unique optimal policy; if trajectory-level gains still fail to reach significance, the reward-ambiguity explanation is not sufficient.","supporting_citations":[{"cited_title":"Using machine teaching to boost novices’ robot teaching skill,","cited_arxiv_id":null,"evidence_quote":"supplies the earlier MT-training protocol for supervised LfD that this paper adapts to RLfD, and is the cited source for Eq. (20)."},{"cited_title":"The teaching dimension of linear learners,","cited_arxiv_id":null,"evidence_quote":"establishes the teaching dimension of linear learners, fixing the teaching budget at N=T=8 demonstrations."},{"cited_title":"Least-squares policy iteration algorithms for robotics: Online, continuous, and automatic,","cited_arxiv_id":null,"evidence_quote":"provides a robotics-oriented instance of least-squares policy iteration that supports the choice of LSPI as the learner."},{"cited_title":"Locally weighted least squares policy iteration for model-free learning in uncertain environments,","cited_arxiv_id":null,"evidence_quote":"supplies the LSPI algorithm whose closed-form parameter update appears in Eq. (12) and drives the MT problem."},{"cited_title":"Curriculum learning,","cited_arxiv_id":null,"evidence_quote":"supports the scaffolding/curriculum-learning sequence that structures the training phases P3-P7."}],"review_version":1}