{"id":"a0b26ee3-fef6-41b2-9592-370391e4f9c4","arxiv_id":"2411.10176","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"People who practiced a simulated task without AI help learned it better than people who got explainable computer or robot assistance.","lead":"In a learning-by-doing experiment, volunteers either managed a simulated power plant alone, with an explainable computer, or with a robot. Those who worked alone later scored higher on a knowledge test, suggesting that AI assistance can reduce exploration and weaken genuine learning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Anomaly count as an exploration proxy is confounded with practice volume; the Self-taught group's superior test score may reflect more actions, not 'failing/exploring.'","rationale":"The reader's weakest assumption identified the same load-bearing point: anomalies are used as a proxy for exploration, and exploration is invoked as the causal mechanism behind the Self-taught group's better test performance. My stress-test agrees and makes the concern more concrete: the paper never validates the anomaly proxy, and the raw anomaly count is entangled with the number of actions performed, which differs systematically between Self-taught and assisted groups. Because the paper already reports that Self-taught participants performed more actions, more energy, more critic steps, and more anomalies, the simplest explanation of the post-test difference is more practice, not more exploration or 'letting people fail.' The paper's own stated limitation about the fixed-time design supports this worry. This concern does not overturn the paper: it is a conditional empirical claim that could be checked with an ANCOVA or a fixed-step replication. Since the reader already returned CONDITIONAL with high confidence, my read does not change the verdict, so UNCHANGED is appropriate while noting that the condition should explicitly include the practice-volume confound.","tokens_in":20321,"tokens_out":4556,"duration_ms":49679,"concrete_test":"Reanalyze the existing data with an ANCOVA on post-experiment test scores (whole test and scenarios subscore) with experimental group as factor and number of training actions as covariate; additionally compare anomaly rate (anomalies per 100 actions) rather than raw anomaly count across groups. If the Self-taught advantage disappears after controlling for action count, or if anomaly rate is not higher in the Self-taught group, the 'failing/exploration' interpretation is unsupported and the result is attributable to practice volume. If the advantage and higher anomaly rate survive, the central claim is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central 'let people fail' claim depends on the interpretive step in Sections 4.6.2 and 5 that the number of anomalies during training is 'a good estimator of participants' degree of exploration,' and that this exploration explains the Self-taught group's superior post-test performance. This step is load-bearing but not secured. Anomalies are damaging outcomes, not a validated measure of exploratory behavior: no independent metric such as state-space coverage, action diversity, or information gain is reported, and no mediation analysis links anomalies to test performance. More concretely, anomaly count is confounded with practice volume. Section 4.6.1 shows the Self-taught group performed significantly more training actions than every assisted group because participants did not spend time querying an agent, giving them more opportunities both to produce anomalies and to learn the task. The paper's own limitations paragraph acknowledges that the fixed-time design makes it unclear whether results would hold 'using a fixed number of steps,' but this confound is not tested. Without controlling for number of actions or training time, the post-test advantage could reflect more hands-on practice rather than autonomous exploration or failure-driven learning, and the exploration-mediated mechanism is not identified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a between-subject user study (N = 55) on learning-by-doing with a simulated nuclear power plant task. Participants either worked autonomously (Self-taught, n = 11) or were assisted by a computer or the iCub robot that could answer 'what' and 'why' queries using either classical (C-XAI) or partner-aware contrastive (A-XAI) explanations derived from a deterministic decision-tree policy. The authors compare behavior in a 30-minute training phase and a 10-minute assessment phase, together with a post-experiment knowledge test. They find that A-XAI made computer users move faster, made the robot more persuasive after 'why' questions, and had no significant effect on final task knowledge. The headline result is that Self-taught participants performed more actions, produced more energy and more anomalies, and scored higher on the post-experiment test, especially on scenario questions, than most assisted groups. The authors interpret the anomaly count as a measure of exploration and argue that unassisted 'failing' and exploring produced better learning, thereby rejecting their hypothesis H2.","tokens_in":20560,"tokens_out":7272,"duration_ms":68488,"significance":"The study's strengths are its direct measurement of learning via a post-experiment knowledge test, its inclusion of a no-assistance baseline, and its concrete manipulation of explanation strategies through a deterministic interpretable decision tree. If the headline effect is robust, it has practical implications for automated tutoring and for designing AI assistants that scaffold rather than replace exploration. The paper also gives useful, explicit credit to the possibility that over-reliance or automation bias explains the assisted groups' lower test scores. However, the central interpretation depends on an unvalidated proxy (anomalies as exploration) and is entangled with a practice-volume confound; in addition, several statistical and reporting issues need correction. With those issues addressed, this could be a valuable contribution to HCI and XAI research.","major_comments":[{"comment":"The conclusion that the Self-taught group's higher post-test scores were caused by greater 'exploration' rests entirely on the assertion in Section 5 that 'the number of anomalies during training is a good estimator of participants' degree of exploration.' This assertion is not secured: anomalies are damaging outcomes, not a validated measure of exploration, and no independent index such as state-space coverage, action diversity, information gain, or query behavior is reported. Because the Self-taught group also performed significantly more actions (Section 4.6.1), the anomaly count is confounded with practice volume. Without a per-action anomaly rate or a mediation analysis linking anomalies to test scores, the data are equally consistent with the alternative explanation that more hands-on practice, rather than autonomous exploration or failure-driven learning, produced the test advantage.","section":"Section 4.6.2 and Section 5"},{"comment":"The fixed-time design directly creates the practice-volume confound: Self-taught participants did not spend time querying an agent, so they could perform more actions in the same 30 minutes. The manuscript acknowledges in the limitations paragraph that it is 'unclear whether we would observe the same results by removing or changing such a constraint, i.e., using a fixed number of steps,' but it does not test this. Since the behavioral differences (actions, energy, anomalies, critic steps) and possibly the test differences depend on this design choice, this limitation is load-bearing rather than peripheral.","section":"Sections 3.1, 4.6.1, and 5"},{"comment":"The baseline knowledge check is only marginally non-significant (chi-square(8) = 14.1, p = .078), and the authors themselves note a 'slight disparity' between the COM group (Germany) and the Self-taught and Robot groups (Italy). The cross-group ANOVAs in Sections 4.6.1 and 4.6.2 do not include site or prior knowledge as a covariate, and the statement that there were 'no differences between these two macro-groups regarding performance in training and assessment' is made without a supporting test. This leaves open the possibility that the Self-taught group's superior test performance partly reflects pre-existing differences rather than the absence of assistance.","section":"Section 4.1 and Section 5"},{"comment":"The manipulation check in Table 1 is based on sixteen independent-samples t-tests without any multiple-comparison correction, so the reported 'five out of eight' and 'four out of eight' significant explanandum differences are not strong evidence that the two explanation strategies systematically differed. Uncorrected t-tests also appear in Sections 4.3-4.5 (for example, t = 2.21, p = .039, and t = 2.226, p = .038). Please report adjusted p-values or explicitly justify these comparisons as exploratory.","section":"Section 4.2 and Table 1"},{"comment":"Reporting the same t-statistic (t = 2.226) and p-value (p = .038) for both the COM and Robot subgroups in the comparison of 'Follow AI' percentages during why-questions is internally inconsistent if this is a single between-group independent-samples test. As written, the reader cannot tell whether this is a paired comparison, a typographical error, or a duplicated statistic. Please clarify the test, report degrees of freedom and effect size, and correct the descriptive statistics.","section":"Section 4.5 and Figure 8"}],"minor_comments":[{"comment":"Section 3.4 lists hypotheses H1-H3, but Section 5 refers to 'our hypothesis H4' and 'partially confirming our hypothesis H4'; please define H4 or correct the label.","section":"Section 3.4 and Section 5"},{"comment":"The caption states that the asterisk refers to a 'p-value < .5' where p < .05 is clearly intended; as written the caption would include all significant and non-significant results.","section":"Figure 7 caption"},{"comment":"Section 4.6.2 reports significant differences between the Self-taught group and the COM C-XAI, COM A-XAI, and Robot A-XAI groups, with Robot C-XAI as the only non-significant comparison; Section 5 restates this as 'all the others but the COM A-XAI one' and attributes the effect to a COM A-XAI outlier. This inconsistency should be resolved.","section":"Section 4.6.2 and Section 5"},{"comment":"The sentence 'the former negatively correlated with the percentage of times participants accepted the robot's suggestions (and not the COM's)' is ambiguous; please state explicitly which correlation is reported for which group and report the correlation coefficient.","section":"Section 4.5"},{"comment":"ANOVA results are reported as F(4) without residual degrees of freedom; reporting the full F(df1, df2) notation, e.g., F(4, 50), would aid verification of the analyses.","section":"Section 4.6"},{"comment":"The paper does not state whether data or analysis scripts are available; a data availability statement would strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"Confidential to the editor: The headline claim is broader than the evidence currently supports. I do not see evidence of deliberate misreporting, but the identical t-statistic in Section 4.5 and the undefined H4 reference suggest the manuscript needs a careful proofreading pass by the authors. The central confound can in principle be addressed with reanalysis (for example, per-action anomaly rates, action diversity, site as a covariate) or by softening the causal language, so I recommend major revision rather than rejection. The paper may be better suited to an HCI or HRI venue than a general cs.AI venue, but that is a scope decision for the editors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for the self-taught group result. In a learning-by-doing task, participants who got no AI assistance scored better on the post-test than those who could query an explainable agent. That's a useful counterexample to the default assumption that expert guidance improves learning.\n\nWhat's new: the comparison of partner-aware contrastive explanations vs classical ones across a virtual agent and a humanoid robot, plus the self-taught control. The task is well-defined and the AI is a transparent decision tree, so the explanation strategies are concrete algorithmic variants. The behavioral results (faster decisions with adaptive explanations on the computer, more persuasion by the robot) are plausible, and the main test-score comparisons use ANOVA with Bonferroni post-hocs, which is appropriate.\n\nThe main soft spot is the interpretive step in Sections 4.6.2 and 5: the paper claims anomaly count is a good estimator of exploration and that this explains the self-taught group's superior test scores. That isn't secured. The self-taught group performed significantly more actions per unit time, so they had more chances to produce anomalies and more hands-on practice. The paper's own limitations paragraph notes the fixed-time design is a confound, but it doesn't test it. A mediation analysis or an additional measure like state-space coverage would be needed to support the \"let people fail\" mechanism. Without that, the result stands as an empirical finding, but the explanation is speculative.\n\nAlso, the duplicated t-statistic for explanation verbosity in Sections 4.3 and 4.4 (t=2.31, p=.03 reported identically for both COM and Robot groups) looks like a copy-paste error. And Table 1 runs eight uncorrected t-tests per group. These are fixable but should be cleaned up.\n\nCredit: the paper is honest about its limitations and doesn't oversell the causal claim—it explicitly says \"In our opinion.\" The empirical design is reasonable given the resource constraints, and the sample sizes, while small, yield effects that survive Bonferroni.\n\nWho this is for: HCI/HRI/XAI researchers interested in automated tutoring and the effects of explanations on learning. The result will likely be cited as evidence that assistance can reduce generalization.\n\nRecommendation: send it to peer review. It deserves referee time. The main revision should address the exploration confound—either with an additional analysis or a stripped-down claim.","headline":"Solid empirical study with a provocative counterintuitive result, but the central interpretation rests on an unvalidated proxy for exploration.","tokens_in":21049,"tokens_out":2232,"would_cite":true,"duration_ms":21278,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that in a learning-by-doing task, participants who learned without any explainable AI agent explored more and ended up with better task knowledge than those assisted by a computer or a humanoid robot.","keywords":["explainable artificial intelligence","human-robot interaction","learning-by-doing","human-AI collaboration","automation bias","contrastive explanations","user study","robotic tutoring"],"falsifier":"Conduct the same learning-by-doing task with a fixed number of training steps instead of a fixed time limit; if the unassisted group no longer shows more anomalies and no longer outscores the assisted groups, the exploration explanation fails. Alternatively, measure exploration directly, for example by counting distinct environment states visited or computing information gain, and check whether that measure, rather than the anomaly count, drives the test-score difference.","tokens_in":20162,"feed_emoji":"🤖","tokens_out":5794,"duration_ms":49812,"temperature":0.7,"pith_summary":"This paper asks whether explainable AI agents actually help people learn a new task, and reports a counterintuitive result: the people who received no help at all learned more. In a simulated nuclear-power-plant management task, participants who worked alone made more mistakes during training, but those mistakes are read as exploration, and the self-taught group went on to score significantly higher on a post-task knowledge test than participants assisted by an explainable computer or humanoid robot. The paper also found that two explanation styles, classical and partner-aware contrastive explanations, produced no difference in final knowledge, but did change behavior: with a computer, adaptive explanations made people faster, while with a robot they made people more likely to follow its suggestions. The authors conclude that explainable assistance can reduce learners' agency and exploration, and suggest that automated tutors should invite people to fail and explore rather than simply follow suggestions.","feed_headline":"Let people fail: self-taught group outperforms AI-assisted learners","feed_subtitle":"In a simulated power-plant task, unassisted participants explored more and scored higher on the final knowledge test.","key_machinery":"The central object is the learning-by-doing assessment task itself: a simulated nuclear power plant with hidden rules that participants must discover by acting, with a 30-minute training phase, a 10-minute assessment phase, and a final knowledge test. The expert agents are driven by a deterministic decision tree trained with the Conservative Q-Improvement reinforcement-learning algorithm, which can answer what and why questions by tracing splits. The two explanation strategies are classical selection of the most relevant feature and partner-aware contrastive selection that compares the agent's suggestion with the user's predicted action. The argument's load-bearing measure is the number of anomalies (conditions that damage the simulated plant) during training, which the paper treats as an estimator of exploration.","core_discovery":"On the paper's own terms, the central discovery is that assisted learning underperforms unassisted learning in a learning-by-doing task. In a between-subjects study, the Self-taught group produced more actions, more energy, more critical steps, and more anomalies during training than all four assisted groups, and then outperformed the COM C-XAI, COM A-XAI, and Robot A-XAI groups on the post-experiment test (ANOVA F(4)=3.99, p=.007), with the effect concentrated in scenario questions (F(4)=6.37, p<.001). The authors interpret the higher anomaly count as evidence of deeper exploration, and argue that interaction with expert explainable agents encouraged over-reliance and reduced autonomous exploration, leaving assisted participants with weaker ability to generalize.","pith_inferences":["Beyond the paper: one could test whether adding a mechanism that forces users to commit to a prediction before seeing the agent's suggestion closes the gap between assisted and unassisted learners, since the paper identifies over-reliance as the likely culprit.","Beyond the paper: the different effects of the same adaptive explanation strategy on computer versus robot interaction suggest that embodiment and social presence may change whether explanations promote reflection or compliance; this could be isolated by matching the agents' wording and nonverbal cues more tightly.","Beyond the paper: the anomaly-as-exploration proxy deserves direct validation; if validated, it gives XAI researchers a simple behavioral marker for exploration in learning tasks, independent of self-reports."],"forward_implications":["If correct, explainable agents that merely answer questions can reduce learning gain in learning-by-doing tasks, even when they improve immediate performance.","Automated tutoring systems should be designed to encourage autonomous exploration, tolerating errors during training as a price for deeper understanding.","Partner-aware contrastive explanations did not improve final knowledge over classical explanations, but they did change behavior: faster decisions with a computer and more compliance with a robot.","The self-taught group's advantage appeared specifically in scenario-based generalization questions, suggesting that unassisted exploration builds transferable causal knowledge rather than rote procedures."],"supporting_citations":[{"why":"Supplies the learning-by-doing theory that motivates the task design.","marker":"Anzai and Simon (1979)"},{"why":"Provides the instructional-design account of learning by doing that grounds the training paradigm.","marker":"Schank et al. (2013)"},{"why":"Supplies the Conservative Q-Improvement algorithm used to train the interpretable decision-tree agents.","marker":"Roth et al. (2019)"},{"why":"Provides the automation-bias concept used to explain assisted participants' over-reliance.","marker":"Vered et al. (2023)"},{"why":"Shows cognitive forcing can reduce over-reliance, used as the comparison point for mitigation strategies.","marker":"Buçinca et al. (2021)"},{"why":"Documents that over-simplified explanations increase mental demand, used to explain the assisted groups' cognitive load.","marker":"Kulesza et al. (2013)"},{"why":"Documents the illusion of explanatory depth, used to explain over-confidence toward agents.","marker":"Chromik et al. (2021)"},{"why":"Identifies trust-calibration errors in users of explainable recommendations, used to contextualize the assisted groups' behavior.","marker":"Naiseh et al. (2021)"}],"fun_headline_variants":["Self-taught learners beat AI-assisted peers in hands-on task","No help, better learning: unassisted group outshines AI-guided in study","AI assistance backfires in learn-by-doing, self-taught excel","Letting people fail: hands-on autonomy tops explainable AI help","Learning by doing: going solo outperforms robot and computer aids"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's argument depends on treating the number of anomalies during training as a good estimator of exploration, and on assuming that exploration, not some other factor like time pressure or prior knowledge, is what gave the unassisted group its higher test scores.","fun_headline_variants_meta":{"raw":{"variants":["Self-taught learners beat AI-assisted peers in hands-on task","No help, better learning: unassisted group outshines AI-guided in study","AI assistance backfires in learn-by-doing, self-taught excel","Letting people fail: hands-on autonomy tops explainable AI help","Learning by doing: going solo outperforms robot and computer aids"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1449,"prompt_tokens":893,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":462}},"tokens_in":509,"tokens_out":556,"duration_ms":5903,"temperature":1.0,"reasoning_tokens":462,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:52:42.759250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the same learning-by-doing task with a fixed number of training steps instead of a fixed time limit; if the unassisted group no longer shows more anomalies and no longer outscores the assisted groups, the exploration explanation fails. Alternatively, measure exploration directly, for example by counting distinct environment states visited or computing information gain, and check whether that measure, rather than the anomaly count, drives the test-score difference.","supporting_citations":[{"cited_title":"Learning by doing","cited_arxiv_id":null,"evidence_quote":"Provides the instructional-design account of learning by doing that grounds the training paradigm."}],"review_version":1}