{"id":"b03672d8-c1a1-4d86-b953-2d68a446fc83","arxiv_id":"2506.12507","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Furhat robot with scripted arguments swayed high-school students' final true/false answers on electric circuits in most disagreements, with higher certainty displays increasing alignment and self-reported LLM experience predicting greater susceptibility to wrong answers.","lead":"In a study with 40 Swedish high-school students, a social robot's arguments changed students' final answers on electric circuit questions in 117 of 139 cases where they initially disagreed, and 30 of 40 students aligned with the robot's wrong answer on at least one easy question. The paper matters because it suggests students, especially frequent LLM users, may over-trust AI tutors, and that displaying uncertainty can reduce this undue influence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 75% 'beyond expected capacity' headline is computed with a 3PL guessing floor (c≥0.25) from three-choice diagnostic items, but the experimental items are binary true/false; the expected probabilities are miscalibrated.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the central phenomenon—students aligning with the robot, including when it is wrong—is credible from the raw event counts. The load-bearing weakness I identify is not the absence of a no-robot control per se, although that also matters; it is that the specific 75% beyond-capacity statistic is generated against an expected-performance model whose guessing parameter is calibrated to a three-choice response format and then applied to binary true/false items. This is an internal inconsistency, not a disagreement with the field's consensus: the same p(θ) formula with c<0.5 gives a lower asymptote that is impossible for a two-choice item. Because the Monte Carlo null is built from these probabilities, the p-values that define 'beyond expectations' are systematically too small for correct answers. A sensitivity analysis with c=0.5 is cheap and would settle whether the headline survives. The event-level claims do not depend on this correction and should retain their weight. Thus the verdict remains CONDITIONAL, with the condition expanded to include this sensitivity check; no move to reject is warranted because the raw alignment data independently support the core finding.","tokens_in":18203,"tokens_out":9006,"duration_ms":112948,"concrete_test":"Recompute the Sec. IV-B Monte Carlo/Fisher analysis with the guessing parameter fixed at c=0.5 for the eight experimental true/false items, keeping the reported a,b values (or re-estimating the 3PL on the preliminary-answer data). Compare the fraction of students classified as beyond expected capacity and the Deception/non-Deception direction. If the 75% drops materially or the above/below pattern reverses, the headline claim needs revision. The authors should release item parameters and per-student abilities to permit this check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim (Sec. IV-B: '75% of the students performed beyond their expected capacity... direct result of robot's influence') is defined against expected correctness probabilities from a 3PL IRT model fit to the diagnostic test (Sec. III-A). The diagnostic test used 19 three-choice items, and the model fixes a guessing floor c≥0.25. The eight experimental problems, however, are TRUE/FALSE statements (Sec. III-D, Fig. 2). For a binary item, the lower asymptote of the response function should be at most 0.5, not 0.25. Using p(θ)=c+(1−c)/(1+e^{−a(θ−b)}) with the diagnostic c understates the expected probability of a correct answer for low-ability students on difficult items by up to 0.25. The Monte Carlo p-values and Fisher combined p-values in steps 4–5 of Sec. IV-B therefore classify answers as 'beyond expectations' too liberally. The same bias enters PA and FA, so the PA 'within expectation' check does not expose it. In-sample fit to the same 36-student cohort and item selection from the same diagnostic data add further overfitting risk. The event-level alignment data (117/139) are compelling, but the quantitatively central 75% claim is not robust until the response-format mismatch is addressed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an experiment with 40 Swedish high-school students who interacted one-on-one with a Furhat robot solving eight true/false electric-circuit statements. The robot argued for the correct answer on six questions and the incorrect answer on two, and it displayed three levels of certainty (Uncertain, Neutral, Certain) across groups. The main reported results are that students aligned with the robot in 117 of 139 cases of initial disagreement, including 34 cases in which the robot was wrong; that alignment was higher for Certain (94.4%) than for Neutral (82.6%) and Uncertain (71.4%); that 75% of students performed “beyond their expected capacity” as defined by a 3PL IRT model calibrated on a diagnostic test; and that self-reported LLM experience was associated with greater alignment on Deception questions. The paper interprets these results as evidence of robot-induced informational influence and discusses implications for educational robotics and AI trust.","tokens_in":18474,"tokens_out":6916,"duration_ms":87036,"significance":"The event-level alignment counts and the certainty manipulation are valuable and largely defensible: the certainty portrayals were validated in separate rating studies, the GLMM analyses control for ability and difficulty, and the post-test self-reports are consistent with the observed alignment patterns. If the central performance-deviation claim could be placed on a sound footing, the paper would make a strong contribution to educational robotics and to the study of trust and overtrust in AI. As it stands, however, the headline “75% beyond expected capacity” result depends on model-based expected probabilities that appear miscalibrated for the binary response format, and it relies on an in-sample calibration with no no-robot control condition. The causal interpretation of the PA-to-FA changes is therefore not yet fully supported, and the abstract-level claims go beyond what the current design can establish.","major_comments":[{"comment":"The 3PL model in Sec. III-A fixes a guessing lower asymptote c >= 0.25 because the diagnostic items had three options, but the experimental items are true/false (Sec. III-D, Fig. 2). For a two-option item, the chance-level success probability is 0.5, so p(theta) = c + (1-c)/(1+e^{-a(theta-b)}) yields expected correct probabilities that are too low for low-ability students. The Monte Carlo simulation and Fisher combined-p-value procedure in Sec. IV-B will therefore classify final-answer outcomes as “beyond expectations” too liberally; applying the same biased model to preliminary answers does not cancel the bias, since both PA and FA are evaluated with the same miscalibrated expectation. The 75% claim is not robust until the expected probabilities are re-estimated with a lower asymptote appropriate for binary items, or the model is recalibrated on the experimental response format.","section":"Sec. III-A and Sec. IV-B"},{"comment":"The interpretation of PA-to-FA changes as “a direct result of robot’s influence” is not warranted by the design. Students answer each question twice, and the second answer is always given after the robot has argued; without a no-robot control condition in which the same questions are answered twice without argumentation, retesting effects, reflection, and demand characteristics are plausible alternative explanations for the 117/139 alignment rate. This is not a minor caveat for the central causal claim: the paper needs either a control condition or a substantially more hedged interpretation of the alignment results.","section":"Sec. IV-B and Sec. II"},{"comment":"Question selection and IRT calibration use diagnostic answers from the same 40-student cohort, and the Monte Carlo p-values treat the estimated item parameters (a_i, b_i, c_i) and abilities (theta_j) as known. Because the same fitted model defines “expected capacity,” the statistical uncertainty in calibration is not propagated, and the risk of in-sample overfitting is substantial. This is especially concerning given the small cohort and the fact that item selection was itself guided by the same diagnostic data. A sensitivity analysis, such as bootstrapping item parameters or using split-half calibration, is needed before the “beyond expectations” count can be regarded as conservative.","section":"Sec. III-A, Sec. III-D, Sec. IV-B"},{"comment":"The LLM-usage finding, which appears in the abstract and conclusions, is selected from a large set of single-factor ANOVAs without multiple-comparison correction. The reported p = 0.038 for Deception alignment would not survive a simple Bonferroni correction across the many examined characteristics, and several other results in Table III are marginal (p values near 0.05–0.09). The abstract-level claim about LLM experience should either be supported by a confirmatory, pre-registered analysis or presented as exploratory.","section":"Sec. IV-D and Table III"}],"minor_comments":[{"comment":"There is a missing space in “c≥0.25to prevent over-fitting,” and the statement that the 3PL model is “particularly appropriate for three-option multiple-choice tests” should be reconciled with its use for binary experimental items.","section":"Sec. III-A"},{"comment":"The no-shows produced unbalanced groups with different raw diagnostic scores, and the paper states that IRT ability will be used as a covariate; the authors should report the actual IRT ability means per group, not only raw DA scores, to support the claim that ability was adequately controlled.","section":"Sec. III-C"},{"comment":"The sentence about the non-alignment rate in Deception questions is duplicated: “The rate of non-alignment in Deception was in fact doubled compared to non-Deception” appears twice in near-identical form in Sec. V; one occurrence should be removed.","section":"Sec. IV-A"},{"comment":"The statement that “98 out of 117 alignments directly contributed to the deviation from expected results” is not defined precisely; the authors should specify what “directly contributed” means in terms of the Monte Carlo and Fisher procedure.","section":"Sec. IV-B"},{"comment":"The GLMM results for Conditioned questions report p-values but no effect sizes or confidence intervals; adding these would help readers assess the magnitude of the certainty effect.","section":"Sec. IV-C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a human-robot interaction audience, and the raw alignment data are compelling. However, the headline 75% claim depends on an IRT calibration with a response-format mismatch, and the causal interpretation lacks a no-robot control. I would recommend asking for a revised version that recalibrates the expected-probability model, adds a sensitivity analysis, and tempers the causal language, before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: the 75% 'beyond expected capacity' statistic is likely inflated. The IRT model that generates those expectations was fitted to three-choice diagnostic items with a guessing floor c≥0.25, but the eight experimental problems are true/false. For a binary item the guessing floor should be 0.5. That means the model under-predicts how well low-ability students should do by chance, so observed correct answers get counted as 'beyond expectations' too easily. The PA/FA comparison doesn't cancel this, because both are scored against the same miscalibrated model.\n\nThat said, the paper is worth reading. The event-level result is straightforward and compelling: in 117 of 139 disagreements, students changed their final answer to match the robot, including 34 times when the robot was wrong. The certainty gradient (94.4% vs 82.6% vs 71.4%) is clean, and the authors took care to validate the U/N/C portrayals with separate raters. The LLM-experience effect, while based on self-report and uncorrected comparisons, is a timely lead for AI trust research. The classroom setting and the intent-based autonomy of the robot are real strengths.\n\nThe soft spots extend beyond the IRT issue: there is no no-robot control, so part of the PA-to-FA shift could be retesting or reflection; the sample is 40 with known group imbalance; and the personality/LLM analyses use many single-factor ANOVAs without correction. Some of these the authors acknowledge in the limitations section, which is honest. But the phrase 'direct result of robot's influence' is too strong given the missing control.\n\nWho should read this: anyone working on educational robots, AI trust, or automation bias. The certainty manipulation methodology is a useful contribution even if the headline number moves.\n\nMy recommendation: send to peer review, not desk reject. The event-level findings and the design deserve referee time. The revision needs to fix or re-derive the expected-performance model for binary items, or at least show a sensitivity analysis with c=0.5, and ideally share the data and code. Without that, the 75% claim shouldn't stand.","headline":"Solid event-level evidence on robot persuasion, but the headline 75% 'beyond expected capacity' is likely inflated by an IRT guessing-parameter mismatch.","tokens_in":19036,"tokens_out":2516,"would_cite":true,"duration_ms":27725,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a social robot's arguments moved 117 of 139 dissenting high-school answers, 34 to wrong answers, and that displayed certainty is a major driver of this alignment.","keywords":["social robots","educational robotics","conformity","persuasion","informational trust","certainty cues","LLM experience","item response theory"],"falsifier":"Give a matched group of students the same eight true/false questions twice with no robot, or with a recorded voice instead of a present robot. If the switch rate and the rate of beyond-expected performance are as high as in the robot condition, the paper's causal reading fails; if they are substantially lower, the robot-influence claim survives.","tokens_in":17964,"feed_emoji":"🤖","tokens_out":9092,"duration_ms":101965,"temperature":0.7,"pith_summary":"The paper tries to establish that high-school students are strongly persuaded by a social robot's arguments on a familiar school subject, electric circuits, even when the robot is wrong. In 117 of 139 cases where a student's preliminary answer disagreed with the robot, the student changed the final answer to match it, and 34 of those switches were to incorrect answers. Using an item-response-theory model of each student's expected performance, the authors report that 75% of students finished beyond their expected capacity, better on questions where the robot was right and worse where it was wrong, and interpret this as direct robot influence. They also report that displayed certainty matters: students aligned 94.4% of the time with a certain robot, 82.6% with a neutral one, and 71.4% with an uncertain one. Students who reported more experience with large language models were more likely to follow the robot's incorrect answers.","feed_headline":"A social robot swayed 84% of dissenting students","feed_subtitle":"Even when the robot was wrong, students aligned; displayed certainty raised compliance from 71% to 94%","key_machinery":"The argument is carried by three linked devices. The first is an item-response-theory model, specifically the three-parameter logistic model, fitted to the students' diagnostic answers to estimate each student's latent ability and each question's difficulty, which defines the expected capacity that the 75% claim is measured against. The second is the alignment variable, a tri-state coding of each interaction event as agreement, resistance, or change toward the robot, with Monte Carlo simulations and Fisher's method used to decide when observed performance is beyond expectation. The third is the multimodal certainty manipulation, differences in wording, speech rate, pauses, gaze, smiles, and head movements, that defines the Certain, Neutral, and Uncertain conditions and produces the graded alignment rates.","core_discovery":"The central discovery is that a social robot does not need to be right to be followed: it persuaded a large majority of 40 high-school students to revise their answers on eight true/false electric-circuit questions, including when its argument was deliberately wrong on the two easiest questions. The authors frame this through the distinction between sense, the students' reliable ability to judge the arguments, and sensibility, their responsiveness to the robot's expressed certainty. They find that 75% of students performed beyond their IRT-predicted capacity, above expectation on the non-deceptive questions and below expectation on the deceptive ones, and that this shift from preliminary to final answers should be read as a direct result of the robot's influence. Displayed certainty was a decisive cue: a robot portrayed as certain drew alignment in 94.4% of disagreements, an uncertain one in 71.4%, and students rated the certain robot as most convincing. Prior experience with large language models increased alignment with incorrect answers, suggesting that familiarity with AI can raise susceptibility rather than critical resistance.","pith_inferences":["Beyond the paper: a matched no-robot control, with students answering the same eight questions twice, would isolate how much of the 117/139 switch rate is a retest or reflection effect rather than robot causation.","Beyond the paper: because the expected-capacity baseline is fitted to the same small cohort, the 75% figure could be rechecked with item parameters estimated from a larger independent sample of the same questions.","Beyond the paper: if the large-language-model experience effect is causal, the same susceptibility may appear with text-based chatbots, which could be tested by replacing the embodied robot with a screen agent while keeping the arguments identical.","Beyond the paper: a practical design rule follows, that robots should display certainty calibrated to their information's reliability, which future systems with LLM reliability metrics could implement directly."],"forward_implications":["Educational robots that argue for answers can override students' own reasoning on familiar material, so designers should treat persuasion as a safety-relevant property rather than a side effect.","Displayed certainty is a practical control knob: tying a robot's confidence signals to the actual reliability of its content could reduce acceptance of wrong information.","Students with heavy large-language-model experience appear more, not less, vulnerable to an AI's wrong answer, so AI-literacy curricula should address trust calibration rather than just tool competence.","Because alignment rates did not depend on measured ability, adjusting question difficulty or selecting strong students will not by itself prevent over-alignment.","The absence of carry-over effects suggests each robot-student exchange is its own persuasion event, so interventions to reduce overtrust may need to operate question by question."],"supporting_citations":[{"why":"Supplies the back-projected social robot head used in the experiment, whose expressive face and speech carry the certainty manipulations.","marker":"[1]"},{"why":"Provides Fisher's method used to combine p-values across Deception and non-Deception questions in the beyond-expectations analysis.","marker":"[19]"},{"why":"Grounds the informational-trust mechanism, since reliability and failure rates determine trust in human-robot interaction.","marker":"[22]"},{"why":"Supports the prosodic manipulation of certainty through speech rate and disfluency effects on perceived confidence.","marker":"[32]"},{"why":"Establishes the review-based claim that persuasive robots mostly use peripheral cues and that argument-based persuasion is understudied.","marker":"[38]"},{"why":"Supplies the Elaboration-Likelihood Model used to interpret why flawed arguments and certainty cues persuade or fail to persuade.","marker":"[51]"},{"why":"Provides the prior demonstration that humans conform to robots, the effect this study extends to educational argumentation.","marker":"[55]"},{"why":"Defines multimodal signals of certainty used to build the robot's facial-expression cues for the Certain and Uncertain conditions.","marker":"[64]"}],"fun_headline_variants":["Certain robots sway 94% of students, even when wrong","Robot confidence boosts student compliance from 71% to 94%","AI-savvy students are more easily misled by robots","Confident robots convince students, reducing critical thinking","Students follow wrong robot answers when it acts sure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the item-response-theory model fitted to the students' own diagnostic answers correctly predicts how they would have scored without the robot, and that students would not have switched answers merely from being asked again.","fun_headline_variants_meta":{"raw":{"variants":["Certain robots sway 94% of students, even when wrong","Robot confidence boosts student compliance from 71% to 94%","AI-savvy students are more easily misled by robots","Confident robots convince students, reducing critical thinking","Students follow wrong robot answers when it acts sure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1708,"prompt_tokens":1053,"completion_tokens":655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":574}},"tokens_in":669,"tokens_out":655,"duration_ms":7309,"temperature":1.0,"reasoning_tokens":574,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:48:23.076518+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a matched group of students the same eight true/false questions twice with no robot, or with a recorded voice instead of a present robot. If the switch rate and the rate of beyond-expected performance are as high as in the robot condition, the paper's causal reading fails; if they are substantially lower, the robot-influence claim survives.","supporting_citations":[{"cited_title":"Furhat: a back-projected human- like robot head for multiparty human-machine interac- tion","cited_arxiv_id":null,"evidence_quote":"Supplies the back-projected social robot head used in the experiment, whose expressive face and speech carry the certainty manipulations."},{"cited_title":"Oliver and Boyd, Edinburgh, UK, 1925","cited_arxiv_id":null,"evidence_quote":"Provides Fisher's method used to combine p-values across Deception and non-Deception questions in the beyond-expectations analysis."},{"cited_title":"A meta-analysis of factors affecting trust in human-robot interaction.Human factors, 53(5):517–527, 2011","cited_arxiv_id":null,"evidence_quote":"Grounds the informational-trust mechanism, since reliability and failure rates determine trust in human-robot interaction."},{"cited_title":"Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence","cited_arxiv_id":null,"evidence_quote":"Supports the prosodic manipulation of certainty through speech rate and disfluency effects on perceived confidence."},{"cited_title":"A systematic review of experimental work on persuasive social robots.International Journal of Social Robotics, 14(6):1339–1378, 2022","cited_arxiv_id":null,"evidence_quote":"Establishes the review-based claim that persuasive robots mostly use peripheral cues and that argument-based persuasion is understudied."},{"cited_title":"Springer, 1986","cited_arxiv_id":null,"evidence_quote":"Supplies the Elaboration-Likelihood Model used to interpret why flawed arguments and certainty cues persuade or fail to persuade."},{"cited_title":"I am definitely certain of this! towards a multimodal repertoire of signals com- municating a high degree of certainty","cited_arxiv_id":null,"evidence_quote":"Defines multimodal signals of certainty used to build the robot's facial-expression cues for the Certain and Uncertain conditions."}],"review_version":1}