{"id":"04f6d204-e9c7-4eeb-a955-4737e385a35f","arxiv_id":"2505.22526","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An RCT reports higher perceived learner control with an AI lecturer than with a human or MOOC, but the learning outcome claim is weakened by inconsistent reporting and post hoc data choices.","lead":"A randomized trial of 125 university students found that those taught by an AI instructional agent reported greater perceived learner control than students taught by a human lecturer or through a MOOC with a chatbot. The paper also claims better test scores, but its own tables do not fully support the comparison against the MOOC group.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PLC scale was constructed and items retained post hoc on the same 125 cases; without measurement invariance across the three arms, the primary outcome is not interpretable as a real group difference.","rationale":"The reader already identified the post-hoc PLC scale as the weakest assumption, and I agree. An invariance test is the decisive check: without it, the primary outcome could be an artifact of item wording and differential interpretation across conditions. The post-test table/text mismatch is real but secondary because the learning-outcome claim versus the human condition survives the correction, whereas the PLC finding is the paper's headline contribution and rests entirely on an unvalidated, post-hoc scale. The post-hoc exclusion of 15 participants and the absence of pre-registration increase the risk of optimism but do not change the central vulnerability. The conditional verdict is appropriate pending the proposed invariance analysis; if the test fails, the primary outcome would be unsupported and the verdict would need to move toward rejection or major revision.","tokens_in":13133,"tokens_out":7089,"duration_ms":90886,"concrete_test":"Obtain the raw item-level PLC responses and run a multi-group confirmatory factor analysis across the three conditions, testing configural, metric, and scalar invariance (e.g., with lavaan in R). If scalar invariance is rejected (conventionally ΔCFI ≤ -0.01 with ΔRMSEA ≥ 0.015), re-estimate the AI-versus-human contrast under partial invariance or with an ordinal item model; if the difference is no longer significant or changes sign, the perceived-learner-control finding is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the four-item Perceived Learner Control scale measures the same latent construct in the human, MOOC, and AI conditions. Section 3.3 says the scale was constructed and the items were retained only after data collection, on the same 125 participants, with no confirmatory factor analysis or measurement-invariance test. The retained items ('decide how to participate', 'decide the pace', 'control the way I learn', 'decide how to allocate my time') nearly restate the features that distinguish the AI condition, so a higher AI-group mean could reflect differential item interpretation or response shift rather than a higher standing on a common PLC construct. If scalar invariance does not hold, the reported ANOVA F(2,122)=12.155 and the Tukey contrasts are not interpretable as differences in perceived control, and the headline claim loses its primary outcome. The post-test reporting mismatch (Table 2 shows MOOC-human, not AI-MOOC, while Section 4.1 reports an AI-MOOC difference) is a separate, correctable reporting concern, but it does not bear on the PLC result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a randomized controlled trial comparing three instructional conditions for a university general-education course: a human teacher, a self-paced MOOC with a separate chatbot, and an AI instructional agent integrated with the MAIC platform. The authors analyze data from 125 students (after excluding 15 with near-uniform response patterns) and report that the AI-agent condition produced significantly higher perceived learner control (PLC) than both the human and MOOC conditions, higher interaction frequency, shorter learning duration, and—according to the Section 4.1 text—significantly higher posttest scores than both comparison groups. A regression model is used to argue that PLC predicts posttest performance. The paper concludes that AI instructional agents, when designed to combine lecture delivery with real-time interactive responsiveness, can improve both students' subjective sense of control and immediate learning outcomes.","tokens_in":13346,"tokens_out":4768,"duration_ms":57552,"significance":"If the findings withstand scrutiny, the paper provides valuable experimental evidence on a timely question: whether integrated AI instructional agents can improve perceived learner control and performance over traditional instruction and over less-integrated online learning. The RCT design, the use of identical lecture scripts and audio across conditions, and the inclusion of behavioral indicators are notable strengths. However, the primary outcome rests on a scale constructed and item-selected after data collection, with no measurement-invariance evidence, and the posttest performance claim is contradicted by the reporting in Table 2. These issues are load-bearing for the paper's central claims, so the significance can only be realized after those concerns are resolved.","major_comments":[{"comment":"The text in Section 4.1 states that 'Students in the AI group achieved significantly higher post-test scores than those in the human teacher group (M difference = 0.128, p < .01), and also outperformed the MOOC group (M difference = 0.112, p < .01).' Table 2, however, reports the posttest pairwise differences as 'MOOC-human 0.112 **' and 'AI-human 0.128 **', with no AI-MOOC row. These two presentations are mutually inconsistent: if the 0.112 difference is MOOC-human, then the text's claim that AI outperformed the MOOC is not supported by the reported pairwise contrasts. Because the posttest result is central to the paper's title and conclusion, the authors must correct this discrepancy, report the actual AI-MOOC contrast if it was tested, and adjust the claims in the abstract, Section 4.1, and Section 5 accordingly.","section":"4.1 / Table 2"},{"comment":"The Perceived Learner Control scale was newly developed for this study, and Section 3.3 states that 'After data collection and reliability and validity analysis, four items were retained.' This item selection was performed on the same 125 participants used for all subsequent analyses, and no confirmatory factor analysis or measurement-invariance test across the three experimental arms is reported. The four retained items—'decide how to participate,' 'decide the pace,' 'control the way I learn,' and 'decide how to allocate my time'—closely mirror the very features that distinguish the AI condition from the others. Without evidence that the scale measures the same construct in the human, MOOC, and AI groups, the ANOVA F(2,122)=12.155 and the Tukey contrasts for PLC cannot be unambiguously interpreted as differences in a common latent construct. The authors should report the full item-development process, test for measurement invariance, or substantially temper the primary-outcome claim.","section":"3.3"},{"comment":"The authors excluded 15 of the 140 participants (10.7%) because their responses showed over 90% identical ratings, and the final valid sample is 41, 43, and 41 for the human, MOOC, and AI groups, respectively. The exclusion is applied after random assignment, but the manuscript does not report how many excluded participants came from each condition. If exclusions are differential across groups, the randomization's integrity could be compromised. The authors should provide per-group exclusion counts and conduct sensitivity analyses that either retain all participants or apply alternative data-quality filters to confirm that the substantive findings are robust.","section":"3.4"},{"comment":"The comparison of learning duration between the human-teacher condition and the self-paced (AI/MOOC) conditions is problematic as evidence of efficiency. In the human condition, students attended a fixed-duration in-person lecture, whereas in the AI and MOOC conditions they studied self-paced videos and could finish at any time. The large duration difference (about 30 minutes) is therefore partly a design artifact rather than a behavioral indicator of more efficient time use. The statement that 'reduced learning time alone... may reflect more efficient time usage and self-management by learners' should be reanalyzed using only the two self-paced conditions, or else qualified to acknowledge the structural difference.","section":"4.1 / Table 2"}],"minor_comments":[{"comment":"The regression model specification lists the predictors 'perceived learner control, gender, age, srl, pretest' but then states 'β1 to β9 are the regression coefficients'; the notation should be β0 to β5 for consistency.","section":"3.4"},{"comment":"The posttest item retention (22 multiple-choice questions reduced to 16 after 'item response analysis') is also a post hoc selection on the same dataset; please report the criteria used to retain or exclude items.","section":"3.3"},{"comment":"The regression model in Section 4.2 does not include condition dummies. Because both PLC and posttest scores differ by condition, the reported coefficient for PLC (β = 0.055) may partially reflect condition effects; adding condition indicators or discussing this limitation would clarify the interpretation.","section":"4.2"},{"comment":"There are several formatting and typographical issues: 'ANOV A' appears instead of 'ANOVA' in Section 3.4; Figure 2 appears to have duplicated panels; and 'self-determined theory' in Section 2.1 should read 'self-determination theory.'","section":"3.3 / Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a topic of current interest and the RCT design is appropriate, but the primary outcome is undermined by the post hoc construction of the PLC scale without invariance testing, and the posttest claim is internally inconsistent between text and table. These are fixable within the manuscript's scope if the authors can supply the missing analyses or correct the claims. I would not recommend rejection at this stage, but the revisions need to be substantive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a genuine randomized comparison of a hybrid AI lecturer against a human teacher and a chatbot-supported MOOC, with the same slides, scripts, and even the same synthesized voice. That is worth something. The behavioral data are also plausible: the AI group asked more questions and finished faster. If the headline claims survive a careful revision, this is a useful datapoint for the AI-in-education literature.\n\nThe soft spots are real but not fatal. First, the perceived learner control scale was constructed and items were retained after data collection, using the same 125 participants. No confirmatory factor analysis, no measurement invariance across the three arms, and the four retained items essentially restate the features that distinguish the AI condition. A skeptic can reasonably say the AI group's higher mean reflects differential item interpretation rather than a higher standing on a common construct. That does not make the PLC result meaningless, but it does make it weaker than the paper claims. The stress-test note has this right.\n\nSecond, the posttest reporting is internally inconsistent. The text says AI outperformed the MOOC group by 0.112 (p < .01), but Table 2 shows the 0.112 difference is MOOC-human, with no AI-MOOC row at all. Given the table also reports AI-human as 0.128, the implied AI-MOOC difference is only 0.016, which would not be significant. One of those numbers is wrong. This is a correctable reporting error, but it is exactly the kind of thing that makes a reader doubt the abstract's strongest claim.\n\nMinor but worth noting: 15 participants were excluded post hoc for straightlining, with no sensitivity analysis; the regression coefficient for PLC is small (β = 0.055) and the within-group PLC-posttest correlations are weak or negative (human r = -0.09, AI r = 0.26, ns), so the mediation story is thinner than the prose suggests.\n\nWho is this for? Researchers in AIED and learning sciences who care about learner control and agent design. It deserves peer review because the design is serious and the question matters, but the revision must address the scale validation and the table-text mismatch before the claims are credible. I would not cite it in its current form.","headline":"A decently designed small RCT whose primary PLC measure was built post hoc on the same sample and whose posttest table contradicts the text, so the conclusions are real but conditional.","tokens_in":13810,"tokens_out":1801,"would_cite":false,"duration_ms":25070,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an AI instructional agent that delivers lectures and answers questions in real time raises students' perceived learner control and post-test scores above both a human-taught lecture and a MOOC-style video.","keywords":["AI instructional agent","perceived learner control","randomized controlled trial","learning outcomes","MAIC","self-paced learning","higher education"],"falsifier":"Run a larger replication that first validates the PLC scale with confirmatory factor analysis and tests measurement invariance across the three conditions, or uses an independently validated autonomy scale; if the AI group's PLC advantage disappears under invariant measurement, or if a pre-registered re-analysis with item-level response-style controls eliminates the group differences, the central claim fails.","tokens_in":12940,"feed_emoji":"🤖","tokens_out":9749,"duration_ms":93857,"temperature":0.7,"pith_summary":"This paper reports a randomized controlled trial in which 125 university students were assigned to one of three lecture formats: a live human teacher, a self-paced MOOC video with a separate chatbot, or an AI instructional agent that delivered the same lecture and answered questions in real time. The central claim is that the AI agent raised perceived learner control above both comparison groups and also produced higher immediate post-test scores. The authors argue this matters because a scalable AI system can combine structured lecture delivery with responsiveness, easing the usual trade-off between scale and personalization. They interpret the result through perceived learner control, which positively predicted post-test performance after controlling for pre-test score, gender, age, and self-regulated learning.","feed_headline":"AI teaching agent beats live lecture on learner control and test scores","feed_subtitle":"In an RCT, students taught by an AI agent felt more in control and scored higher than human or MOOC groups.","key_machinery":"The load-bearing mechanism is the AI instructional agent built on the MAIC platform: an LLM-driven virtual 'Teacher L.' that delivers lecture content with a synthesized voice and responds to students' real-time questions, letting them pause and control pacing. Because the PowerPoint slides, lecture scripts, and audio were identical across conditions, the design isolates the delivery-and-interaction mode as the source of group differences. The study also relies on the four-item Perceived Learner Control scale (pace, participation, method, and time allocation) as the subjective measure, supported by behavioral indicators—question frequency and learning duration—and by a regression model linking perceived control to post-test score.","core_discovery":"The paper's central discovery is that the mode of instructional delivery changes both felt autonomy and measured learning even when content, slides, and the teacher's voice are held constant. Students in the AI instructional agent group reported significantly higher perceived learner control than the human teacher group (M difference = 0.732, $p < .001$) and the MOOC group (M difference = 0.416, $p < .05$). They also scored higher on the 16-item post-test than both the human teacher group (M difference = 0.128, $p < .01$) and the MOOC group (M difference = 0.112, $p < .01$). Behavioral data aligned with the subjective reports: the AI group asked more questions and finished the learning task faster, and perceived learner control predicted post-test scores in a regression that controlled for prior knowledge and background ($\\beta = 0.055$, $p < .05$).","pith_inferences":["Because the human-teacher condition had a fixed classroom period, part of the AI group's shorter completion time is structural; a self-paced recorded-human condition would separate the freedom to stop early from the agent's interactivity.","The PLC scale was constructed and pruned after data collection, so the clearest confirmation of the subjective-control claim is a replication using a pre-validated autonomy scale or a formal measurement-invariance analysis across the three conditions.","If perceived control is the active mechanism, then independently varying the agent's pause, question, and pacing features should reproduce or split the observed effects; this is a direct design experiment the paper does not run.","The test-score differences are modest in proportion-correct units and measured only immediately after the lesson; whether they compound over a full course or vanish with harder, more discussion-heavy material is an open empirical question."],"forward_implications":["An AI agent that delivers lectures and fields questions in real time can produce higher perceived learner control than a live human lecturer or a self-paced MOOC while content, slides, and voice are held constant.","The same agent produced higher immediate post-test scores than both comparison conditions (mean difference = 0.128 vs. human teacher, 0.112 vs. MOOC, both $p < .01$).","Students who felt more in control completed the task faster and asked more questions, and perceived learner control predicted post-test score after controlling for pre-test, gender, age, and self-regulated learning.","The advantages are demonstrated for immediate outcomes in a moderate-difficulty lecture course; the paper does not claim long-term retention, transfer, or effects in problem-based or inquiry courses."],"supporting_citations":[{"why":"Introduces the MAIC platform and its LLM-driven lecture agents, the intervention technology tested in this study.","marker":"Yu et al., 2024"},{"why":"Defines perceived learner control and grounds the scale's focus on pace, sequence, and content.","marker":"Kraiger & Jerden, 2007"},{"why":"Provides the multidimensional account of learner control that informs the construction of the PLC measure.","marker":"Karim & Behrend, 2014"},{"why":"Supplies the three-dimensional engagement scale used to measure cognitive, emotional, and behavioral engagement.","marker":"Reeve & Tseng, 2011"},{"why":"Offers a comparison case of AI-generated lecture video without real-time interaction, which found no performance gain over human lecturers.","marker":"Arkün-Kocadere & Özhan, 2024"},{"why":"Shows that interactive presence in instructional videos improves comprehension, the design hypothesis the AI agent operationalizes.","marker":"Wu et al., 2024"},{"why":"Documents that unguided learner control can hurt novices, motivating the structured-control design of the agent.","marker":"Hasler, Kersten & Sweller, 2007"},{"why":"Provides the self-determination theory account that perceived autonomy drives motivation and effort.","marker":"Ryan & Deci, 2020"}],"fun_headline_variants":["AI agent heightens learner control and lifts test scores in RCT","RCT: AI teaching agent tops lecture and MOOC on control and learning","AI instructor raises perceived control and post-test performance","More learner control, better results with AI teaching agent","AI agent study: students feel more in control, score higher"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The four-item Perceived Learner Control scale was constructed and pruned only after data collection, with no reported measurement-invariance check or external validation showing the items capture the same construct across the human, MOOC, and AI conditions; if the scale measures different things in different modes, the reported PLC differences could be artifacts of the measure.","fun_headline_variants_meta":{"raw":{"variants":["AI agent heightens learner control and lifts test scores in RCT","RCT: AI teaching agent tops lecture and MOOC on control and learning","AI instructor raises perceived control and post-test performance","More learner control, better results with AI teaching agent","AI agent study: students feel more in control, score higher"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1354,"prompt_tokens":887,"completion_tokens":467,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":384}},"tokens_in":503,"tokens_out":467,"duration_ms":5512,"temperature":1.0,"reasoning_tokens":384,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:04:13.299535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger replication that first validates the PLC scale with confirmatory factor analysis and tests measurement invariance across the three conditions, or uses an independently validated autonomy scale; if the AI group's PLC advantage disappears under invariant measurement, or if a pre-registered re-analysis with item-level response-style controls eliminates the group differences, the central claim fails.","supporting_citations":[{"cited_title":"& Özhan, ¸ S","cited_arxiv_id":null,"evidence_quote":"Offers a comparison case of AI-generated lecture video without real-time interaction, which found no performance gain over human lecturers."}],"review_version":1}