{"id":"47ded66b-6463-4534-97b6-d772ecee1042","arxiv_id":"2412.10267","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across six advocacy lessons, multiple-choice practice questions produced posttest scores statistically indistinguishable from open-response or combined practice, in about one-third less time.","lead":"A randomized experiment with 234 tutors found no significant learning differences between multiple-choice, open-response, or combined practice questions in six online advocacy lessons, while multiple-choice took the least time. The result suggests well-designed multiple-choice practice can be as effective as open responses for tutor training at scale, with GPT-based autograding of open responses showing mixed accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'as effective' claim relies on a composite posttest score; without subscale analyses, the assertion that MCQ practice transfers to open-ended performance is unsupported.","rationale":"The reader's weakest assumption focuses on the independence of repeated lesson completions from the same tutor. That is a real methodological concern about standard errors, but it is unlikely to change the substantive conclusion: the learning outcome F is extremely small (F=0.27), so even a design-effect correction would almost certainly leave the result non-significant; the time efficiency difference is large enough to remain significant. The subscale aggregation issue is more load-bearing because it directly threatens the paper's headline interpretation. The paper explicitly claims that MCQ practice transfers to open-ended performance, yet the only evidence is a null result on a composite of MCQ and open-response posttest items. Without reporting condition effects on the open-response posttest items alone, the paper cannot distinguish between true transfer and a format-specific trade-off that happens to cancel in the aggregate. This is not a matter of over-interpreting a null result generally; it is a missing analysis for a specific claim that the paper itself advances. The concrete test is straightforward with the existing data and would settle the concern. Because the paper otherwise reports a randomized design, careful time measurement, and a useful efficiency finding, the appropriate verdict remains conditional on the additional subscale analysis, matching the reader's CONDITIONAL verdict. I therefore set verdict_should_be to UNCHANGED, meaning the reader's verdict is not altered, though the rationale now includes this additional condition.","tokens_in":15585,"tokens_out":10155,"duration_ms":78632,"concrete_test":"Re-run the RQ1 factorial ANOVA (condition, lesson, and their interaction) using as the outcome (a) accuracy on the two open-response posttest items only and (b) accuracy on the two MCQ posttest items only, using the same data and model as Section 4.1. If the condition main effect is significant on the open-response subscale, or if its confidence interval excludes a pre-specified equivalence bound, then the claim that MCQ practice transfers to open-ended performance is not supported. If both subscale analyses show null condition effects (or CIs within an equivalence margin), the central claim gains direct support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MCQs are 'as effective' as open-response practice, and specifically that 'MCQ practice transfers to open-ended performance' (Section 6), is based on a null ANOVA for the total posttest score (F(2,717)=0.27, p=.765, Section 4.1). The posttest is composed of two multiple-choice and two open-response items, yet the paper reports only the aggregated accuracy and never analyzes the open-response and MCQ posttest items separately. If the MCQ condition performs better on MCQ posttest items and worse on open-response items, the composite could show no difference even though the transfer claim is false. This is a concrete, testable omission: the paper's own theoretical conclusion ('MCQ practice transfers to open-ended performance') requires evidence on the open-ended subscale, which the current analysis does not provide. The clustering concern raised by the reader is valid but secondary; even with cluster-robust inference, the overall null is likely to persist because F=0.27 is far from significance, whereas the subscale issue could reverse the central claim if a format-specific trade-off exists. The paper's recommendation to substitute MCQs for open-response practice in tutor training depends on transfer to open-ended outcomes, so this missing analysis is the most load-bearing weakness.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a posttest-only randomized experiment comparing three learning-by-doing conditions---multiple-choice questions only, open-response questions only, and a combination of both---embedded in six scenario-based tutor-training lessons on advocacy. With 234 tutors contributing 790 lesson completions, the authors find no significant main effect of condition on a composite posttest score (F(2,717)=0.27, p=.765), a significant condition-by-lesson interaction (F(10,717)=2.20, p=.012), and significantly shorter instruction time in the MCQ-only condition (about 29% less than open-response only and 35% less than both). The paper also evaluates GPT-4o and GPT-4-turbo as automated graders of open responses, reporting accuracy between 57% and 91% depending on lesson and question type, and argues that these models are adequate for low-stakes assessment. The authors conclude that MCQs are as effective as, and more efficient than, open-response tasks for learning when practice time is limited.","tokens_in":15819,"tokens_out":4604,"duration_ms":39391,"significance":"If the equivalence claim were properly supported, this would be a practically valuable result for scalable tutor training and for the broader debate about whether MCQ practice transfers to open-ended performance. The randomized design, the use of six content lessons in a less-structured domain, the time-on-task measurement, the reporting of human inter-rater reliability, and the stated commitment to sharing data, rubrics, and prompts are genuine strengths. The LLM-grading comparison is also relevant to the growing use of generative AI in low-stakes assessment. However, the central 'as effective' claim currently rests on a non-significant p-value rather than on an equivalence test or an effect-size bound, and the missing subscale analysis leaves the transfer-to-open-ended claim unsupported. These are load-bearing gaps that can be addressed with additional analyses.","major_comments":[{"comment":"The central claim that MCQs are 'as effective' as open-response practice rests on a null ANOVA result (F(2,717)=0.27, p=.765), but no equivalence test, effect-size estimate, or confidence interval for the condition differences is reported. A non-significant p-value cannot support the assertion of equivalent outcomes; report a TOST procedure or a confidence interval for the mean differences and specify the smallest effect size the design can rule out. The significant condition-by-lesson interaction (F(10,717)=2.20, p=.012) also needs to be addressed before a general equivalence claim is made.","section":"4.1, 5.1"},{"comment":"The posttest score aggregates two multiple-choice and two open-response items, yet the paper reports only this composite and then concludes in Section 6 that 'MCQ practice transfers to open-ended performance.' That conclusion requires an analysis of the open-response subscale separately (and ideally the MCQ subscale as well); otherwise a format-specific trade-off (MCQ practice helping MCQ posttest items but not open-ended items) could produce the same null composite. Please add subscale analyses by condition and lesson, with the same equivalence/effect-size evidence requested for the composite.","section":"3.4.1, 4.1, 6"},{"comment":"The 790 lesson completions come from only 234 tutors, with many tutors completing multiple lessons, but the ANOVA in Section 3.4 treats all observations as independent ('all factors being between subjects'). If tutor-level ability or motivation is correlated across lessons, the standard errors and p-values for both posttest performance and completion time are misestimated. Report the distribution of lessons per tutor, the intraclass correlation, and a mixed-effects model with a random tutor intercept (or cluster-robust standard errors) to verify the null result and the time differences.","section":"3.1, 3.4"},{"comment":"The claim that GPT models 'demonstrate proficiency' for low-stakes assessment is not well supported by Tables 6 and 7, which include AUC values as low as 0.17 (GPT-4-turbo, predict, Helping Students Manage Inequity) and 0.43-0.45 in several other lesson-by-question-type cells. Moreover, the human rubrics were developed by the same research team and the prompts were iteratively refined on the same responses, so the reported accuracy is a development-sample estimate. Please report the range of AUC/F1 as evidence of variability and clarify whether any holdout or cross-validation procedure was used.","section":"4.3, 5.3"}],"minor_comments":[{"comment":"The sample size is reported inconsistently: the abstract and Section 3.1 say 234 tutors, while Section 6 says n=235.","section":"Abstract, Section 6"},{"comment":"The caption states 'no overall significant differences in instruction time prior to posttest were found between the conditions,' which contradicts the significant main effect of condition reported in Section 4.2 (F(2,716)=12.56, p<.001).","section":"Figure 3 caption"},{"comment":"The text reports F(10,716)=13.46, p=.199 for the condition-by-lesson interaction on time; an F of 13.46 with those degrees of freedom would have a far smaller p-value, so this appears to be a typo (possibly F=1.346).","section":"4.2"},{"comment":"The first sentence is garbled: it says there was no significant interaction between instruction time and condition on learning outcomes and then says there was a significant interaction between condition and lesson; please rewrite to match Section 4.2.","section":"5.2"},{"comment":"The sentence saying GPT-4-turbo 'demonstrated proficiency across all lessons with poorer performance relative to the other lessons' for Helping Students Manage Inequity is self-contradictory and should be revised.","section":"4.3"},{"comment":"The observation that the Both condition 'took students less time than the sum of the open and MCQ conditions' is trivially true and not an informative result; consider removing or reframing it.","section":"4.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look, but the headline should be split in two. The efficiency finding is solid: randomized three-condition comparison, 790 lesson completions, and a 29–35% time saving for MCQ-only practice with no posttest difference. That is a practically useful result for tutor training at scale. The paper also ships lesson logs, rubrics, and prompts, which is exactly the kind of reproducibility this area needs. The LLM grading comparison is a reasonable addition, though the results are mixed.\n\nThe soft spot is the 'as effective' claim. F(2,717)=0.27, p=.765 is a null result, not evidence of equivalence. No equivalence test, no effect-size bound. That alone would be a moderate issue. But the stress-test note lands harder: the posttest is a composite of two MCQ and two open-response items, and the paper never reports the subscales separately. The Section 6 claim that 'MCQ practice transfers to open-ended performance' requires evidence on the open-ended items. Until that analysis exists, a format-specific trade-off could be hiding behind the aggregate. This is concrete, testable, and load-bearing. The authors should run the subscale ANOVA or mixed model before anyone takes the transfer claim seriously.\n\nThe clustering concern is real but secondary. Treating 790 completions from 234 tutors as independent observations likely overstates precision, especially for the time effects and the condition-by-lesson interaction. A mixed-effects model or cluster-robust errors should be standard here. For the main null, F=0.27 is so far from significant that clustering probably won't flip it, but the paper should still report the corrected inference. The significant condition-by-lesson interaction is handled honestly, if a bit hand-wavy; the authors admit the pattern lacks a clean explanation.\n\nThe LLM grader comparison has a milder version of the same fitting problem: prompts were iteratively refined on the same responses used for evaluation, and some AUC values are poor (0.17–0.45). Calling that 'proficiency' is generous. A held-out prompt development procedure would fix it.\n\nWho gets value from this: people designing tutor training, researchers comparing practice formats, and anyone thinking about LLM autograding for low-stakes assessment. The empirical direction is plausible and worth building on, but the central inference needs repair. This deserves peer review, not a desk reject, and the revision should require the subscale analysis, equivalence bounds, and clustered inference.","headline":"The time-saving result is real, but the 'as effective' claim rests on a null p-value and a composite posttest score that never tests the paper's own transfer claim.","tokens_in":16374,"tokens_out":1601,"would_cite":true,"duration_ms":572630,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replacing open-response practice with multiple-choice practice yields equivalent posttest learning outcomes while saving roughly a third of instruction time.","keywords":["multiple-choice questions","open-response questions","posttest-only randomized controlled trial","tutor training","learning-by-doing","generative AI assessment","GPT-4","learning efficiency"],"falsifier":"Re-running the same ANOVA on the released lesson log data with tutor as a random effect or with clustered standard errors would settle it: if the condition effect becomes statistically significant, or the 29-35% time savings shrink to non-significance, the core claims of equivalent learning and time efficiency would not withstand the correction.","tokens_in":15371,"feed_emoji":"🎓","tokens_out":5970,"duration_ms":50390,"temperature":0.7,"pith_summary":"The paper asks whether multiple-choice questions (MCQs) can serve as learning tasks, not just assessment tools, now that generative AI makes open-response grading scalable. In a posttest-only randomized trial with 234 tutors and 790 lesson completions across six scenario-based advocacy lessons, the authors compared three learning-by-doing conditions: MCQ only, open-response only, and both. They found no statistically significant difference in posttest performance across conditions ($F(2,717)=0.27$, $p=.765$), but the MCQ-only condition was significantly faster, taking about 29% less time than open-response-only and 35% less than the combined condition. The authors conclude that MCQs are as effective and more efficient than open-response tasks for learning when practice time is limited, and that GPT-4o and GPT-4-turbo can grade open responses with sufficient proficiency for low-stakes assessment. This matters because, if true, large-scale training and homework could shift toward MCQs without sacrificing learning outcomes.","feed_headline":"MCQ practice matches open-response learning at 29% less time","feed_subtitle":"A 234-tutor randomized trial found equal posttest scores, so time-saving MCQs may be a sound replacement for open response.","key_machinery":"The load-bearing mechanism is a posttest-only randomized design in which each of six scenario-based lessons embeds one of three learning-by-doing conditions: MCQ only, open-response only, or both, followed by identical instruction and a common posttest that mixes both question types. The comparison of posttest accuracy across conditions, via ANOVA with lesson, condition, and scenario order as between-subjects factors, combined with completion times from lesson log data, is what produces the equivalence-and-efficiency result. The auxiliary machinery is a prompt-engineering pipeline for LLM autograding, using few-shot examples, chain-of-thought rationale requests, temperature 0, and JSON output, which is what makes scalable open-response grading plausible.","core_discovery":"The central discovery, stated on the paper's own terms, is that practice format does not reliably change learning: posttest scores were statistically indistinguishable across MCQ-only, open-response-only, and combined conditions ($F(2,717)=0.27$, $p=.765$), despite a significant condition-by-lesson interaction ($F(10,717)=2.20$, $p=.012$) that the authors attribute mostly to random variability. The MCQ-only condition completed instruction in $M = 3.83$ minutes, versus $M = 5.38$ minutes for open-response-only and $M = 5.87$ minutes for both, yielding 29% and 35% time savings. The authors interpret this as evidence that MCQ practice transfers to open-ended posttest performance at least as well as open-response practice, and they explicitly note that this result is inconsistent with the ICAP framework's general prediction that constructive tasks (open response) should beat active tasks (MCQ with feedback). In the auxiliary LLM evaluation, GPT-4o and GPT-4-turbo, prompted with few-shot examples and chain-of-thought, achieved 71-91% accuracy on predict responses and 71-87% accuracy on explain responses, with notable exceptions such as an AUC of 0.17 for one lesson, indicating that LLM grading is usable but not yet universally reliable.","pith_inferences":["The paper's null result is consistent with the view that MCQ selection and open-response construction both require the learner to retrieve and evaluate the same underlying rule; what differs is the response mode, not the memory retrieval, which would explain the equivalent transfer.","Because the time savings came from the learning-by-doing phase, and follow-up instruction was actually shortest in the combined condition, a natural testable extension is to vary the ratio of MCQ to open-response items within a fixed practice budget to locate the efficiency frontier.","The 3 of 18 significant pairwise contrasts, versus roughly 1 in 20 expected by chance, could be probed by pre-registering lesson-specific hypotheses about which advocacy skills rely on constructed justification.","The public release of log data would allow a reanalysis with tutor-level random effects, which is exactly the precision check the design needs to rule out clustering artifacts."],"forward_implications":["If MCQs are as effective and faster, homework and tutor-training platforms could substitute MCQ practice for open-response practice without expecting learning losses, freeing learner time for other content.","The transfer of MCQ practice to open-ended posttest performance supports using MCQs as learning tasks even when the final assessment is open response.","The null result challenges the general ICAP prediction that constructive open-response tasks outperform active MCQ tasks, at least for advocacy content, and invites theory refinement.","The significant condition-by-lesson interaction suggests the equivalence may not be uniform across content; two lessons showed significant pairwise contrasts, so content-treatment interactions deserve direct study.","GPT-4o and GPT-4-turbo autograding, with accuracy mostly in the 71-91% range, could support low-stakes open-response assessment at scale, though the low-AUC cases mark where human review remains necessary."],"supporting_citations":[{"why":"Supplies the central debate about whether MCQs are effective learning tools or merely efficient assessment tools, framing the research question.","marker":"[3]"},{"why":"Provides the ICAP framework, whose prediction that constructive tasks should outperform active tasks is the theoretical baseline the null result contradicts.","marker":"[7]"},{"why":"Recent empirical comparison of multiple-choice versus fill-in problems, establishing the trade-off between scalability and learning that this study extends.","marker":"[16]"},{"why":"Establishes the learning-by-doing methodology on which the scenario-based lessons are built.","marker":"[23]"},{"why":"Prior demonstration that GPT models can score open-ended responses with few-shot prompting, the approach the LLM autograding pipeline builds on.","marker":"[25]"},{"why":"The authors' earlier scenario-based lesson study with a pre-post design, which grounds the lesson structure and confirms that the assessments can detect learning gains.","marker":"[35]"},{"why":"Chain-of-thought prompting technique used in the LLM scoring prompts to elicit rationale and improve grading accuracy.","marker":"[40]"}],"fun_headline_variants":["MCQ matches open-response learning, saves 29% time","Equal learning, 29% faster: MCQs beat open response in RCT","RCT: MCQs as effective as open response, but 29% quicker","Multiple choice matches open response with 29% time saved","MCQ practice equals open response, but takes 29% less time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis treats each of the 790 lesson completions as an independent observation, even though the same tutor often completed several lessons, meaning that if a tutor's skill or motivation links their lessons together, the reported error bars and p-values could be misestimated.","fun_headline_variants_meta":{"raw":{"variants":["MCQ matches open-response learning, saves 29% time","Equal learning, 29% faster: MCQs beat open response in RCT","RCT: MCQs as effective as open response, but 29% quicker","Multiple choice matches open response with 29% time saved","MCQ practice equals open response, but takes 29% less time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2306,"prompt_tokens":1056,"completion_tokens":1250,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":1156}},"tokens_in":672,"tokens_out":1250,"duration_ms":8346,"temperature":1.0,"reasoning_tokens":1156,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:59:46.514316+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the same ANOVA on the released lesson log data with tutor as a random effect or with clustered standard errors would settle it: if the condition effect becomes statistically significant, or the 29-35% time savings shrink to non-significance, the core claims of equivalent learning and time efficiency would not withstand the correction.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ICAP framework, whose prediction that constructive tasks should outperform active tasks is the theoretical baseline the null result contradicts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Recent empirical comparison of multiple-choice versus fill-in problems, establishing the trade-off between scalability and learning that this study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the learning-by-doing methodology on which the scenario-based lessons are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior demonstration that GPT models can score open-ended responses with few-shot prompting, the approach the LLM autograding pipeline builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The authors' earlier scenario-based lesson study with a pre-post design, which grounds the lesson structure and confirms that the assessments can detect learning gains."}],"review_version":1}