{"id":"31ad04a9-5a48-4614-8efa-a9df42641e34","arxiv_id":"1908.01028","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-year teaching case study found that most first-year mathematics students perceived online proof-validation quizzes as helpful for engaging with lecture notes and noticing small details in proofs.","lead":"A first-year university mathematics course added bi-weekly online quizzes to its assignments, and students reported that the quizzes helped them start studying earlier and think about proof details. The paper is a concrete case study on using e-assessment as a nudge toward regular, deeper engagement with pure mathematics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on self-report surveys with 59%/37% response rates and no objective outcome; logged BlackBoard behaviour would test whether perceived engagement matches actual engagement.","rationale":"The reader's weakest assumption identifies self-report validity as the key vulnerability, and I agree: the conclusion that the quizzes 'successfully achieved' their goals is supported mainly by perception items, not by direct measurement of engagement or proof behaviour. The low and unequal response rates (59% and 37%) create a real risk of non-response bias, and the paper's own objective indicator—unchanged average marks—does not corroborate the stronger language in the abstract. The paper is honest about many of these limitations, and the quiz design itself is thoughtful, so this is not a reason to reject the work. It does, however, justify keeping the verdict at CONDITIONAL rather than ACCEPT, because a straightforward log-based check could substantially raise or lower confidence in the central claim. My read does not move the reader's verdict, so I set verdict_should_be to UNCHANGED while noting the specific test that would settle the concern.","tokens_in":12304,"tokens_out":5389,"duration_ms":60972,"concrete_test":"Request BlackBoard server logs for both years. For every enrolled student, record: date/time of first quiz access relative to material release; number of quiz attempts; whether feedback pages were opened after each deadline; score on proof-validation questions; and final exam script or relevant exam question scores. Then test three things: (1) whether survey respondents differ from non-respondents on these logged behaviours; (2) whether self-reported 'read feedback' matches logged feedback views; (3) whether logged feedback use predicts quiz proof-validation scores or exam proof-detail performance after controlling for prior attainment (e.g., A-level grade). If logs agree with self-reports and predict outcomes, the central claim is strengthened; if they diverge, the questionnaire-only conclusion should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 6 ('the online quizzes successfully achieved their intended goals') rests entirely on end-of-module questionnaires analysed in Section 4. The response rates are 59% (Year 1) and 37% (Year 2), and the items measure students' perceptions ('found the quizzes helped them to engage', 'thought that the online quizzes helped them to pay attention to the small details'), not actual engagement or proof behaviour. If responders are more engaged than non-responders, or if students over-report perceived benefit, the headline conclusion is unsupported. The paper's own objective evidence cuts the other way: Section 5 reports that average student marks were 'not affected' in either comparison, and only 22-23% of students said feedback helped with subsequent written assignments. The Year 2 decrease in the small-details item (69% to 59%) and student comments such as 'It's quite easy to spot mistakes but much harder to write a whole correct proof from scratch' further weaken the claim that the quizzes successfully emphasised proof details. Because the quizzes were the only assessment change and were also summative (10%), the design cannot separate e-assessment effects from periodic-assessment or grading effects. The study is a useful case description, but the central claim should be treated as conditional pending objective behavioural data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a two-year case study in which bi-weekly online quizzes (10% of the final grade, with automatic feedback) were introduced into a first-year pure mathematics module, 'Introduction to Proofs.' The quizzes were intended to motivate early engagement with learning material and to draw attention to small details in proofs. The evaluation draws on end-of-module student questionnaires (response rates 59% and 37%) and on the author's teaching observations. The paper reports that majorities of respondents perceived the quizzes as supporting engagement and proof-detail awareness, and it concludes that the quizzes 'successfully achieved their intended goals.'","tokens_in":12508,"tokens_out":5054,"duration_ms":45058,"significance":"If the central claim could be supported by objective evidence, the paper would provide a useful model for integrating e-assessment into proof-based instruction, an area where automated assessment is rare. The manuscript's concrete strengths are its detailed and replicable description of quiz design, worked examples of questions and tailored feedback, and a thoughtful set of practical recommendations in Section 6. Its principal weakness is that the evidence is entirely self-reported, with low response rates, no control group, and no objective measure of engagement or learning; the paper's own exam-mark data show no change. The contribution is therefore best read as a practitioner case description rather than a strong empirical demonstration, though the implementation details are valuable to the teaching community.","major_comments":[{"comment":"The central conclusion in Section 6 rests on student self-reports from end-of-module questionnaires with response rates of 59% (Year 1) and 37% (Year 2). Because non-respondents may differ systematically from respondents, and because self-reports of perceived benefit are not measures of actual engagement or proof-writing behaviour, the data as presented cannot support the unqualified claim that the quizzes 'successfully achieved their intended goals.' The manuscript should either present objective behavioural records (such as BlackBoard logs of quiz attempts and feedback views) or explicitly restrict the conclusion to perceived benefits.","section":"Section 4, response rates"},{"comment":"The paper reports that average student marks were 'not affected' by the quizzes in either comparison. This objective outcome is acknowledged but not weighed in the Conclusion. For a claim that quizzes improved learning or proof-detail awareness, unchanged exam marks are a relevant contrary indicator. The Conclusion in Section 6 should either confront this evidence directly or temper the claim to state that the perceived benefits were achieved even though measurable performance did not change.","section":"Section 5, exam marks"},{"comment":"In Year 2, online quizzes were introduced in all but one other first-year mathematics module at the institution. The attribution of the Year 2 increases in the engagement items (72% versus 55% and 75% versus 61%) to the author's 10% change in question type is therefore confounded by a system-wide change in assessment practice. The simultaneous 10-percentage-point drop in the 'small details' item (69% to 59%) is also not explained; this pattern is difficult to reconcile with the conclusion that both goals were achieved in both years.","section":"Section 4, Year 2 confound"},{"comment":"The manuscript states that 'The data presented in the previous chapter confirms that both of these goals were successfully achieved.' This is an overstatement: the questionnaire items measure students' perceptions ('thought that the online quizzes helped them...'), not demonstrated engagement or proof-detail behaviour. The language in the Discussion and Conclusion should be revised throughout to distinguish perceived benefit from demonstrated effect.","section":"Section 5, 'confirms' language"}],"minor_comments":[{"comment":"The sentence 'Figure 2 shows that 69% of students in Year 1 thought that the online quizzes helped them to pay attention to the small details when writing proofs' should refer to Figure 3, which is the figure captioned 'Emphasis of the small details within proofs.'","section":"Section 4, figure references"},{"comment":"The text states '69% of students in Year 1 and 77% of students in Year 2 indicated that the feedback helped them to understand why their particular answer choice was incorrect,' but Section 4 reports 68% for Year 1, not 69%.","section":"Section 5, numeric inconsistency"},{"comment":"The text states 'only 59% of students, who had read the available feedback, claimed that the quizzes helped them to understand the small details when writing proofs,' but Section 4 reports this statistic as 56%. The discrepancy should be resolved.","section":"Section 5, further numeric inconsistency"},{"comment":"The phrase 'feedback seems to be the one of the main tools' should read 'one of the main tools.'","section":"Section 2, typo"},{"comment":"The claim that the effect of guessing was reduced because 'not a single student commented on being able to guess answers in Year 2' is an inference from absence of evidence; it should be presented as a weaker, anecdotal observation.","section":"Section 5, inference from absence of comments"}],"recommendation":"major_revision","confidential_remarks":"This is a practitioner case study whose evidence base is too weak for the strong empirical claim made in the conclusion. If the journal publishes teaching-focused case studies, major revision is appropriate; the authors should reframe the conclusion as reporting perceived benefits and explicitly discuss the confounds and self-report limitations. The author's dual role as module designer, instructor, and evaluator is a structural source of bias that should be acknowledged in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a straightforward teaching case study: the author designed bi-weekly Blackboard quizzes for a first-year proofs module, collected student feedback over two years, and reports that the quizzes motivated early engagement and attention to proof details. The good part is genuine. The quiz items are carefully built around known proof-writing pitfalls (unstated notation, wrong induction base case, proving only one direction), and the exemplar feedback is specific and pedagogically sound. The paper is also transparent: it reports response rates, quotes negative comments, and admits exam marks did not move.\n\nThe soft spot is the gap between the conclusion and the evidence. The headline claim that the quizzes 'successfully achieved' their goals rests on end-of-module questionnaires with 59% and 37% response rates, with no control group, no statistical analysis, and no objective measure of engagement or proof-writing skill. The author's own data cut the other way on the details point: the small-details item fell from 69% to 59% in Year 2, and only 22–23% of students said feedback helped with subsequent written assignments. Student comments like 'quite easy to spot mistakes but much harder to write a whole correct proof' suggest proof validation is not transferring to proof construction, which the author acknowledges as a possible 'gap.' On exam marks, the author reports no noticeable difference, which is honest but weakens the 'powerful complementary tool' framing.\n\nI don't see misconduct or even carelessness. This is the normal shape of a single-instructor case study in education research. The Year 2 comparison is confounded by the introduction of quizzes in other modules, and the self-report assumption is real but not fatal given the modest claims. The author is careful in the Discussion to hedge some interpretations. The abstract and Conclusion could be toned down to say 'students perceived' rather than 'successfully achieved.'\n\nBottom line: the paper is a useful, honest case description with good materials for practitioners. The evaluation evidence is thinner than the conclusion suggests. I'd send it to peer review with a request to soften the claims and, ideally, to add whatever Blackboard log data exists (attempt counts, time-on-task) to back the engagement claim. The citation pattern and literature review are solid. Not groundbreaking, but a legitimate contribution for the math education audience.","headline":"Honest, well-designed case study of proof-validation quizzes whose central claim outruns the self-report evidence; still worth a serious referee.","tokens_in":13006,"tokens_out":2013,"would_cite":false,"duration_ms":19711,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports evidence that periodic open-book online quizzes, counting for 10 percent of the grade, give first-year proof students an extrinsic motivation to engage early with lecture notes and make the small details of proofs hard…","keywords":["e-assessment","mathematical proofs","undergraduate mathematics education","online quizzes","assessment for learning","proof validation","student feedback","transition to university mathematics"],"falsifier":"In a matched comparison of two first-year proof modules, one with bi-weekly open-book quizzes and one without, a blind-scored end-of-term proof question would show whether quiz students produce more rigorous proofs (all notation defined, all cases covered). If the quiz cohort's proofs are no more rigorous, the claim that the quizzes emphasized small proof details is not supported.","tokens_in":1430,"feed_emoji":"📝","tokens_out":1362,"duration_ms":79255,"temperature":0.7,"pith_summary":"The paper reports a two-year case study in which bi-weekly, open-book online quizzes were added to a first-year pure mathematics module, Introduction to Proofs, alongside written homework and an end-of-term exam. The quizzes were deliberately built as assessment for learning rather than assessment of learning: they gave automatic, individually tailored feedback and tested the kind of proof details that written homework often misses, such as defining notation, covering all cases, and forming correct contrapositives. Based on end-of-module questionnaires, the author argues that the quizzes successfully gave students an extrinsic reason to engage with lecture notes early and sharpened their attention to small details within proofs, while noting that only a minority of students connected the quizzes to their own proof writing. If this is right, e-assessment is not a replacement for written homework but a workable complement to it for large first-year proof modules.","feed_headline":"Quiz-based assessment can teach proof habits early","feed_subtitle":"Two years of survey data show periodic online tests boost early engagement and attention to proof details","key_machinery":"The core mechanism is the proof-validation question: a multiple-choice, matching, or fill-in-the-blank item that presents a proof and asks students to judge its correctness or arrange its steps. Because the chosen answer is right or wrong as a whole, one missing detail—say, failing to define x ∈ Z before using divisibility—collapses an otherwise valid-looking proof. The automatic feedback then states exactly which detail broke the proof and how to repair it, sometimes changing N to Z or inserting one line. This couplet of all-or-nothing proof validation plus tailored repair feedback is what carries the paper's argument.","core_discovery":"The central discovery is that assessment for learning in pure mathematics does not have to wait for written proofs. In the author's words, the online quizzes successfully achieved their intended goals: they gave first-year students an extrinsic motivation to engage with lecture notes early, and they made the small details inside proofs—defining variables, covering all cases, forming contrapositives—difficult to ignore. Survey data from two cohorts showed 55 percent then 72 percent of students crediting the quizzes with early engagement with lecture notes, 69 percent then 59 percent crediting them with sharper attention to small details when writing proofs, and feedback-reading rising from 44 percent to 64 percent after the author repeatedly taught students where to find it. The author also reports a persistent gap: only about a fifth of students felt quiz feedback helped them write their own proofs later, which she explains as a transfer gap rather than a failure of the quizzes.","pith_inferences":["An implication the author leaves implicit: the transfer gap could be attacked directly by adding a short free-response follow-up to each quiz in which students fix the flawed proof they just evaluated; the feedback examples already provide the repair language.","A plausible extension: systematically vary the proportion of multiple-choice versus fill-in-the-blank and matching items to find the mix that keeps guessing low without sacrificing the small-detail emphasis that all-or-nothing validation provides.","A natural corroborating study: use platform logs of quiz attempts, submission timing, and feedback clicks to measure early engagement directly instead of relying on end-of-module self-reports."],"forward_implications":["Students report engaging with lecture notes earlier in both years of quizzes than in the year before, with the second-year cohort reporting stronger engagement after repeated feedback instructions.","All-or-nothing proof-validation questions make a single omitted detail—such as an undefined variable or an unexamined case—break the whole proof, which is exactly the lesson written homework often leaves implicit.","Changing 10 percent of quiz questions from multiple-choice to fill-in-the-blank and matching forms eliminated student complaints about guessing and coincided with higher reported engagement with lecture notes.","Feedback reading rose from 44 percent to 64 percent after the instructor repeatedly showed students where feedback lived, indicating that engagement with feedback is teachable.","The persistent gap between validating proofs and writing proofs (only about 22–23 percent of students credited quizzes with helping their own written work) tells the author that transfer needs to be taught explicitly, not assumed."],"supporting_citations":[{"why":"supplies the design principles (encoding common misconceptions, random parameters, tailored feedback) that the quizzes follow.","marker":"[12]"},{"why":"grounds the claim that proof validation exercises can improve students' proof writing.","marker":"[26]"},{"why":"defines proof validation as reading and reflecting on proofs, the activity the harder quiz items demand.","marker":"[30]"},{"why":"finds that regular online testing improves learning in numerical sciences, the prior result this study extends.","marker":"[1]"},{"why":"shows formative online quizzes can raise summative exam scores, supporting the pedagogical mechanism.","marker":"[7]"},{"why":"supports the open-book format by showing it lets students answer by reading and thinking rather than memorizing.","marker":"[15]"},{"why":"supports periodic scheduling as a way to keep students continuously engaged with material.","marker":"[29]"},{"why":"establishes feedback as a critical component of effective e-assessment.","marker":"[9]"}],"fun_headline_variants":["Online quizzes nudge first-year math students to start early and inspect proof details","E-quizzes drive early study and proof-detail focus in pure math","Quizzes as learning: early engagement, sharper proof details","Two-year case study: e-quizzes boost early proof habits"],"cache_read_input_tokens":15232,"weakest_assumption_plain":"The central claim rests on students' self-reported survey answers, from 59% and 37% response rates, accurately reflecting their engagement and attention to proof details.","fun_headline_variants_meta":{"raw":{"variants":["Online quizzes nudge first-year math students to start early and inspect proof details","E-quizzes drive early study and proof-detail focus in pure math","Quizzes as learning: early engagement, sharper proof details","Two-year case study: e-quizzes boost early proof habits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001531,"raw_usage":{"total_tokens":6063,"prompt_tokens":814,"completion_tokens":5249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":5174}},"tokens_in":430,"tokens_out":5249,"duration_ms":35084,"temperature":1.0,"reasoning_tokens":5174,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:24:44.649143+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a matched comparison of two first-year proof modules, one with bi-weekly open-book quizzes and one without, a blind-scored end-of-term proof question would show whether quiz students produce more rigorous proofs (all notation defined, all cases covered). If the quiz cohort's proofs are no more rigorous, the claim that the quizzes emphasized small proof details is not supported.","supporting_citations":[{"cited_title":"3, 117–137","cited_arxiv_id":null,"evidence_quote":"supplies the design principles (encoding common misconceptions, random parameters, tailored feedback) that the quizzes follow."},{"cited_title":"4, 501–514","cited_arxiv_id":null,"evidence_quote":"grounds the claim that proof validation exercises can improve students' proof writing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines proof validation as reading and reflecting on proofs, the activity the harder quiz items demand."},{"cited_title":"2, 255–272","cited_arxiv_id":null,"evidence_quote":"finds that regular online testing improves learning in numerical sciences, the prior result this study extends."},{"cited_title":"4, 297–302","cited_arxiv_id":null,"evidence_quote":"shows formative online quizzes can raise summative exam scores, supporting the pedagogical mechanism."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supports the open-book format by showing it lets students answer by reading and thinking rather than memorizing."},{"cited_title":"4, 423–431","cited_arxiv_id":null,"evidence_quote":"supports periodic scheduling as a way to keep students continuously engaged with material."},{"cited_title":"3, 117–132","cited_arxiv_id":null,"evidence_quote":"establishes feedback as a critical component of effective e-assessment."}],"review_version":1}