REVIEW 4 major objections 5 minor 38 references
Evaluating the use of e-assessment in a first-year pure mathematics module
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper reports evidence that periodic open-book online quizzes, counting for 10 percent of the grade, give first-year proof students an extrinsic motivation to engage early with lecture notes and make the small details of proofs hard…
desk verdict Honest, well-designed case study of proof-validation quizzes whose central claim outruns the self-report evidence; still worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core mechanism is the proof-validation question: a multiple-choice, matching, or fill-in-the-blank item that presents a proof and asks students to judge its correctness or arrange its steps. Because the chosen answer is right or wrong as a whole, one missing detail—say, failing to define x ∈ Z before using divisibility—collapses an otherwise valid-looking proof. The automatic feedback then states exactly which detail broke the proof and how to repair it, sometimes changing N to Z or inserting one line. This couplet of all-or-nothing proof validation plus tailored repair feedback is what carries the paper's argument.
What would settle it
In a matched comparison of two first-year proof modules, one with bi-weekly open-book quizzes and one without, a blind-scored end-of-term proof question would show whether quiz students produce more rigorous proofs (all notation defined, all cases covered). If the quiz cohort's proofs are no more rigorous, the claim that the quizzes emphasized small proof details is not supported.
Extended reading notes
Core claim
The central discovery is that assessment for learning in pure mathematics does not have to wait for written proofs. In the author's words, the online quizzes successfully achieved their intended goals: they gave first-year students an extrinsic motivation to engage with lecture notes early, and they made the small details inside proofs—defining variables, covering all cases, forming contrapositives—difficult to ignore. Survey data from two cohorts showed 55 percent then 72 percent of students crediting the quizzes with early engagement with lecture notes, 69 percent then 59 percent crediting them with sharper attention to small details when writing proofs, and feedback-reading rising from 44 percent to 64 percent after the author repeatedly taught students where to find it. The author also reports a persistent gap: only about a fifth of students felt quiz feedback helped them write their own proofs later, which she explains as a transfer gap rather than a failure of the quizzes.
Load-bearing premise
The central claim rests on students' self-reported survey answers, from 59% and 37% response rates, accurately reflecting their engagement and attention to proof details.
Editorial extensions
If this is right
- Students report engaging with lecture notes earlier in both years of quizzes than in the year before, with the second-year cohort reporting stronger engagement after repeated feedback instructions.
- All-or-nothing proof-validation questions make a single omitted detail—such as an undefined variable or an unexamined case—break the whole proof, which is exactly the lesson written homework often leaves implicit.
- Changing 10 percent of quiz questions from multiple-choice to fill-in-the-blank and matching forms eliminated student complaints about guessing and coincided with higher reported engagement with lecture notes.
- Feedback reading rose from 44 percent to 64 percent after the instructor repeatedly showed students where feedback lived, indicating that engagement with feedback is teachable.
- The persistent gap between validating proofs and writing proofs (only about 22–23 percent of students credited quizzes with helping their own written work) tells the author that transfer needs to be taught explicitly, not assumed.
Reading between the lines
- An implication the author leaves implicit: the transfer gap could be attacked directly by adding a short free-response follow-up to each quiz in which students fix the flawed proof they just evaluated; the feedback examples already provide the repair language.
- A plausible extension: systematically vary the proportion of multiple-choice versus fill-in-the-blank and matching items to find the mix that keeps guessing low without sacrificing the small-detail emphasis that all-or-nothing validation provides.
- A natural corroborating study: use platform logs of quiz attempts, submission timing, and feedback clicks to measure early engagement directly instead of relying on end-of-module self-reports.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a two-year case study in which bi-weekly online quizzes (10% of the final grade, with automatic feedback) were introduced into a first-year pure mathematics module, 'Introduction to Proofs.' The quizzes were intended to motivate early engagement with learning material and to draw attention to small details in proofs. The evaluation draws on end-of-module student questionnaires (response rates 59% and 37%) and on the author's teaching observations. The paper reports that majorities of respondents perceived the quizzes as supporting engagement and proof-detail awareness, and it concludes that the quizzes 'successfully achieved their intended goals.'
Significance. If the central claim could be supported by objective evidence, the paper would provide a useful model for integrating e-assessment into proof-based instruction, an area where automated assessment is rare. The manuscript's concrete strengths are its detailed and replicable description of quiz design, worked examples of questions and tailored feedback, and a thoughtful set of practical recommendations in Section 6. Its principal weakness is that the evidence is entirely self-reported, with low response rates, no control group, and no objective measure of engagement or learning; the paper's own exam-mark data show no change. The contribution is therefore best read as a practitioner case description rather than a strong empirical demonstration, though the implementation details are valuable to the teaching community.
major comments (4)
- [Section 4, response rates] The central conclusion in Section 6 rests on student self-reports from end-of-module questionnaires with response rates of 59% (Year 1) and 37% (Year 2). Because non-respondents may differ systematically from respondents, and because self-reports of perceived benefit are not measures of actual engagement or proof-writing behaviour, the data as presented cannot support the unqualified claim that the quizzes 'successfully achieved their intended goals.' The manuscript should either present objective behavioural records (such as BlackBoard logs of quiz attempts and feedback views) or explicitly restrict the conclusion to perceived benefits.
- [Section 5, exam marks] The paper reports that average student marks were 'not affected' by the quizzes in either comparison. This objective outcome is acknowledged but not weighed in the Conclusion. For a claim that quizzes improved learning or proof-detail awareness, unchanged exam marks are a relevant contrary indicator. The Conclusion in Section 6 should either confront this evidence directly or temper the claim to state that the perceived benefits were achieved even though measurable performance did not change.
- [Section 4, Year 2 confound] In Year 2, online quizzes were introduced in all but one other first-year mathematics module at the institution. The attribution of the Year 2 increases in the engagement items (72% versus 55% and 75% versus 61%) to the author's 10% change in question type is therefore confounded by a system-wide change in assessment practice. The simultaneous 10-percentage-point drop in the 'small details' item (69% to 59%) is also not explained; this pattern is difficult to reconcile with the conclusion that both goals were achieved in both years.
- [Section 5, 'confirms' language] The manuscript states that 'The data presented in the previous chapter confirms that both of these goals were successfully achieved.' This is an overstatement: the questionnaire items measure students' perceptions ('thought that the online quizzes helped them...'), not demonstrated engagement or proof-detail behaviour. The language in the Discussion and Conclusion should be revised throughout to distinguish perceived benefit from demonstrated effect.
minor comments (5)
- [Section 4, figure references] The sentence 'Figure 2 shows that 69% of students in Year 1 thought that the online quizzes helped them to pay attention to the small details when writing proofs' should refer to Figure 3, which is the figure captioned 'Emphasis of the small details within proofs.'
- [Section 5, numeric inconsistency] The text states '69% of students in Year 1 and 77% of students in Year 2 indicated that the feedback helped them to understand why their particular answer choice was incorrect,' but Section 4 reports 68% for Year 1, not 69%.
- [Section 5, further numeric inconsistency] The text states 'only 59% of students, who had read the available feedback, claimed that the quizzes helped them to understand the small details when writing proofs,' but Section 4 reports this statistic as 56%. The discrepancy should be resolved.
- [Section 2, typo] The phrase 'feedback seems to be the one of the main tools' should read 'one of the main tools.'
- [Section 5, inference from absence of comments] The claim that the effect of guessing was reduced because 'not a single student commented on being able to guess answers in Year 2' is an inference from absence of evidence; it should be presented as a weaker, anecdotal observation.
Circularity Check
No circularity: the paper's conclusions are empirical survey-based claims, not results that reduce to their inputs by construction.
full rationale
This is an education case study, not a mathematical derivation or a fitted model, so the usual circularity patterns (fitted parameters renamed as predictions, uniqueness theorems imported from authors, ansatz smuggled via citation) do not apply. The central claims that periodic online quizzes motivated engagement and emphasized proof details are supported by end-of-module questionnaire responses, which are independent empirical data, however limited in validity. The paper does not define the outcome in terms of the conclusion, nor does it fit a parameter to a subset and then predict that same subset. There are no load-bearing self-citations; the author cites external literature for background and does not invoke her own prior work to justify the intervention's effectiveness. The paper even reports countervailing evidence, including that average student marks were not affected and that only 22-23% of students said feedback helped with subsequent written assignments, which shows the conclusions are not forced by the design. The main weaknesses, low response rates and reliance on self-report, are methodological validity concerns rather than circularity, and the manuscript itself acknowledges the gap between perceived learning and proof-writing performance. Under the stated rules, absence of a derivation chain means the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Student self-reported questionnaire responses are valid indicators of engagement with learning material and attention to proof details.
- domain assumption Differences in self-reports between Year 1 and Year 2 are caused by changes in the quiz design and feedback promotion rather than by other changes in the students' environment.
- domain assumption The recollection of the year before the quizzes, where students 'rarely approached' with worries, is a reliable baseline.
Cite this review
Pith. "Pith review of Evaluating the use of e-assessment in a first-year pure mathematics module." pith.science (2026). https://pith.science/paper/4ZS4W6OO
@misc{pith2026190801028,
author = {Pith},
title = {Pith review of: Evaluating the use of e-assessment in a first-year pure mathematics module},
year = {2026},
howpublished = {\url{https://pith.science/paper/4ZS4W6OO}},
note = {Machine review of arXiv:1908.01028}
}
read the original abstract
This article presents the findings of a case study which introduced online quizzes as a form of assessment in pure mathematics. Rather than being designed as an assessment of learning, these quizzes were designed to be an assessment for learning; they aimed to academically support students in their transition from A-Level mathematics to university-level pure mathematics by providing an extrinsic motivation to engage them with their learning material early on and to emphasise the small details within proofs, such as defining notation, which are not necessarily emphasised by written homework assignments. The results obtained during the two-year study using online quizzes show e-assessment to be a powerful complementary tool to traditional written homework assignments.
Figures
Reference graph
Works this paper leans on
-
[1]
Simon D Angus and Judith Watson, Does regular online testing enhance student learning in the numerical sciences? robust evidence from a large data se t, British Journal of Educational Technology 40 (2009), no. 2, 255–272
work page 2009
-
[2]
Glenda Anthony, Factors influencing first-year students’ success in mathema tics, International Journal of Mathematical Education in Science and Technology 31 (2000), no. 1, 3–14
work page 2000
-
[3]
Leopold Bayerlein, Students’ feedback preferences: how do students react to ti mely and automat- ically generated assessment feedback? , Assessment & Evaluation in Higher Education 39 (2014), no. 8, 916–931
work page 2014
- [4]
-
[5]
Bruno Buchberger, Adrian Crˇ aciun, Tudor Jebelean, Laura Ko v´ acs, Temur Kutsia, Koji Naka- gawa, Florina Piroi, Nikolaj Popov, Judit Robu, Markus Rosenkranz, et al., Theorema: Towards computer-aided mathematical theory exploration , Journal of Applied Logic 4 (2006), no. 4, 470– 504
work page 2006
-
[6]
Eros DeSouza and Matthew Fleming, A comparison of in-class and online quizzes on student exam performance, Journal of Computing in Higher Education 14 (2003), no. 2, 121–134
work page 2003
-
[7]
John L Dobson, The use of formative online quizzes to enhance class prepara tion and scores on summative exams , Advances in Physiology Education 32 (2008), no. 4, 297–302
work page 2008
-
[8]
John F Feldhusen, An evaluation of college students’ reactions to open book ex aminations, Edu- cational and Psychological Measurement 21 (1961), no. 3, 637–646
work page 1961
Show all 38 references
-
[9]
3, 117–132
Jorge Gaytan and Beryl C McEwen, Effective online instructional and assessment strategies , The American Journal of Distance Education 21 (2007), no. 3, 117–132
2007
-
[10]
1, 40–45
Sharon Gedye, Formative assessment and feedback: A review , Planet 23 (2010), no. 1, 40–45
2010
-
[11]
4, 2333–2351
Joyce Wangui Gikandi, Donna Morrow, and Niki E Davis, Online formative assessment in higher education: A review of the literature , Computers & education 57 (2011), no. 4, 2333–2351
2011
-
[12]
3, 117–137
Martin Greenhow, Effective computer-aided assessment of mathematics; princ iples, practice and results, Teaching Mathematics and its Applications: An International Jour nal of the IMA 34 (2015), no. 3, 117–137
2015
-
[13]
5, IEEE, 2008, pp
Susanne Gruttmann, Dominik B¨ ohm, and Herbert Kuchen, E-assessment of mathematical proofs: chances and challenges for students and tutors , 2008 International Conference on Computer Sci- ence and Software Engineering, vol. 5, IEEE, 2008, pp. 612–615
2008
-
[14]
Lee Harvey, Student feedback [1] , Quality in higher education 9 (2003), no. 1, 3–20
2003
-
[15]
Marjolein Heijne-Penninga, Open-book tests assessed: quality, learning behaviour tes t time and performance, Tijdschrift voor Medisch Onderwijs 30 (2011), no. 3, 109
2011
-
[16]
Timothy Hetherington, Assessing proofs in pure mathematics , Mapping university mathemat- ics assessment practices (Paola Iannone and Adrian Simpson, eds.) , University of East Anglia Norwich, 2012, pp. 83–93. 12
2012
-
[17]
4, 186–196
Paola Iannone and Adrian Simpson, The summative assessment diet: how we assess in mathe- matics degrees, Teaching Mathematics and its Applications: An International Jour nal of the IMA 30 (2011), no. 4, 186–196
2011
-
[18]
Alastair Irons, Enhancing learning through formative assessment and feedb ack, Routledge, 2007
2007
-
[19]
2, 125–129
Jonathan D Kibble, Teresa R Johnson, Mohammed K Khalil, Loren D Nelson, Garrett H Riggs, Jose L Borrero, and Andrew F Payer, Insights gained from the analysis of performance and participation in online formative assessment , Teaching and learning in medicine 23 (2011), no. 2, 125–129
2011
-
[20]
London Mathematical Society, Mathematics degrees, their teaching and assessment , Press Release, 2010, LMS Teaching Position Statement
2010
-
[21]
Mayrhofer, S
G. Mayrhofer, S. Saminger, and W. Windsteiger, CreaComp: Computer-Supported Experiments and Automated Proving in Learning and Teaching Mathematics , Proceedings of ICTMT8, 2007
2007
-
[22]
Juan Pablo Mejia-Ramos, Evan Fuller, Keith Weber, Kathryn Rho ads, and Aron Samkoff, An assessment model for proof comprehension in undergraduate mathematics, Educational Studies in Mathematics 79 (2012), no. 1, 3–18
2012
-
[23]
, Alberta journal of educational research 19 (1973), no
Shirleyanne Michaels and TR Kieren, An investigation of open-book and closed-book examination s in mathematics. , Alberta journal of educational research 19 (1973), no. 3, 202–07
1973
-
[24]
David Miller, Nicole Infante, and Keith Weber, How mathematicians assign points to student proofs, The Journal of Mathematical Behavior 49 (2018), 24–34
2018
-
[25]
2, 246–278
Robert C Moore, Mathematics professors’ evaluation of students’ proofs: A complex teaching practice, International Journal of Research in Undergraduate Mathema tics Education 2 (2016), no. 2, 246–278
2016
-
[26]
4, 501–514
Robert A Powers, Cathleen Craviotto, and Richard M Grassl, Impact of proof validation on proof writing in abstract algebra , International Journal of Mathematical Education in Science and Technology 41 (2010), no. 4, 501–514
2010
-
[27]
, Journal of Technology and Science Education 2 (2012), no
Lorenzo Salas-Morera, Antonio Arauzo-Azofra, and Laura G arc ´ ıa-Hern´ andez,Analysis of online quizzes as a teaching and assessment tool. , Journal of Technology and Science Education 2 (2012), no. 1, 39–45
2012
-
[28]
Chris Sangwin, Computer aided assessment of mathematics , OUP Oxford, 2013
2013
-
[29]
4, 423–431
Carina Savander-Ranne, Olli-Pekka Lund´ en, and Samuli Kolari, An alternative teaching method for electrical engineering courses , IEEE Transactions on Education 51 (2008), no. 4, 423–431
2008
-
[30]
Annie Selden and John Selden, Validations of proofs considered as texts: Can undergradua tes tell whether an argument proves a theorem? , Journal for research in mathematics education 34 (2003), no. 1, 4–36
2003
-
[31]
Glenn Gordon Smith and David Ferguson, Student attrition in mathematics e-learning , Aus- tralasian Journal of Educational Technology 21 (2005), no. 3
2005
-
[32]
, ERIC, 2006
Lynn Arthur Steen, Supporting assessment in undergraduate mathematics. , ERIC, 2006
2006
-
[33]
4, 379– 393
Christos Theophilides and Mary Koutselini, Study behavior in the closed-book and the open-book examination: A comparative analysis , Educational Research and Evaluation 6 (2000), no. 4, 379– 393. 13
2000
-
[34]
1, 101–119
Keith Weber, Student difficulty in constructing proofs: The need for strat egic knowledge, Educa- tional studies in mathematics 48 (2001), no. 1, 101–119
2001
-
[35]
2, 311–329
Thomas Wolsey, Efficacy of instructor feedback on written work in an online pr ogram, Interna- tional Journal on E-learning 7 (2008), no. 2, 311–329
2008
-
[36]
4, 477–501
Mantz Yorke, Formative assessment in higher education: Moves towards th eory and the enhance- ment of pedagogic practice , Higher education 45 (2003), no. 4, 477–501
2003
-
[37]
3, 137–166
Xinming Zhu and Herbert A Simon, Learning mathematics from examples and by doing , Cognition and instruction 4 (1987), no. 3, 137–166. APPENDIX Below are some examples of the different question types used for th e online quizzes. Example 3 is an example of a multiple-choice que...
1987
-
[38]
Define f : N → Z by f (x) = x2 − 4.Please arrange the following proof by contradiction in the correct order
⋃ i∈ N (Ai ∪ Bi) = c) 2 N − 1 d) ∅ Example 6. Define f : N → Z by f (x) = x2 − 4.Please arrange the following proof by contradiction in the correct order. 14 Step 1 a) Therefore, f (x) ⁄= f (y). Step 2 b) Take x, y∈ N such that x = y. Step 3 c) We will show that f is injective....
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.