Pith. sign in

REVIEW 4 major objections 5 minor 38 references

Evaluating the use of e-assessment in a first-year pure mathematics module

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper reports evidence that periodic open-book online quizzes, counting for 10 percent of the grade, give first-year proof students an extrinsic motivation to engage early with lecture notes and make the small details of proofs hard…

desk verdict Honest, well-designed case study of proof-validation quizzes whose central claim outruns the self-report evidence; still worth a serious referee. read the letter →

arxiv 1908.01028 v5 pith:4ZS4W6OO submitted 2019-08-02 math.HO

classification math.HO
keywords e-assessmentmathematicalproofsundergraduatemathematicseducationonlinequizzesassessmentforlearningproofvalidationstudentfeedbacktransitiontouniversity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports a two-year case study in which bi-weekly, open-book online quizzes were added to a first-year pure mathematics module, Introduction to Proofs, alongside written homework and an end-of-term exam. The quizzes were deliberately built as assessment for learning rather than assessment of learning: they gave automatic, individually tailored feedback and tested the kind of proof details that written homework often misses, such as defining notation, covering all cases, and forming correct contrapositives. Based on end-of-module questionnaires, the author argues that the quizzes successfully gave students an extrinsic reason to engage with lecture notes early and sharpened their attention to small details within proofs, while noting that only a minority of students connected the quizzes to their own proof writing. If this is right, e-assessment is not a replacement for written homework but a workable complement to it for large first-year proof modules.

What carries the argument

The core mechanism is the proof-validation question: a multiple-choice, matching, or fill-in-the-blank item that presents a proof and asks students to judge its correctness or arrange its steps. Because the chosen answer is right or wrong as a whole, one missing detail—say, failing to define x ∈ Z before using divisibility—collapses an otherwise valid-looking proof. The automatic feedback then states exactly which detail broke the proof and how to repair it, sometimes changing N to Z or inserting one line. This couplet of all-or-nothing proof validation plus tailored repair feedback is what carries the paper's argument.

What would settle it

In a matched comparison of two first-year proof modules, one with bi-weekly open-book quizzes and one without, a blind-scored end-of-term proof question would show whether quiz students produce more rigorous proofs (all notation defined, all cases covered). If the quiz cohort's proofs are no more rigorous, the claim that the quizzes emphasized small proof details is not supported.

Watch

Extended reading notes

Core claim

The central discovery is that assessment for learning in pure mathematics does not have to wait for written proofs. In the author's words, the online quizzes successfully achieved their intended goals: they gave first-year students an extrinsic motivation to engage with lecture notes early, and they made the small details inside proofs—defining variables, covering all cases, forming contrapositives—difficult to ignore. Survey data from two cohorts showed 55 percent then 72 percent of students crediting the quizzes with early engagement with lecture notes, 69 percent then 59 percent crediting them with sharper attention to small details when writing proofs, and feedback-reading rising from 44 percent to 64 percent after the author repeatedly taught students where to find it. The author also reports a persistent gap: only about a fifth of students felt quiz feedback helped them write their own proofs later, which she explains as a transfer gap rather than a failure of the quizzes.

Load-bearing premise

The central claim rests on students' self-reported survey answers, from 59% and 37% response rates, accurately reflecting their engagement and attention to proof details.

Editorial extensions

If this is right

  • Students report engaging with lecture notes earlier in both years of quizzes than in the year before, with the second-year cohort reporting stronger engagement after repeated feedback instructions.
  • All-or-nothing proof-validation questions make a single omitted detail—such as an undefined variable or an unexamined case—break the whole proof, which is exactly the lesson written homework often leaves implicit.
  • Changing 10 percent of quiz questions from multiple-choice to fill-in-the-blank and matching forms eliminated student complaints about guessing and coincided with higher reported engagement with lecture notes.
  • Feedback reading rose from 44 percent to 64 percent after the instructor repeatedly showed students where feedback lived, indicating that engagement with feedback is teachable.
  • The persistent gap between validating proofs and writing proofs (only about 22–23 percent of students credited quizzes with helping their own written work) tells the author that transfer needs to be taught explicitly, not assumed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the author leaves implicit: the transfer gap could be attacked directly by adding a short free-response follow-up to each quiz in which students fix the flawed proof they just evaluated; the feedback examples already provide the repair language.
  • A plausible extension: systematically vary the proportion of multiple-choice versus fill-in-the-blank and matching items to find the mix that keeps guessing low without sacrificing the small-detail emphasis that all-or-nothing validation provides.
  • A natural corroborating study: use platform logs of quiz attempts, submission timing, and feedback clicks to measure early engagement directly instead of relying on end-of-module self-reports.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a two-year case study in which bi-weekly online quizzes (10% of the final grade, with automatic feedback) were introduced into a first-year pure mathematics module, 'Introduction to Proofs.' The quizzes were intended to motivate early engagement with learning material and to draw attention to small details in proofs. The evaluation draws on end-of-module student questionnaires (response rates 59% and 37%) and on the author's teaching observations. The paper reports that majorities of respondents perceived the quizzes as supporting engagement and proof-detail awareness, and it concludes that the quizzes 'successfully achieved their intended goals.'

Significance. If the central claim could be supported by objective evidence, the paper would provide a useful model for integrating e-assessment into proof-based instruction, an area where automated assessment is rare. The manuscript's concrete strengths are its detailed and replicable description of quiz design, worked examples of questions and tailored feedback, and a thoughtful set of practical recommendations in Section 6. Its principal weakness is that the evidence is entirely self-reported, with low response rates, no control group, and no objective measure of engagement or learning; the paper's own exam-mark data show no change. The contribution is therefore best read as a practitioner case description rather than a strong empirical demonstration, though the implementation details are valuable to the teaching community.

major comments (4)
  1. [Section 4, response rates] The central conclusion in Section 6 rests on student self-reports from end-of-module questionnaires with response rates of 59% (Year 1) and 37% (Year 2). Because non-respondents may differ systematically from respondents, and because self-reports of perceived benefit are not measures of actual engagement or proof-writing behaviour, the data as presented cannot support the unqualified claim that the quizzes 'successfully achieved their intended goals.' The manuscript should either present objective behavioural records (such as BlackBoard logs of quiz attempts and feedback views) or explicitly restrict the conclusion to perceived benefits.
  2. [Section 5, exam marks] The paper reports that average student marks were 'not affected' by the quizzes in either comparison. This objective outcome is acknowledged but not weighed in the Conclusion. For a claim that quizzes improved learning or proof-detail awareness, unchanged exam marks are a relevant contrary indicator. The Conclusion in Section 6 should either confront this evidence directly or temper the claim to state that the perceived benefits were achieved even though measurable performance did not change.
  3. [Section 4, Year 2 confound] In Year 2, online quizzes were introduced in all but one other first-year mathematics module at the institution. The attribution of the Year 2 increases in the engagement items (72% versus 55% and 75% versus 61%) to the author's 10% change in question type is therefore confounded by a system-wide change in assessment practice. The simultaneous 10-percentage-point drop in the 'small details' item (69% to 59%) is also not explained; this pattern is difficult to reconcile with the conclusion that both goals were achieved in both years.
  4. [Section 5, 'confirms' language] The manuscript states that 'The data presented in the previous chapter confirms that both of these goals were successfully achieved.' This is an overstatement: the questionnaire items measure students' perceptions ('thought that the online quizzes helped them...'), not demonstrated engagement or proof-detail behaviour. The language in the Discussion and Conclusion should be revised throughout to distinguish perceived benefit from demonstrated effect.
minor comments (5)
  1. [Section 4, figure references] The sentence 'Figure 2 shows that 69% of students in Year 1 thought that the online quizzes helped them to pay attention to the small details when writing proofs' should refer to Figure 3, which is the figure captioned 'Emphasis of the small details within proofs.'
  2. [Section 5, numeric inconsistency] The text states '69% of students in Year 1 and 77% of students in Year 2 indicated that the feedback helped them to understand why their particular answer choice was incorrect,' but Section 4 reports 68% for Year 1, not 69%.
  3. [Section 5, further numeric inconsistency] The text states 'only 59% of students, who had read the available feedback, claimed that the quizzes helped them to understand the small details when writing proofs,' but Section 4 reports this statistic as 56%. The discrepancy should be resolved.
  4. [Section 2, typo] The phrase 'feedback seems to be the one of the main tools' should read 'one of the main tools.'
  5. [Section 5, inference from absence of comments] The claim that the effect of guessing was reduced because 'not a single student commented on being able to guess answers in Year 2' is an inference from absence of evidence; it should be presented as a weaker, anecdotal observation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's conclusions are empirical survey-based claims, not results that reduce to their inputs by construction.

full rationale

This is an education case study, not a mathematical derivation or a fitted model, so the usual circularity patterns (fitted parameters renamed as predictions, uniqueness theorems imported from authors, ansatz smuggled via citation) do not apply. The central claims that periodic online quizzes motivated engagement and emphasized proof details are supported by end-of-module questionnaire responses, which are independent empirical data, however limited in validity. The paper does not define the outcome in terms of the conclusion, nor does it fit a parameter to a subset and then predict that same subset. There are no load-bearing self-citations; the author cites external literature for background and does not invoke her own prior work to justify the intervention's effectiveness. The paper even reports countervailing evidence, including that average student marks were not affected and that only 22-23% of students said feedback helped with subsequent written assignments, which shows the conclusions are not forced by the design. The main weaknesses, low response rates and reliance on self-report, are methodological validity concerns rather than circularity, and the manuscript itself acknowledges the gap between perceived learning and proof-writing performance. Under the stated rules, absence of a derivation chain means the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on assumptions about the validity of self-report data, the comparability of two academic years, and the reliability of anecdotal baseline observations. No free parameters or invented entities appear because the paper is an educational case study, not a quantitative model.

assumptions (3)
  • domain assumption Student self-reported questionnaire responses are valid indicators of engagement with learning material and attention to proof details.
    All key results in Section 4 come from end-of-module surveys; no objective measure of engagement or proof-writing skill is used.
  • domain assumption Differences in self-reports between Year 1 and Year 2 are caused by changes in the quiz design and feedback promotion rather than by other changes in the students' environment.
    The author attributes increases in engagement and feedback reading to her own modifications, but she also notes that in Year 2 all but one other first-year module introduced quizzes, a major contextual change not controlled for.
  • domain assumption The recollection of the year before the quizzes, where students 'rarely approached' with worries, is a reliable baseline.
    Section 5 uses this anecdotal memory to support the conclusion about reduced stress; no data from that prior year are presented.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating the use of e-assessment in a first-year pure mathematics module." pith.science (2026). https://pith.science/paper/4ZS4W6OO

@misc{pith2026190801028,
  author       = {Pith},
  title        = {Pith review of: Evaluating the use of e-assessment in a first-year pure mathematics module},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4ZS4W6OO}},
  note         = {Machine review of arXiv:1908.01028}
}
read the original abstract

This article presents the findings of a case study which introduced online quizzes as a form of assessment in pure mathematics. Rather than being designed as an assessment of learning, these quizzes were designed to be an assessment for learning; they aimed to academically support students in their transition from A-Level mathematics to university-level pure mathematics by providing an extrinsic motivation to engage them with their learning material early on and to emphasise the small details within proofs, such as defining notation, which are not necessarily emphasised by written homework assignments. The results obtained during the two-year study using online quizzes show e-assessment to be a powerful complementary tool to traditional written homework assignments.

Figures

Figures reproduced from arXiv: 1908.01028 by the authors.

Figure 1
Figure 1. Student engagement with lecture notes [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 3
Figure 3. Emphasis of the small details within proofs [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 38 canonical work pages

  1. [1]

    2, 255–272

    Simon D Angus and Judith Watson, Does regular online testing enhance student learning in the numerical sciences? robust evidence from a large data se t, British Journal of Educational Technology 40 (2009), no. 2, 255–272

  2. [2]

    Glenda Anthony, Factors influencing first-year students’ success in mathema tics, International Journal of Mathematical Education in Science and Technology 31 (2000), no. 1, 3–14

  3. [3]

    8, 916–931

    Leopold Bayerlein, Students’ feedback preferences: how do students react to ti mely and automat- ically generated assessment feedback? , Assessment & Evaluation in Higher Education 39 (2014), no. 8, 916–931

  4. [4]

    1, 61–67

    Menucha Birenbaum, Klaus Breuer, Eduardo Cascallar, Filip Dochy , Yehudit Dori, Jim Ridg- way, Rolf Wiesemes, and Goele Nickmans, A learning integrated assessment system , Educational Research Review 1 (2006), no. 1, 61–67

  5. [5]

    4, 470– 504

    Bruno Buchberger, Adrian Crˇ aciun, Tudor Jebelean, Laura Ko v´ acs, Temur Kutsia, Koji Naka- gawa, Florina Piroi, Nikolaj Popov, Judit Robu, Markus Rosenkranz, et al., Theorema: Towards computer-aided mathematical theory exploration , Journal of Applied Logic 4 (2006), no. 4, 470– 504

  6. [6]

    2, 121–134

    Eros DeSouza and Matthew Fleming, A comparison of in-class and online quizzes on student exam performance, Journal of Computing in Higher Education 14 (2003), no. 2, 121–134

  7. [7]

    4, 297–302

    John L Dobson, The use of formative online quizzes to enhance class prepara tion and scores on summative exams , Advances in Physiology Education 32 (2008), no. 4, 297–302

  8. [8]

    3, 637–646

    John F Feldhusen, An evaluation of college students’ reactions to open book ex aminations, Edu- cational and Psychological Measurement 21 (1961), no. 3, 637–646

Show all 38 references
  1. [9]

    3, 117–132

    Jorge Gaytan and Beryl C McEwen, Effective online instructional and assessment strategies , The American Journal of Distance Education 21 (2007), no. 3, 117–132

  2. [10]

    1, 40–45

    Sharon Gedye, Formative assessment and feedback: A review , Planet 23 (2010), no. 1, 40–45

  3. [11]

    4, 2333–2351

    Joyce Wangui Gikandi, Donna Morrow, and Niki E Davis, Online formative assessment in higher education: A review of the literature , Computers & education 57 (2011), no. 4, 2333–2351

  4. [12]

    3, 117–137

    Martin Greenhow, Effective computer-aided assessment of mathematics; princ iples, practice and results, Teaching Mathematics and its Applications: An International Jour nal of the IMA 34 (2015), no. 3, 117–137

  5. [13]

    5, IEEE, 2008, pp

    Susanne Gruttmann, Dominik B¨ ohm, and Herbert Kuchen, E-assessment of mathematical proofs: chances and challenges for students and tutors , 2008 International Conference on Computer Sci- ence and Software Engineering, vol. 5, IEEE, 2008, pp. 612–615

  6. [14]

    Lee Harvey, Student feedback [1] , Quality in higher education 9 (2003), no. 1, 3–20

  7. [15]

    Marjolein Heijne-Penninga, Open-book tests assessed: quality, learning behaviour tes t time and performance, Tijdschrift voor Medisch Onderwijs 30 (2011), no. 3, 109

  8. [16]

    Timothy Hetherington, Assessing proofs in pure mathematics , Mapping university mathemat- ics assessment practices (Paola Iannone and Adrian Simpson, eds.) , University of East Anglia Norwich, 2012, pp. 83–93. 12

  9. [17]

    4, 186–196

    Paola Iannone and Adrian Simpson, The summative assessment diet: how we assess in mathe- matics degrees, Teaching Mathematics and its Applications: An International Jour nal of the IMA 30 (2011), no. 4, 186–196

  10. [18]

    Alastair Irons, Enhancing learning through formative assessment and feedb ack, Routledge, 2007

  11. [19]

    2, 125–129

    Jonathan D Kibble, Teresa R Johnson, Mohammed K Khalil, Loren D Nelson, Garrett H Riggs, Jose L Borrero, and Andrew F Payer, Insights gained from the analysis of performance and participation in online formative assessment , Teaching and learning in medicine 23 (2011), no. 2, 125–129

  12. [20]

    London Mathematical Society, Mathematics degrees, their teaching and assessment , Press Release, 2010, LMS Teaching Position Statement

  13. [21]

    Mayrhofer, S

    G. Mayrhofer, S. Saminger, and W. Windsteiger, CreaComp: Computer-Supported Experiments and Automated Proving in Learning and Teaching Mathematics , Proceedings of ICTMT8, 2007

  14. [22]

    Juan Pablo Mejia-Ramos, Evan Fuller, Keith Weber, Kathryn Rho ads, and Aron Samkoff, An assessment model for proof comprehension in undergraduate mathematics, Educational Studies in Mathematics 79 (2012), no. 1, 3–18

  15. [23]

    , Alberta journal of educational research 19 (1973), no

    Shirleyanne Michaels and TR Kieren, An investigation of open-book and closed-book examination s in mathematics. , Alberta journal of educational research 19 (1973), no. 3, 202–07

  16. [24]

    David Miller, Nicole Infante, and Keith Weber, How mathematicians assign points to student proofs, The Journal of Mathematical Behavior 49 (2018), 24–34

  17. [25]

    2, 246–278

    Robert C Moore, Mathematics professors’ evaluation of students’ proofs: A complex teaching practice, International Journal of Research in Undergraduate Mathema tics Education 2 (2016), no. 2, 246–278

  18. [26]

    4, 501–514

    Robert A Powers, Cathleen Craviotto, and Richard M Grassl, Impact of proof validation on proof writing in abstract algebra , International Journal of Mathematical Education in Science and Technology 41 (2010), no. 4, 501–514

  19. [27]

    , Journal of Technology and Science Education 2 (2012), no

    Lorenzo Salas-Morera, Antonio Arauzo-Azofra, and Laura G arc ´ ıa-Hern´ andez,Analysis of online quizzes as a teaching and assessment tool. , Journal of Technology and Science Education 2 (2012), no. 1, 39–45

  20. [28]

    Chris Sangwin, Computer aided assessment of mathematics , OUP Oxford, 2013

  21. [29]

    4, 423–431

    Carina Savander-Ranne, Olli-Pekka Lund´ en, and Samuli Kolari, An alternative teaching method for electrical engineering courses , IEEE Transactions on Education 51 (2008), no. 4, 423–431

  22. [30]

    Annie Selden and John Selden, Validations of proofs considered as texts: Can undergradua tes tell whether an argument proves a theorem? , Journal for research in mathematics education 34 (2003), no. 1, 4–36

  23. [31]

    Glenn Gordon Smith and David Ferguson, Student attrition in mathematics e-learning , Aus- tralasian Journal of Educational Technology 21 (2005), no. 3

  24. [32]

    , ERIC, 2006

    Lynn Arthur Steen, Supporting assessment in undergraduate mathematics. , ERIC, 2006

  25. [33]

    4, 379– 393

    Christos Theophilides and Mary Koutselini, Study behavior in the closed-book and the open-book examination: A comparative analysis , Educational Research and Evaluation 6 (2000), no. 4, 379– 393. 13

  26. [34]

    1, 101–119

    Keith Weber, Student difficulty in constructing proofs: The need for strat egic knowledge, Educa- tional studies in mathematics 48 (2001), no. 1, 101–119

  27. [35]

    2, 311–329

    Thomas Wolsey, Efficacy of instructor feedback on written work in an online pr ogram, Interna- tional Journal on E-learning 7 (2008), no. 2, 311–329

  28. [36]

    4, 477–501

    Mantz Yorke, Formative assessment in higher education: Moves towards th eory and the enhance- ment of pedagogic practice , Higher education 45 (2003), no. 4, 477–501

  29. [37]

    3, 137–166

    Xinming Zhu and Herbert A Simon, Learning mathematics from examples and by doing , Cognition and instruction 4 (1987), no. 3, 137–166. APPENDIX Below are some examples of the different question types used for th e online quizzes. Example 3 is an example of a multiple-choice que...

  30. [38]

    Define f : N → Z by f (x) = x2 − 4.Please arrange the following proof by contradiction in the correct order

    ⋃ i∈ N (Ai ∪ Bi) = c) 2 N − 1 d) ∅ Example 6. Define f : N → Z by f (x) = x2 − 4.Please arrange the following proof by contradiction in the correct order. 14 Step 1 a) Therefore, f (x) ⁄= f (y). Step 2 b) Take x, y∈ N such that x = y. Step 3 c) We will show that f is injective....

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.