REVIEW 3 major objections 4 minor 1 cited by
Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read In an introductory real analysis course, a pre-/post-assessment cycle raised mean quiz scores from 6.39 to 11.11 and produced positive gains in engagement and self-efficacy.
desk verdict A useful teaching narrative with an unsupported causal headline; worth a review if reframed as a case study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the paired pre-assessment and post-assessment cycle, with immediate self-grading, a written reflection prompt, and instructor-led solution discussion. The pre-assessment quiz is a short diagnostic given at the start of the limit unit; students grade their own responses, articulate a learning goal, and then receive five class sessions of proof-based instruction. The post-assessment mirrors the pre-assessment in structure and content, so the pre-to-post score difference is presented as a direct measure of learning gain. The reflection questions and self-grading are the active components that the paper credits with fostering metacognition and engagement.
What would settle it
Run the same unit with a control section that receives identical instruction but takes only the post-assessment; if the control post-assessment mean is statistically indistinguishable from the pre/post group's post-assessment mean, the claimed causal effect fails. Alternatively, give a delayed post-assessment several weeks after the unit: if the gain vanishes, the result reflects short-term familiarity rather than durable learning.
Extended reading notes
Core claim
The paper's central claim is that a structured, self-graded pre-assessment followed by instruction and a parallel post-assessment improves student learning, engagement, self-efficacy, and conceptual mastery in a real analysis unit on limits. It reports that among the 18 to 20 participating students, the mean score rose from 6.39 to 11.11, the median rose from 6.5 to 11.0, and the paired t-test gave $p = 0.0013$ with Cohen's $d = 2.36$. Students who initially struggled with concepts such as the uniqueness of a limit answered those questions correctly on the post-assessment, and feedback highlighted the value of reviewing definitions, self-grading, and instructor explanation. The author interprets these results as evidence that pre-assessments act as learning roadmaps and post-assessments as measures of growth, helping students transition from procedural calculus thinking to proof-based understanding.
Load-bearing premise
The paper assumes that the measured pre-to-post score gain is caused by the pre/post assessment activity itself, and not by the five class sessions of instruction, by practice effects from seeing similar questions twice, or by regression toward the mean.
Editorial extensions
If this is right
- Instructors of proof-based courses can reuse the paired-quiz design as a no-stakes diagnostic that gives students a preview of unit content and a concrete target for study.
- Because the quizzes are self-graded, the activity adds minimal grading burden and turns the assessment moment itself into a teaching opportunity.
- If the claimed gain is real, the practice should generalize to other abstract units where definitions and proofs are the bottleneck, such as continuity or differentiability.
- The positive feedback about confidence and exam preparation suggests the activity may carry motivational benefits separate from raw score gains.
Reading between the lines
- Because the study lacks a control group, a fair causal test would compare a section doing the pre/post cycle with a section receiving identical instruction and only the post-assessment; the paper's design cannot separate the assessment activity from the five lectures.
- The self-grading and reflection prompt are plausible active ingredients; a follow-up that varies or removes those components would show whether the gain comes from the quizzes themselves or from the surrounding discussion.
- A delayed post-assessment, given weeks after the unit, would test whether the gains reflect durable understanding or short-term familiarity with the mirrored questions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a classroom implementation of a structured pre-assessment, immediate self-grading and reflection, five lectures on limits, and a post-assessment in an introductory real analysis course (N=28). It presents descriptive statistics (pre mean 6.39, post mean 11.11), a paired t-test (p=0.0013), Cohen's d=2.36, and qualitative student feedback, and interprets these as evidence that the pre/post assessment activity improved learning outcomes, engagement, self-efficacy, and conceptual understanding. The appendices include the actual assessment instruments.
Significance. If the causal claims were supported, the study would provide a useful contribution to formative assessment in proof-based mathematics. The paper's strengths are the detailed description of the intervention, the inclusion of the actual quizzes in the appendices, and the honest reporting of some practical caveats, such as the difficulty of scheduling and participation (Section 4). However, the evidence is descriptive, and the inferential statistics do not identify the effect of the assessment activity because the five instructional sessions between the two measurements are a complete confound. The paper can be a valuable case study of how students respond to such an activity, but it cannot support the causal language used in the title, abstract, and discussion.
major comments (3)
- [Abstract; §2.1; §3 Results] The central claim that pre- and post-assessments "shape learning outcomes" and "have an impact" is not supported by the data. The pre-assessment was followed immediately by instructor-led solution discussion and reflection, then by five 50-minute lectures on exactly the tested material (§2.1). The post-assessment scores therefore reflect the combined effect of the assessment activity, the instruction, practice from seeing similar questions twice, and regression to the mean. A paired t-test (p=0.0013) and Cohen's d=2.36 only test whether the mean change differs from zero; they do not identify the assessment activity as the cause. The manuscript should remove causal language from the title, abstract, and Section 4.2, or provide a comparison condition; with the current design, the descriptive case-study framing is the only defensible one.
- [§3, Table 1] The statistical reporting is incomplete. The paper states that 18 of 28 students took the pre-assessment and 20 of 28 took the post-assessment, but it does not report how many students took both, the number of matched pairs, or the degrees of freedom for the paired t-test. If the means in Table 1 are computed from different student groups, a paired test is invalid; if only the matched subset was used, the missing-data mechanism needs discussion. The 73.91% improvement is computed from group means and may not represent a per-student gain. The paper should report the paired differences (mean, SD, 95% CI), the number of pairs, and a comparison of participants and nonparticipants. The effect size also needs a confidence interval.
- [§3.1.2, §3.1.3] Claims about engagement, self-efficacy, and conceptual understanding rest entirely on self-reported student comments and on pie-chart percentages whose denominators are not stated (e.g., "56.3% Much better" in Fig. 2). Self-reports collected by the instructor are susceptible to demand characteristics and do not constitute validated measures of the constructs named in the abstract. The number of respondents for each feedback question and the response rate should be reported, and the claims should be limited to "students perceived" rather than "the activity enhanced."
minor comments (4)
- [§1.1] There is a typo in the text: "assessment plays" is split as "play s"; please proofread the manuscript carefully.
- [Figure 1] Figure 1 lacks a descriptive caption and the x-axis categories are not clearly labeled; readers cannot tell which bars correspond to pre- and post-assessment without referring to the table.
- [References] Some references are incomplete, e.g., Berry (2008) lacks page numbers or a DOI, and the Balan (2012) entry would benefit from fuller bibliographic details consistent with the other references.
- [§4] The caveats section mentions absent students and timing issues but does not acknowledge the fundamental absence of a control condition; this should be stated explicitly as a limitation of the study.
Circularity Check
No circularity found: the pre/post gain is an empirical observation that could have differed, and no claimed result is defined into existence or forced by a fit or self-citation.
full rationale
This paper makes no mathematical derivation and contains no fitted parameters, normalized quantities, or self-citations that could force its conclusions. The central quantitative result, 'The mean pre-assessment score was 6.39, while the mean post-assessment score increased to 11.11, indicating a 73.91% improvement' (Section 3), is an empirical before/after comparison: the post-assessment scores are not defined in terms of the pre-assessment scores, and the reported paired t-test (p=0.0013) and Cohen's d (2.36) are statistics computed from observations that could have come out differently. Nothing in the paper defines 'learning outcome' as 'post-assessment score' and then re-presents that score as evidence for itself; the claim is that the activity contributed to the gain, which is a causal inference, not a definitional identity. The main threats to that inference, namely the five intervening instruction sessions, practice effects from closely mirrored quizzes, and lack of a control group, are internal-validity limitations that are acknowledged only partially in Section 4's caveats about timing and absent students, but they are not circularity: they do not make the observed score change equal to the paper's input by construction. There are no self-citations, no imported uniqueness theorems, no ansatz smuggled in by citation, and no renaming of a known empirical pattern as an organizing result. The study is self-contained as an observational case study; its weakness is causal identification, not circular derivation. Therefore no circular step is identified and the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Scores on the instructor-designed pre/post quizzes validly measure students' conceptual understanding of limits.
- domain assumption The students who completed the assessments are representative of the class and the paired t-test is based on actual pre/post pairs.
- domain assumption Any improvement between pre- and post-assessment is attributable to the formative assessment activity rather than to normal instruction, practice effects, or regression to the mean.
- domain assumption Student self-reports and reflections accurately capture engagement, self-efficacy, and learning benefits.
Cite this review
Pith. "Pith review of Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course." pith.science (2026). https://pith.science/paper/BGRIX26M
@misc{pith2026250522479,
author = {Pith},
title = {Pith review of: Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course},
year = {2026},
howpublished = {\url{https://pith.science/paper/BGRIX26M}},
note = {Machine review of arXiv:2505.22479}
}
read the original abstract
This study explores how pre- and post-assessments shape learning outcomes in an Introductory Real Analysis course. Pre-assessments act as learning roadmaps, highlighting prior knowledge and guiding student focus, while post-assessments measure growth and conceptual mastery. By analyzing student performance and feedback, we assess their impact on engagement, self-efficacy, and deeper mathematical understanding. The findings offer valuable insights for enhancing instructional strategies and fostering a more effective, student-centered learning experience in advanced mathematics.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Pre/Post-Assessment Cycles in Calculus: Supporting Student Preparation, Reflection, and Learning
Repeated low-stakes pre/post-assessment cycles in honors Calculus I yielded large unit gains, full participation, and student reports of better preparation and self-monitoring.
Reference graph
Works this paper leans on
-
[1]
Balan, A. (2012). Assessment for learning: A case study in mathematics education (Doctoral dissertation, Malmö högskola, Fakulteten för lärande och samhälle). Berry, T. D. (2008). Pre-test assessment. American Journal of Business Education,
work page 2012
-
[2]
Rämö, J., Häsä, J., & Yan, Z. (2023, July). Students' self-assessment predictors and practices in an undergraduate mathematics course. In Thirteenth Congress of the European Society for Research in Mathematics Education (CERME13) (No. 20). Alfréd Rényi Institute of Mathematics; ERME. Sanders, S. (2019). A Brief Guide to Selecting and Using Pre-Post Assess...
work page 2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.