Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read In an introductory real analysis course, a pre-/post-assessment cycle raised mean quiz scores from 6.39 to 11.11 and produced positive gains in engagement and self-efficacy.

desk verdict A useful teaching narrative with an unsupported causal headline; worth a review if reframed as a case study. read the letter →

arxiv 2505.22479 v1 pith:BGRIX26M submitted 2025-05-28 math.HO

classification math.HO
keywords pre-assessmentpost-assessmentformativeassessmentrealanalysisundergraduatemathematicseducationstudentself-efficacylimitslearningoutcomes
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that pairing an ungraded pre-assessment quiz with a matching post-assessment quiz helps students in an introductory real analysis course focus their study, reinforce key definitions, and see their own progress. The reported quantitative evidence is a mean score increase from 6.39 to 11.11, a 73.91% improvement, with a paired t-test p-value of 0.0013 and Cohen's d of 2.36. Student feedback describes better preparation, stronger confidence, and a clearer sense of what the course expects. The author presents the activity as a practical formative-assessment tool for proof-heavy mathematics, where conceptual grounding and self-efficacy are hard to build.

What carries the argument

The central mechanism is the paired pre-assessment and post-assessment cycle, with immediate self-grading, a written reflection prompt, and instructor-led solution discussion. The pre-assessment quiz is a short diagnostic given at the start of the limit unit; students grade their own responses, articulate a learning goal, and then receive five class sessions of proof-based instruction. The post-assessment mirrors the pre-assessment in structure and content, so the pre-to-post score difference is presented as a direct measure of learning gain. The reflection questions and self-grading are the active components that the paper credits with fostering metacognition and engagement.

What would settle it

Run the same unit with a control section that receives identical instruction but takes only the post-assessment; if the control post-assessment mean is statistically indistinguishable from the pre/post group's post-assessment mean, the claimed causal effect fails. Alternatively, give a delayed post-assessment several weeks after the unit: if the gain vanishes, the result reflects short-term familiarity rather than durable learning.

Watch

Extended reading notes

Core claim

The paper's central claim is that a structured, self-graded pre-assessment followed by instruction and a parallel post-assessment improves student learning, engagement, self-efficacy, and conceptual mastery in a real analysis unit on limits. It reports that among the 18 to 20 participating students, the mean score rose from 6.39 to 11.11, the median rose from 6.5 to 11.0, and the paired t-test gave $p = 0.0013$ with Cohen's $d = 2.36$. Students who initially struggled with concepts such as the uniqueness of a limit answered those questions correctly on the post-assessment, and feedback highlighted the value of reviewing definitions, self-grading, and instructor explanation. The author interprets these results as evidence that pre-assessments act as learning roadmaps and post-assessments as measures of growth, helping students transition from procedural calculus thinking to proof-based understanding.

Load-bearing premise

The paper assumes that the measured pre-to-post score gain is caused by the pre/post assessment activity itself, and not by the five class sessions of instruction, by practice effects from seeing similar questions twice, or by regression toward the mean.

Editorial extensions

If this is right

  • Instructors of proof-based courses can reuse the paired-quiz design as a no-stakes diagnostic that gives students a preview of unit content and a concrete target for study.
  • Because the quizzes are self-graded, the activity adds minimal grading burden and turns the assessment moment itself into a teaching opportunity.
  • If the claimed gain is real, the practice should generalize to other abstract units where definitions and proofs are the bottleneck, such as continuity or differentiability.
  • The positive feedback about confidence and exam preparation suggests the activity may carry motivational benefits separate from raw score gains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the study lacks a control group, a fair causal test would compare a section doing the pre/post cycle with a section receiving identical instruction and only the post-assessment; the paper's design cannot separate the assessment activity from the five lectures.
  • The self-grading and reflection prompt are plausible active ingredients; a follow-up that varies or removes those components would show whether the gain comes from the quizzes themselves or from the surrounding discussion.
  • A delayed post-assessment, given weeks after the unit, would test whether the gains reflect durable understanding or short-term familiarity with the mirrored questions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a classroom implementation of a structured pre-assessment, immediate self-grading and reflection, five lectures on limits, and a post-assessment in an introductory real analysis course (N=28). It presents descriptive statistics (pre mean 6.39, post mean 11.11), a paired t-test (p=0.0013), Cohen's d=2.36, and qualitative student feedback, and interprets these as evidence that the pre/post assessment activity improved learning outcomes, engagement, self-efficacy, and conceptual understanding. The appendices include the actual assessment instruments.

Significance. If the causal claims were supported, the study would provide a useful contribution to formative assessment in proof-based mathematics. The paper's strengths are the detailed description of the intervention, the inclusion of the actual quizzes in the appendices, and the honest reporting of some practical caveats, such as the difficulty of scheduling and participation (Section 4). However, the evidence is descriptive, and the inferential statistics do not identify the effect of the assessment activity because the five instructional sessions between the two measurements are a complete confound. The paper can be a valuable case study of how students respond to such an activity, but it cannot support the causal language used in the title, abstract, and discussion.

major comments (3)
  1. [Abstract; §2.1; §3 Results] The central claim that pre- and post-assessments "shape learning outcomes" and "have an impact" is not supported by the data. The pre-assessment was followed immediately by instructor-led solution discussion and reflection, then by five 50-minute lectures on exactly the tested material (§2.1). The post-assessment scores therefore reflect the combined effect of the assessment activity, the instruction, practice from seeing similar questions twice, and regression to the mean. A paired t-test (p=0.0013) and Cohen's d=2.36 only test whether the mean change differs from zero; they do not identify the assessment activity as the cause. The manuscript should remove causal language from the title, abstract, and Section 4.2, or provide a comparison condition; with the current design, the descriptive case-study framing is the only defensible one.
  2. [§3, Table 1] The statistical reporting is incomplete. The paper states that 18 of 28 students took the pre-assessment and 20 of 28 took the post-assessment, but it does not report how many students took both, the number of matched pairs, or the degrees of freedom for the paired t-test. If the means in Table 1 are computed from different student groups, a paired test is invalid; if only the matched subset was used, the missing-data mechanism needs discussion. The 73.91% improvement is computed from group means and may not represent a per-student gain. The paper should report the paired differences (mean, SD, 95% CI), the number of pairs, and a comparison of participants and nonparticipants. The effect size also needs a confidence interval.
  3. [§3.1.2, §3.1.3] Claims about engagement, self-efficacy, and conceptual understanding rest entirely on self-reported student comments and on pie-chart percentages whose denominators are not stated (e.g., "56.3% Much better" in Fig. 2). Self-reports collected by the instructor are susceptible to demand characteristics and do not constitute validated measures of the constructs named in the abstract. The number of respondents for each feedback question and the response rate should be reported, and the claims should be limited to "students perceived" rather than "the activity enhanced."
minor comments (4)
  1. [§1.1] There is a typo in the text: "assessment plays" is split as "play s"; please proofread the manuscript carefully.
  2. [Figure 1] Figure 1 lacks a descriptive caption and the x-axis categories are not clearly labeled; readers cannot tell which bars correspond to pre- and post-assessment without referring to the table.
  3. [References] Some references are incomplete, e.g., Berry (2008) lacks page numbers or a DOI, and the Balan (2012) entry would benefit from fuller bibliographic details consistent with the other references.
  4. [§4] The caveats section mentions absent students and timing issues but does not acknowledge the fundamental absence of a control condition; this should be stated explicitly as a limitation of the study.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the pre/post gain is an empirical observation that could have differed, and no claimed result is defined into existence or forced by a fit or self-citation.

full rationale

This paper makes no mathematical derivation and contains no fitted parameters, normalized quantities, or self-citations that could force its conclusions. The central quantitative result, 'The mean pre-assessment score was 6.39, while the mean post-assessment score increased to 11.11, indicating a 73.91% improvement' (Section 3), is an empirical before/after comparison: the post-assessment scores are not defined in terms of the pre-assessment scores, and the reported paired t-test (p=0.0013) and Cohen's d (2.36) are statistics computed from observations that could have come out differently. Nothing in the paper defines 'learning outcome' as 'post-assessment score' and then re-presents that score as evidence for itself; the claim is that the activity contributed to the gain, which is a causal inference, not a definitional identity. The main threats to that inference, namely the five intervening instruction sessions, practice effects from closely mirrored quizzes, and lack of a control group, are internal-validity limitations that are acknowledged only partially in Section 4's caveats about timing and absent students, but they are not circularity: they do not make the observed score change equal to the paper's input by construction. There are no self-citations, no imported uniqueness theorems, no ansatz smuggled in by citation, and no renaming of a known empirical pattern as an organizing result. The study is self-contained as an observational case study; its weakness is causal identification, not circular derivation. Therefore no circular step is identified and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim depends on the quiz scores being valid measures of learning, on the participants being representative, on the absence of a control group being inconsequential, and on student self-reports being truthful. These are all domain assumptions that the paper does not justify.

assumptions (4)
  • domain assumption Scores on the instructor-designed pre/post quizzes validly measure students' conceptual understanding of limits.
    The quizzes (Appendices A and B) were not validated against an external benchmark, yet all learning-gain claims are based on their scores.
  • domain assumption The students who completed the assessments are representative of the class and the paired t-test is based on actual pre/post pairs.
    Only 18 of 28 students took the pre-assessment and 20 took the post-assessment; the paper does not state how many students took both or handle missing data.
  • domain assumption Any improvement between pre- and post-assessment is attributable to the formative assessment activity rather than to normal instruction, practice effects, or regression to the mean.
    The study has no control group, yet Section 3.1.2 and Section 4.2 interpret the gain as evidence of the activity's impact.
  • domain assumption Student self-reports and reflections accurately capture engagement, self-efficacy, and learning benefits.
    The feedback form was administered by the instructor, and responses are quoted as evidence for improved confidence and understanding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course." pith.science (2026). https://pith.science/paper/BGRIX26M

@misc{pith2026250522479,
  author       = {Pith},
  title        = {Pith review of: Impact of Pre-Assessment and Post-Assessment in an Introductory Real Analysis Course},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BGRIX26M}},
  note         = {Machine review of arXiv:2505.22479}
}
read the original abstract

This study explores how pre- and post-assessments shape learning outcomes in an Introductory Real Analysis course. Pre-assessments act as learning roadmaps, highlighting prior knowledge and guiding student focus, while post-assessments measure growth and conceptual mastery. By analyzing student performance and feedback, we assess their impact on engagement, self-efficacy, and deeper mathematical understanding. The findings offer valuable insights for enhancing instructional strategies and fostering a more effective, student-centered learning experience in advanced mathematics.

Figures

Figures reproduced from arXiv: 2505.22479 by the authors.

Figure 1
Figure 1. Pre vs. Post Assessment Scores 3.1 Discussion 3.1.1 Results from the Pre-Assessment Quiz and the Reflection Question The pre-assessment quiz provided valuable insights into students’ prior knowledge and helped shape the approach to the course. Here are some key findings: (I) Self-Assessment and Reflection: Pre-Assessment Post-Assessment Scores 0 2 4 6 8 10 12 14 16 Standard Deviation 2.06 1.95 Effect Size 2.36 Paire… view at source ↗
Figure 2
Figure 2. Performance on the post-assessment Q2: Did the post-assessment quiz help reinforce your understanding of limits? [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Effect on understanding Q3: Did the pre-assessment and post-assessment quizzes together help you learn better? 56.30% 43.80% 0% Much better Somewhat better About the same Worse 47.10% 47.10% 5.80% Yes, a lot Yes, somewhat No [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Feedback on the overall activity Here is the summary of student feedback: (I) Preparation and Focus: One common theme in the feedback was that the assessment activity helped students get a clearer understanding of the course material and allowed them to focus their att…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pre/Post-Assessment Cycles in Calculus: Supporting Student Preparation, Reflection, and Learning

    math.HO 2026-07 conditional novelty 3.5 of 10

    Repeated low-stakes pre/post-assessment cycles in honors Calculus I yielded large unit gains, full participation, and student reports of better preparation and self-monitoring.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages · cited by 1 Pith paper

  1. [1]

    Balan, A. (2012). Assessment for learning: A case study in mathematics education (Doctoral dissertation, Malmö högskola, Fakulteten för lärande och samhälle). Berry, T. D. (2008). Pre-test assessment. American Journal of Business Education,

  2. [2]

    (2023, July)

    Rämö, J., Häsä, J., & Yan, Z. (2023, July). Students' self-assessment predictors and practices in an undergraduate mathematics course. In Thirteenth Congress of the European Society for Research in Mathematics Education (CERME13) (No. 20). Alfréd Rényi Institute of Mathematics; ERME. Sanders, S. (2019). A Brief Guide to Selecting and Using Pre-Post Assess...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.