Pith. sign in

REVIEW 2 major objections 6 minor 34 references

When students can make or take a one-page cheat sheet, they choose between trust, personalization, and efficiency more than between better and worse scores.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:24 UTC pith:HZIKUUBC

load-bearing objection Solid free-choice longitudinal survey on cheat-sheet agency; the qualitative core holds, but the policy lean toward self-created sheets overreaches self-selected prep-time confounds. the 2 major comments →

arxiv 2607.24736 v1 pith:HZIKUUBC submitted 2026-07-27 cs.HC

Make or Take: How Students Navigate Self-Created and Instructor-Provided Cheat Sheets

classification cs.HC
keywords cheat sheetsexam preparationstudent choicestudent agencyself-regulated learningassessment designcomputing educationHCI
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks what happens when senior software-requirements students may use either an instructor-provided one-page cheat sheet or one they build themselves for midterm and final. Across three survey waves, choices turn on trust in the instructor’s sense of what matters, desire to personalize content and layout, and how much preparation time students want to spend. Self-created sheets went with longer midterm study time and higher perceived coverage on the final, and most students treated the sheet as an ongoing reference rather than a last resort. Exam scores did not differ significantly by format. The authors argue that cheat-sheet policy is less a technical fairness fix than a pedagogical signal about agency, support, and what students are expected to construct for themselves.

Core claim

Given a free choice between instructor-provided and self-created one-page cheat sheets, students’ preferences are shaped by recurring tensions—trust in instructor judgment versus their own knowledge needs, coverage versus clarity, and effort versus payoff—and those choices track preparation time, usage habits, and perceived coverage more clearly than statistically significant differences in exam performance.

What carries the argument

A longitudinal three-wave survey design in one senior requirements course, with naturalistic free choice of sheet format at midterm and final, combining closed items on time, use, and coverage with open-ended thematic coding of rationales, shifts, and constraints.

Load-bearing premise

The particular instructor sheet in this course—slide-faithful, intentionally unpolished, and withheld until exam day—stands in for instructor-provided cheat sheets in general, so student reactions and null score gaps can be read as format effects rather than reactions to this sheet’s layout and coverage.

What would settle it

In a comparable course, give students the same free choice but use a clearly exam-aligned, well-laid-out instructor sheet (or run an artifact comparison of content overlap and density); if preferences, preparation-time gaps, and the null performance contrast flip or vanish, the format-choice story does not generalize beyond this implementation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Cheat-sheet policy should be treated as a design choice that grants or withholds student agency, not only as a fairness or cognitive-load rule.
  • Encouraging self-created sheets within tight page limits can function as structured study work even when scores do not rise.
  • Usage patterns matter: most students use sheets as continuous backup aids, so layout and legibility constrain real benefit.
  • Inclusive variants (typed sheets, other accommodations) are needed so personalization is not only available to students who can handwrite densely.
  • Null score differences do not mean the formats are interchangeable; they support different preparation contracts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If AI tools start drafting ‘self-created’ sheets, the learning value the paper ties to selection and compression may shrink unless courses redesign what counts as making the sheet.
  • The midterm-only preparation-time gap suggests format effects may be strongest early, before students recalibrate after feedback on the first exam.
  • Courses that want instructor sheets to compete on trust may need to show scaffold quality without turning the sheet into an answer key—the friction the instructor intended is itself a design variable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper reports a longitudinal survey study (three waves, 53/50/44 responses, 41-person complete cohort) in a senior undergraduate software requirements course where students could freely choose, for both midterm and final exams, between an instructor-provided cheat sheet and a self-created one (max one double-sided page). RQ1 is answered qualitatively: choices are shaped by trust in instructor expertise vs. personalization and control, and by effort/payoff reasoning, with attitudes shifting over the term (several students abandoned the instructor sheet after finding it sparse and topically narrow). RQ2 is answered quantitatively: self-creators prepared significantly longer for the midterm (χ²=12.58, p=0.0056) but not the final; self-created sheets had higher perceived final-exam coverage; exam scores did not differ significantly, though the final showed a non-significant trend favoring self-created sheets (69.5 vs 61.8, d=−0.63, p=0.13). The discussion frames cheat-sheet policy as a pedagogical design choice and suggests instructors consider encouraging self-created sheets. The qualitative core is well executed and the three-tension account is a real contribution; the quantitative-to-policy bridge is where the manuscript overreaches.

Significance. If the quantitative framing is repaired, this is a useful contribution to computing-education research on assessment design. The cheat-sheet literature has almost entirely treated format as an experimenter-fixed variable; treating it as a student choice, and documenting the trust/efficiency/personalization trade-offs students articulate, is genuinely new and practically relevant to instructors setting policy. Strengths worth naming: the naturalistic longitudinal design across two high-stakes exams; qualitative themes that are well grounded in participant quotes; honest reporting of null performance differences with effect sizes rather than significance-fishing; and an unusually candid §6.2 that discloses the instructor artifact's design intent, giving readers the material to judge generalizability. The study is a single-course, single-instrument convenience sample, so its contribution is descriptive and hypothesis-generating rather than prescriptive — the manuscript should be brought into line with that.

major comments (2)
  1. [§5 Finding 6; §6.3] Finding 6 (§5) and §6.3: the final-exam contrast (69.5 vs 61.8, t=−1.65, p=0.13, d=−0.63) is described as a 'moderate effect size' suggesting self-created sheets 'may have been more beneficial,' and §6.3 converts this into the paper's only actionable recommendation ('instructors may consider encouraging students to create their own cheat sheets'). This is not supported by the design. Students self-selected format, and the paper's own Finding 1 shows selection is confounded with the key behavior: self-creators prepared significantly longer for the midterm (χ²=12.58, p=0.0056). The observed final-exam gap is exactly what one would predict if more diligent students both self-create and score higher; no format effect is needed. No analysis controls for preparation time or midterm score (midterm score is available and would be a natural covariate — the midterm groups were nearly identical at
  2. [§6.2; Findings 1, 3, 6] The instructor-provided sheet in this offering is an unusual exemplar of its category: §6.2 reveals it was intentionally unpolished, slide-faithful, OCL/scaffold-focused, and — per P15 — 'basically only covered one topic (which was also the easiest topic)'; P30 and P48 echo the misalignment. This matters in two places. First, Finding 3 (self-created sheets had greater perceived coverage) may be a property of this artifact's sparse, single-topic layout rather than of instructor-provided sheets generally. Second, the qualitative theme 'trust in instructor judgment' was elicited under a design where the artifact was withheld until exam day; students were trusting a promise, not an artifact. §7 gestures at this ('we cannot separate students' reactions to the instructor-provided format from their reactions to the particular implementation'), but Findings 1, 3, and 6 and the §6.3 advice are st
minor comments (6)
  1. [§5 Finding 4] The 73%/23%/13% figures are presented as if they partition respondents ('Other strategies... were much less common'), but Survey 3 Q6 is 'Select ALL that apply,' so these are marginal selection rates from non-exclusive options. A respondent could select both 'continuously' and 'skimmed at the beginning.' Please reword to make clear these are rates of endorsement per option, and note the implications for interpreting the 64%/50% 'occasional use' contrast with the 73% 'continuous reference' figure, which appear to be in mild tension.
  2. [§5 Finding 1] Preparation time is ordinal (four bins), and the Chi-square test of independence discards the ordering. An ordinal alternative (e.g., Mantel–Haenszel linear association or an ordinal logistic model) would be more powerful and more informative about direction. Relatedly, with cell sizes this small, please report the underlying contingency counts, not just χ² and p.
  3. [§3.1] The coding process (§3.1) is described at a high level. Given the Braun & Clarke citations, the authors presumably adopt a reflexive thematic analysis stance under which inter-rater reliability is not required — but this stance should be stated explicitly, along with the number of coders, how disagreements were resolved, and a codebook excerpt or code counts so readers can gauge theme prevalence (e.g., how many participants voiced 'trust in instructor' vs 'exam alignment').
  4. [§5 Finding 6; Figure 2] Cohen's d is reported as negative (d=−0.63) without stating the sign convention; given 'self-created' is the second group in the comparison, the sign presumably reflects instructor minus self-created, but this should be made explicit once. Also, Figure 2 would benefit from per-group n's in the caption.
  5. [References] References [32] and [33] are the same paper in two publication states (online-first 2024 and the 2025 issue version). Please keep only one.
  6. [§3.2] Within the cohort of 41, gender counts are 20 men, 19 women, 1 questioning, 1 undisclosed; the course enrolled 55. A brief note on whether respondents differed from non-respondents (e.g., by grade or prior cheat-sheet experience) would help assess self-selection into the survey itself, distinct from self-selection into format.

Circularity Check

0 steps flagged

Empirical survey study with no derivation-by-construction; findings are coded observations and group comparisons, not fitted inputs renamed as predictions.

full rationale

This paper is a longitudinal survey study (three waves; n up to 53/50/44) of student choice between instructor-provided and self-created cheat sheets in one software-requirements course. RQ1–RQ2 are answered by thematic coding of open responses and descriptive/inferential comparisons (χ² on prep-time bins; t-tests on exam scores; reported usage and coverage frequencies). Nothing in the chain defines a quantity in terms of the target it then “predicts,” fits a parameter to data and re-labels the fit as an out-of-sample prediction, or rests a uniqueness/forced-choice claim on a self-citation. The instructor’s design rationale (§6.2) and the single-course context are scope limitations, not circular reductions. Self-selection and the prep-time confound noted by the skeptic affect causal attribution of score trends, which is a validity concern outside the circularity taxonomy. No circular steps identified.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

Load-bearing premises are standard social-science and course-study assumptions: voluntary survey responses (with grade incentive) validly reflect decision factors; thematic codes capture stable constructs; one RE course’s exams and one instructor sheet instantiate the two ‘formats’; self-selected sheet type groups are comparable enough for descriptive and simple inferential contrasts. No physical constants or fitted theoretical parameters; free parameters are design knobs (page limit, bonus marks, sheet withheld until exam).

free parameters (3)
  • cheat_sheet_page_limit = 1 double-sided page
    Maximum one double-sided page constrains density and fairness arguments; chosen by course policy, not estimated from data.
  • participation_bonus_marks = 5 marks max
    Up to 5 bonus marks across three surveys (1+1+3) can shift who responds and how carefully; fixed by study design.
  • instructor_sheet_content_policy = slide-faithful scaffold; withheld pre-exam
    Sheet intentionally mirrors class slides with minimal editing and is hidden pre-exam; this implementation choice drives many ‘misalignment’ and trust responses.
axioms (5)
  • domain assumption Self-reported preparation time, referral frequency, coverage percentages, and open rationales are sufficiently valid indicators of student decision-making and exam behavior for the stated RQs.
    Underpins all Findings 1–5 and qualitative §4; no behavioral proctoring or sheet-use logs.
  • domain assumption Collaborative thematic coding of open responses yields themes that generalize within the longitudinal cohort beyond idiosyncratic wording.
    §3.1 describes deductive+inductive coding without reported IRR; themes carry RQ1.
  • domain assumption Students who choose different sheet types are comparable on unmeasured confounds (prior ability, engagement, note quality) enough that prep-time and grade contrasts are interpretable as related to sheet choice.
    No random assignment (§3.2); t-tests/χ² treat groups as observational arms.
  • standard math Standard chi-square tests of independence and two-sample t-tests with Cohen’s d are appropriate for the ordinal prep bins and grade outcomes reported.
    §5 Findings 1 and 6; small cell counts possible given ~20–30% instructor-sheet users.
  • ad hoc to paper Withholding the instructor sheet until the exam isolates ‘format preference’ from copying the official artifact during preparation.
    Explicit design choice §3.2; shapes ecological validity and trust/personalization themes.

pith-pipeline@v1.2.0-grok45-kimik3 · 19867 in / 3469 out tokens · 84654 ms · 2026-07-31T06:24:25.037852+00:00 · methodology

0 comments
read the original abstract

The use of cheat sheets in exams is often framed as a way to reduce cognitive load and support student performance. However, little is known about how students choose between self-created and instructor-provided cheat sheets, or how these choices relate to their broader approaches to exam preparation. We conducted a longitudinal study in a senior-level undergraduate software requirements course, where students could use either an instructor-provided or a self-created cheat sheet for both the midterm and final exams. Across three survey waves, we received 53, 50, and 44 responses, respectively. 41 students completed all three surveys and formed the longitudinal cohort used to examine how choices and experiences evolved over time, while exam-specific analyses used all available responses from the corresponding wave. Our findings identify several considerations that shaped students' choices, including trust in instructor expertise, the desire for personalization, and preparation efficiency. We further show how students' attitudes shifted over time and how their preferences were reflected in patterns of cheat sheet use, perceived content coverage, and challenges encountered during the exams.

Figures

Figures reproduced from arXiv: 2607.24736 by Helen Weixu Chen, Lesley Istead, Victoria Sakhnini.

Figure 1
Figure 1. Figure 1: Distribution of preparation time, frequency of cheat sheet use, and final exam coverage by cheat sheet type. [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Grade distributions for midterm and final exams by cheat sheet type. [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 17 canonical work pages

  1. [1]

    Daniel Berry, Ricardo Gacitua, Pete Sawyer, and Sri Fatimah Tjong. 2012. The case for dumb requirements engineering tools. InProceedings of the 18th International Conference on Requirements Engineering: Foundation for Software Quality(Essen, Germany)(REFSQ’12). Springer-Verlag, Berlin, Heidelberg, 211–217. doi:10.1007/978-3-642-28714-5_18

  2. [2]

    Barry Boehm, Alexander Egyed, Dan Port, Archita Shah, Julie Kwan, and Ray Madachy. 1998. A stakeholder win–win approach to software engineering education.Annals of Software Engineering6, 1 (1998), 295–321. doi:10.1023/A:1018988827405

  3. [3]

    David Boniface. 1985. Candidates’ use of notes and textbooks during an open-book examination.Educational Research27, 3 (1985), 201–209. arXiv:https://doi.org/10.1080/0013188850270307 doi:10.1080/0013188850270307

  4. [4]

    Virginia Braun and Victoria Clarke. 2021. One size fits all? What counts as quality practice in (reflexive) thematic analysis?Qualitative Research in Psychology18, 3 (2021), 328–352. arXiv:https://doi.org/10.1080/14780887.2020.1769238 doi:10.1080/14780887.2020.1769238

  5. [5]

    2014.Thematic Analysis

    Victoria Clarke and Virginia Braun. 2014.Thematic Analysis. Springer New York, New York, NY, 1947–1952. doi:10.1007/978-1-4614-5583-7_311

  6. [6]

    2019.How Do I Handle Academic Integrity Issues?Springer Singapore, Singapore, 689–726

    Ray Cooksey and Gael McDonald. 2019.How Do I Handle Academic Integrity Issues?Springer Singapore, Singapore, 689–726. doi:10.1007/978-981- 13-7747-1_15

  7. [7]

    Shant Aram Danielian and Natascha Trellinger Buswell. 2019. Do support sheets actually support students? A content analysis of student support sheets for exams. In2019 Pacific Southwest Section Meeting. ASEE Conferences, California State University, Los Angeles , California. https://peer.asee.org/31824

  8. [8]

    Michael de Raadt. 2012. Student created cheat-sheets in examinations: impact on student outcomes. InProceedings of the Fourteenth Australasian Computing Education Conference - Volume 123(Melbourne, Australia)(ACE ’12). Australian Computer Society, Inc., AUS, 71–76

  9. [9]

    Deterding and Mary C

    Nicole M. Deterding and Mary C. Waters. 2021. Flexible Coding of In-depth Interviews: A Twenty-first-century Approach.Sociological Methods & Research50, 2 (2021), 708–739. arXiv:https://doi.org/10.1177/0049124118799377 doi:10.1177/0049124118799377

  10. [10]

    Laurie Dickson and Michelle D

    K. Laurie Dickson and Michelle D. Miller. 2005. Authorized Crib Cards Do Not Improve Exam Performance.Teaching of Psychology32, 4 (2005), 230–233. arXiv:https://doi.org/10.1207/s15328023top3204_6 doi:10.1207/s15328023top3204_6

  11. [11]

    Brigitte Erbe. 2007. Reducing Test Anxiety While Increasing Learning: The Cheat Sheet.College Teaching55, 3 (2007), 96–98. arXiv:https://doi.org/10.3200/CTCH.55.3.96-98 doi:10.3200/CTCH.55.3.96-98

  12. [12]

    Feldhusen

    John F. Feldhusen. 1961. An Evaluation of College Students’ Reactions to Open Book Examinations.Educational and Psychological Measurement21, 3 (1961), 637–646. arXiv:https://doi.org/10.1177/001316446102100310 doi:10.1177/001316446102100310

  13. [13]

    Funk and K

    Steven C. Funk and K. Laurie Dickson. 2011. Crib Card Use During Tests: Helpful or a Crutch?Teaching of Psychology38, 2 (2011), 114–117. arXiv:https://doi.org/10.1177/0098628311401584 doi:10.1177/0098628311401584

  14. [14]

    Afshin Gharib and William Phillips. 2013. Test Anxiety, Student Preferences and Performance on Different Exam Types in Introductory Psychology. International Journal of e-Education, e-Business, e-Management and e-Learning3, 1 (2013), 1–6. Manuscript submitted to ACM Chen et al

  15. [15]

    Afshin Gharib, William Phillips, and Noelle Mathew. 2012. Cheat Sheet or Open-Book? A Comparison of the Effects of Exam Types on Performance, Retention, and Anxiety.Online Submission2, 8 (2012), 469–478. doi:10.17265/2159-5542/2012.08.004

  16. [16]

    Sally Hamouda and Clifford A Shaffer. 2016. Crib sheets and exam performance in a data structures course.Computer Science Education26, 1 (2016), 1–26. doi:10.1080/08993408.2016.1140427

  17. [17]

    Douglas C. Hindman. 1980. Crib Notes in the Classroom: Cheaters Never Win.Teaching of Psychology7, 3 (1980), 166–168. arXiv:https://doi.org/10.1207/s15328023top0703_10 doi:10.1207/s15328023top0703_10

  18. [18]

    Lynne Kendall. 2018. Supporting students with disabilities within a UK university: Lecturer perspectives.Innovations in Education and Teaching International55, 6 (2018), 694–703. arXiv:https://doi.org/10.1080/14703297.2017.1299630 doi:10.1080/14703297.2017.1299630

  19. [19]

    Nurassyl Kerimbayev, Zhanat Umirzakova, Rustam Shadiev, and Vladimir Jotsov. 2023. A student-centered approach using modern technologies in distance learning: a systematic review of the literature.Smart Learning Environments10, 1 (2023), 61. doi:10.1186/s40561-023-00280-8

  20. [20]

    Karen Larwin. 2012. Student Prepared Testing Aids: A Low-Tech Method of Encouraging Student Engagement.Journal of Instructional Psychology 39, 2 (2012). https://link.gale.com/apps/doc/A321057800/AONE?u=anon~e8d61d35&sid=googleScholar&xid=da8cff10

  21. [21]

    Pam A Mueller and Daniel M Oppenheimer. 2014. The pen is mightier than the keyboard: Advantages of longhand over laptop note taking. Psychological science25, 6 (2014), 1159–1168. doi:10.1177/0956797614524581

  22. [22]

    Norheim, Eric Rebentisch, Dekai Xiao, Lorenz Draeger, Alain Kerbrat, and Olivier L

    Johannes J. Norheim, Eric Rebentisch, Dekai Xiao, Lorenz Draeger, Alain Kerbrat, and Olivier L. de Weck. 2024. Challenges in applying large language models to requirements engineering tasks.Design Science10 (2024), e16. doi:10.1017/dsj.2024.8

  23. [23]

    Paul R Pintrich. 2002. The role of metacognitive knowledge in learning, teaching, and assessing.Theory into practice41, 4 (2002), 219–225. doi:10.1207/s15430421tip4104_3

  24. [24]

    Michael Prince. 2004. Does Active Learning Work? A Review of the Research.Journal of Engineering Education93, 3 (2004), 223–231. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/j.2168-9830.2004.tb00809.x doi:10.1002/j.2168-9830.2004.tb00809.x

  25. [25]

    Settlage and Jim R

    Daniel M. Settlage and Jim R. Wollscheid. 2019. An Analysis of the Effect of Student Prepared Notecards on Exam Performance.College Teaching67, 1 (2019), 15–22. arXiv:https://doi.org/10.1080/87567555.2018.1514485 doi:10.1080/87567555.2018.1514485

  26. [26]

    Jaswinder Pal Singh, Neha Mishra, and Krishna Kumar Mishra. 2025. Promoting Academic Integrity: Strategies and Challenges in Higher Education. InHigher Education and Quality Assurance Practices. IGI Global Scientific Publishing, 191–224. doi:10.4018/979-8-3693-6765-0.ch007

  27. [27]

    Long, Bruce W

    Murali Sitaraman, Timothy J. Long, Bruce W. Weide, E. James Harpner, and Liqing Wang. 2001. A formal approach to component-based software engineering: education and evaluation. InProceedings of the 23rd International Conference on Software Engineering(Toronto, Ontario, Canada)(ICSE ’01). IEEE Computer Society, USA, 601–609

  28. [28]

    Raymond L Smith and Henry D Lester. 2019. Instructor and Student Perceptions of the Authorized, Self-prepared Reference Sheet for Examinations. In2019 ASEE Annual Conference & Exposition. doi:10.18260/1-2--32977

  29. [29]

    Yang Song and David Thuente. 2015. A quantitative case study in engineering of the efficacy of quality cheat-sheets. In2015 IEEE Frontiers in Education Conference (FIE). 1–7. doi:10.1109/FIE.2015.7344082

  30. [30]

    Valentin Sorescu and Diana Andone. 2025. Developing Intelligent Tutors with Artificial Intelligence for Digital Education.2025 IEEE Digital Education and MOOCS Conference (DEMOcon)(2025), 1–6. http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11282632

  31. [31]

    Antony Tang. 2007. A rationale-based model for architecture design reasoning. (1 2007). doi:10.25916/sut.26284492.v1

  32. [33]

    Sophie Whitehouse, John Bredican, Jayne Heaford, Anouk De Regt, and Kirk Plangger. 2025. Implications of crib sheets for student well-being and exam performance: a mixed methods study.Studies in Higher Education50, 12 (2025), 2838–2856. arXiv:https://doi.org/10.1080/03075079.2024.2435563 doi:10.1080/03075079.2024.2435563

  33. [34]

    Wünsche, Dominik Lange-Nawka, Zixuan Wang, Steffan Hooper, Samuel E

    Burkhard C. Wünsche, Dominik Lange-Nawka, Zixuan Wang, Steffan Hooper, Samuel E. R. Thompson, and Tony Haoran Feng. 2025. Characteristics and Effectiveness of Cheat Sheets for a Third-year Computer Graphics and Image Processing Course. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2(Pittsburgh, PA, USA)(SIGCSETS 2025)....

  34. [35]

    Zimmerman

    Barry J. Zimmerman. 2002. Becoming a Self-Regulated Learner: An Overview.Theory Into Practice41, 2 (2002), 64–70. arXiv:https://doi.org/10.1207/s15430421tip4102_2 doi:10.1207/s15430421tip4102_2 Manuscript submitted to ACM Make or Take: How Students Navigate Self-Created and Instructor-Provided Cheat Sheets A Survey Questions A.1 Survey 1 (1) Please select...