{"id":"8d638b04-2328-4ed8-932c-efc8f3adf8d1","arxiv_id":"2607.24736","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Given a free choice of instructor-provided vs self-created cheat sheets, students decide mainly on trust, personalization, and efficiency; formats change prep and use more than exam scores.","lead":"Students in a software requirements course chose between instructor-made and self-made exam cheat sheets; choices tracked trust, personalization, and efficiency more than grades. The study maps how that choice shapes preparation and use, not which format wins on scores.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The policy-leaning suggestion in §6.3 (\"encourage self-created sheets\") rests on contrasts between self-selected groups, and the paper's own Finding 1 shows the groups differ systematically on preparation time — so the observed final-exam trend cannot be attributed to format.","rationale":"The reader's weakest_assumption targeted the unpolished instructor sheet as a poor stand-in for \"instructor-provided sheets generally.\" That is a real generalizability caveat, and the paper partially concedes it in §7 (\"Nor can we separate students' reactions to the instructor-provided format from their reactions to the particular implementation\"). I judge it less load-bearing than selection confounding, for two reasons. First, the paper's headline claims are already hedged to be descriptive (\"choices relate more clearly to preparation... than to score differences\"), so the artifact critique mainly threatens an interpretation the paper does not strongly make. Second, the artifact critique threatens the null result, but the genuinely risky inference runs the other way: the non-null trend and the §6.3 recommendation depend on comparing self-selected groups, a threat the paper does not address anywhere — §7's limitations list course context, artifact analysis, and AI, but not confounding by preparation/ability, despite Finding 1 demonstrating the confound in the paper's own data. The reader's rationale does cite \"selection\" as a correctness-risk driver, so this is a sharpening of a concern the reader gestured at rather than a new axis: partial agreement. Because the reader already set CONDITIONAL with medium correctness risk, my read does not move the verdict; I recommend UNCHANGED, with the specific condition being that §6.3's encouragement of self-created sheets either be reframed as a study-habit hypothesis or be supported by the controlled analysis described in concrete_test. The qualitative contribution (RQ1 tensions, attitude shifts) stands independently of this and is the paper's real value.","tokens_in":15551,"tokens_out":2291,"duration_ms":90350,"concrete_test":"With the existing data (n≈44 for the final), re-run the type-vs-score comparison as an OLS or ANCOVA of final score on cheat sheet type controlling for (a) midterm score and (b) reported final prep-time bin. If the type coefficient shrinks toward zero or reverses once midterm score/prep time enter, the 7.7-point trend is composition, not format, and §6.3's recommendation should be softened. Supplement with a stratified check: within the \"11–20 hrs\" prep bin alone, compare means by sheet type; and report per-strategy percentages for Finding 4 out of total selections, not respondents, so the \"Select ALL\" data are not read as exclusive choices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claim (RQ1, the three tensions) is well-supported by the qualitative data and needs no attack. The soft spot is in the RQ2 quantitative contrasts that feed §6.3's recommendation. Students self-selected their format, and Finding 1 shows this selection is not random with respect to the key behavior: self-creators prepared significantly longer for the midterm (χ²=12.58, p=0.0056). Preparation time is therefore simultaneously the strongest \"effect\" the paper finds and a confounder for every other contrast. Finding 6's final-exam gap (69.5 vs 61.8, d=−0.63, p=0.13, with only ~9 instructor-sheet users among 44 respondents) is presented as a \"moderate effect size\" suggesting self-created sheets \"may have been more beneficial,\" and §6.3 converts this into advice that instructors \"encourage students to create their own.\" But the same gap is exactly what you would predict if more diligent students both self-create and score higher — no format effect required. The paper never reports a control for midterm score or prep time in the score comparisons. A secondary, smaller slip: Finding 4 reports strategy percentages (73%/23%/13%) as if mutually exclusive, yet Survey 3 Q6 is \"Select ALL that apply,\" so these are marginal rates, not a strategy distribution. Neither flaw touches the qualitative core, but the first one is load-bearing for the paper's only actionable prescription.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper reports a longitudinal survey study (three waves, 53/50/44 responses, 41-person complete cohort) in a senior undergraduate software requirements course where students could freely choose, for both midterm and final exams, between an instructor-provided cheat sheet and a self-created one (max one double-sided page). RQ1 is answered qualitatively: choices are shaped by trust in instructor expertise vs. personalization and control, and by effort/payoff reasoning, with attitudes shifting over the term (several students abandoned the instructor sheet after finding it sparse and topically narrow). RQ2 is answered quantitatively: self-creators prepared significantly longer for the midterm (χ²=12.58, p=0.0056) but not the final; self-created sheets had higher perceived final-exam coverage; exam scores did not differ significantly, though the final showed a non-significant trend favoring self-created sheets (69.5 vs 61.8, d=−0.63, p=0.13). The discussion frames cheat-sheet policy as a pedagogical design choice and suggests instructors consider encouraging self-created sheets. The qualitative core is well executed and the three-tension account is a real contribution; the quantitative-to-policy bridge is where the manuscript overreaches.","tokens_in":15874,"tokens_out":2651,"duration_ms":43359,"significance":"If the quantitative framing is repaired, this is a useful contribution to computing-education research on assessment design. The cheat-sheet literature has almost entirely treated format as an experimenter-fixed variable; treating it as a student choice, and documenting the trust/efficiency/personalization trade-offs students articulate, is genuinely new and practically relevant to instructors setting policy. Strengths worth naming: the naturalistic longitudinal design across two high-stakes exams; qualitative themes that are well grounded in participant quotes; honest reporting of null performance differences with effect sizes rather than significance-fishing; and an unusually candid §6.2 that discloses the instructor artifact's design intent, giving readers the material to judge generalizability. The study is a single-course, single-instrument convenience sample, so its contribution is descriptive and hypothesis-generating rather than prescriptive — the manuscript should be brought into line with that.","major_comments":[{"comment":"Finding 6 (§5) and §6.3: the final-exam contrast (69.5 vs 61.8, t=−1.65, p=0.13, d=−0.63) is described as a 'moderate effect size' suggesting self-created sheets 'may have been more beneficial,' and §6.3 converts this into the paper's only actionable recommendation ('instructors may consider encouraging students to create their own cheat sheets'). This is not supported by the design. Students self-selected format, and the paper's own Finding 1 shows selection is confounded with the key behavior: self-creators prepared significantly longer for the midterm (χ²=12.58, p=0.0056). The observed final-exam gap is exactly what one would predict if more diligent students both self-create and score higher; no format effect is needed. No analysis controls for preparation time or midterm score (midterm score is available and would be a natural covariate — the midterm groups were nearly identical at","section":"§5 Finding 6; §6.3"},{"comment":"The instructor-provided sheet in this offering is an unusual exemplar of its category: §6.2 reveals it was intentionally unpolished, slide-faithful, OCL/scaffold-focused, and — per P15 — 'basically only covered one topic (which was also the easiest topic)'; P30 and P48 echo the misalignment. This matters in two places. First, Finding 3 (self-created sheets had greater perceived coverage) may be a property of this artifact's sparse, single-topic layout rather than of instructor-provided sheets generally. Second, the qualitative theme 'trust in instructor judgment' was elicited under a design where the artifact was withheld until exam day; students were trusting a promise, not an artifact. §7 gestures at this ('we cannot separate students' reactions to the instructor-provided format from their reactions to the particular implementation'), but Findings 1, 3, and 6 and the §6.3 advice are st","section":"§6.2; Findings 1, 3, 6"}],"minor_comments":[{"comment":"The 73%/23%/13% figures are presented as if they partition respondents ('Other strategies... were much less common'), but Survey 3 Q6 is 'Select ALL that apply,' so these are marginal selection rates from non-exclusive options. A respondent could select both 'continuously' and 'skimmed at the beginning.' Please reword to make clear these are rates of endorsement per option, and note the implications for interpreting the 64%/50% 'occasional use' contrast with the 73% 'continuous reference' figure, which appear to be in mild tension.","section":"§5 Finding 4"},{"comment":"Preparation time is ordinal (four bins), and the Chi-square test of independence discards the ordering. An ordinal alternative (e.g., Mantel–Haenszel linear association or an ordinal logistic model) would be more powerful and more informative about direction. Relatedly, with cell sizes this small, please report the underlying contingency counts, not just χ² and p.","section":"§5 Finding 1"},{"comment":"The coding process (§3.1) is described at a high level. Given the Braun & Clarke citations, the authors presumably adopt a reflexive thematic analysis stance under which inter-rater reliability is not required — but this stance should be stated explicitly, along with the number of coders, how disagreements were resolved, and a codebook excerpt or code counts so readers can gauge theme prevalence (e.g., how many participants voiced 'trust in instructor' vs 'exam alignment').","section":"§3.1"},{"comment":"Cohen's d is reported as negative (d=−0.63) without stating the sign convention; given 'self-created' is the second group in the comparison, the sign presumably reflects instructor minus self-created, but this should be made explicit once. Also, Figure 2 would benefit from per-group n's in the caption.","section":"§5 Finding 6; Figure 2"},{"comment":"References [32] and [33] are the same paper in two publication states (online-first 2024 and the 2025 issue version). Please keep only one.","section":"References"},{"comment":"Within the cohort of 41, gender counts are 20 men, 19 women, 1 questioning, 1 undisclosed; the course enrolled 55. A brief note on whether respondents differed from non-respondents (e.g., by grade or prior cheat-sheet experience) would help assess self-selection into the survey itself, distinct from self-selection into format.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The instructor-provided sheet and its design rationale (§6.2) come from the authors' own course offering; readers may wish to know the relationship between the instructional team and the author team, which is never stated explicitly. This does not invalidate the work but bears on how neutrally the \"instructor intent\" material in §6.2 can be read. The venue fit (ACM TOCE-style computing education empirical work) seems good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real contribution here is the free-choice design plus three-wave tracking in a senior RE course. Prior work mostly fixes the policy and studies sheet quality or open-book comparisons; this paper lets students pick instructor-provided vs self-created and follows how they reason, switch, and use the sheets. That gap-fill is real and cleanly executed.\n\nWhat they do well: the qualitative themes (trust vs personalization, coverage vs clarity, effort vs payoff) are grounded in quotes and evolve across the term in a believable way. Methods match the claims—naturalistic choice, mixed items, team coding, chi-square on ordinal prep time, t-tests with Cohen’s d, honest nulls on grades. Writing is clear; limitations section is forthright about single-course scope and missing artifact comparison. No circularity, no invented constructs.\n\nSoft spots are real but bounded. Finding 1 shows self-creators prepared significantly longer for the midterm; that selection is never controlled when they later discuss the non-significant final-exam trend (d≈0.63, tiny instructor-sheet n) or when §6.3 suggests instructors “encourage” self-created sheets. The stress-test is right that diligence can explain the pattern without a format effect. Also minor: strategy percentages are reported as if exclusive while the item was multi-select. Neither sinks the descriptive core; both weaken the only actionable prescription. The particular instructor sheet (slide-faithful, withheld, intentionally unpolished) is another limit they partly own in §6.2, so generalizing “instructor-provided” is cautious at best.\n\nThis is for computing-education and assessment-design readers who care about student agency signals, not for anyone hunting a proven performance intervention. Math/stats are basic and correctly applied; citations cover the right prior strands. I would send it to peer review—publishable after tightening the causal language around RQ2 and dialing back the policy claim. Worth a look if you work on exam design or self-regulated learning in CS; skip if you only care about outcome RCTs.","headline":"Solid free-choice longitudinal survey on cheat-sheet agency; the qualitative core holds, but the policy lean toward self-created sheets overreaches self-selected prep-time confounds.","tokens_in":16831,"tokens_out":516,"would_cite":false,"duration_ms":12736,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When students can make or take a one-page cheat sheet, they choose between trust, personalization, and efficiency more than between better and worse scores.","keywords":["cheat sheets","exam preparation","student choice","student agency","self-regulated learning","assessment design","computing education","HCI"],"falsifier":"In a comparable course, give students the same free choice but use a clearly exam-aligned, well-laid-out instructor sheet (or run an artifact comparison of content overlap and density); if preferences, preparation-time gaps, and the null performance contrast flip or vanish, the format-choice story does not generalize beyond this implementation.","tokens_in":16689,"feed_emoji":"📝","tokens_out":858,"duration_ms":18116,"temperature":0.7,"pith_summary":"This paper asks what happens when senior software-requirements students may use either an instructor-provided one-page cheat sheet or one they build themselves for midterm and final. Across three survey waves, choices turn on trust in the instructor’s sense of what matters, desire to personalize content and layout, and how much preparation time students want to spend. Self-created sheets went with longer midterm study time and higher perceived coverage on the final, and most students treated the sheet as an ongoing reference rather than a last resort. Exam scores did not differ significantly by format. The authors argue that cheat-sheet policy is less a technical fairness fix than a pedagogical signal about agency, support, and what students are expected to construct for themselves.","feed_headline":"Students pick cheat sheets for trust, not higher scores","feed_subtitle":"Make-or-take choice tracks study time and coverage more than exam grades","key_machinery":"A longitudinal three-wave survey design in one senior requirements course, with naturalistic free choice of sheet format at midterm and final, combining closed items on time, use, and coverage with open-ended thematic coding of rationales, shifts, and constraints.","core_discovery":"Given a free choice between instructor-provided and self-created one-page cheat sheets, students’ preferences are shaped by recurring tensions—trust in instructor judgment versus their own knowledge needs, coverage versus clarity, and effort versus payoff—and those choices track preparation time, usage habits, and perceived coverage more clearly than statistically significant differences in exam performance.","pith_inferences":["If AI tools start drafting ‘self-created’ sheets, the learning value the paper ties to selection and compression may shrink unless courses redesign what counts as making the sheet.","The midterm-only preparation-time gap suggests format effects may be strongest early, before students recalibrate after feedback on the first exam.","Courses that want instructor sheets to compete on trust may need to show scaffold quality without turning the sheet into an answer key—the friction the instructor intended is itself a design variable."],"forward_implications":["Cheat-sheet policy should be treated as a design choice that grants or withholds student agency, not only as a fairness or cognitive-load rule.","Encouraging self-created sheets within tight page limits can function as structured study work even when scores do not rise.","Usage patterns matter: most students use sheets as continuous backup aids, so layout and legibility constrain real benefit.","Inclusive variants (typed sheets, other accommodations) are needed so personalization is not only available to students who can handwrite densely.","Null score differences do not mean the formats are interchangeable; they support different preparation contracts."],"fun_headline_variants":["Students choose cheat sheets for trust and prep, not grades","Make or take: cheat sheet picks track study time over scores","Trust vs personalization drives student cheat sheet choices","Cheat sheet preferences mirror prep habits more than exam results","Students weigh trust, coverage, effort in make-or-take sheets"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The particular instructor sheet in this course—slide-faithful, intentionally unpolished, and withheld until exam day—stands in for instructor-provided cheat sheets in general, so student reactions and null score gaps can be read as format effects rather than reactions to this sheet’s layout and coverage.","fun_headline_variants_meta":{"raw":{"variants":["Students choose cheat sheets for trust and prep, not grades","Make or take: cheat sheet picks track study time over scores","Trust vs personalization drives student cheat sheet choices","Cheat sheet preferences mirror prep habits more than exam results","Students weigh trust, coverage, effort in make-or-take sheets"]},"model":"grok-4.5","effort":"low","cost_usd":0.001768,"raw_usage":{"total_tokens":815,"prompt_tokens":726,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":17684000,"prompt_tokens_details":{"text_tokens":726,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":726,"tokens_out":69,"duration_ms":2392,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T06:24:25.037852+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a comparable course, give students the same free choice but use a clearly exam-aligned, well-laid-out instructor sheet (or run an artifact comparison of content overlap and density); if preferences, preparation-time gaps, and the null performance contrast flip or vanish, the format-choice story does not generalize beyond this implementation.","supporting_citations":[],"review_version":1}