{"id":"97766ee5-9765-4bed-bbde-766dba745bd7","arxiv_id":"2607.16475","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Instructors' AI policies are mostly assessment-protection measures (paper exams, exam-heavy grading) that create policing burden and relationship strain, with learning-oriented alternatives available.","lead":"This qualitative study interviewed 13 U.S. computer science instructors about their generative-AI course policies. It finds that most policies focus on protecting exams from AI rather than on helping students learn, and that policing AI use strains instructor-student relationships.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'primarily' claim about assessment-oriented AI policies is a population-frequency statement unsupported by the 13-instructor non-representative sample; Section 3.3 explicitly disclaims representativeness.","rationale":"The paper is a thoughtful qualitative study with rich quotes and an honest reflexive methodology. The reader's weakest assumption—that the analysis rests on unvalidated self-reports from a small, self-selected sample—captures a genuine limitation. My stress-test sharpens this: the most load-bearing issue is not whether self-reports are accurate, but whether the abstract's 'primarily' claim is a legitimate inference from a purposive sample that Section 3.3 explicitly says is not representative. Even if every interviewee were perfectly accurate, the study cannot establish that assessment-oriented policies are the dominant response among CS instructors generally. This matters because the contribution's headline is precisely that dominance claim. The study's qualitative insights—the learning/assessment harm distinction, second-order costs, learning-oriented alternatives—do not depend on exact prevalence and remain valuable. Therefore the verdict should be CONDITIONAL: the paper should be accepted after the abstract and conclusion revise the frequency language to be explicitly sample-bound, or after a broader quantitative test supports the prevalence claim. This is a concrete, good-faith concern about the boundary between hypothesis generation and empirical prevalence.","tokens_in":24493,"tokens_out":4826,"duration_ms":56848,"concrete_test":"Administer a pre-registered survey (or systematic syllabus analysis) to a broad, non-snowball sample of ~200 US undergraduate CS instructors, coding each course's AI policy using the paper's assessment-oriented vs. learning-oriented rubric. If assessment-oriented policies (e.g., paper exams, AI bans, exam-passing requirements) do not clearly outnumber learning-oriented or mixed policies in the 95% confidence interval (say, below 60%), then the abstract's 'primarily' claim should be revised to 'in this sample' or 'among interviewed instructors.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—'AI policies primarily seek to AI-proof assessments' (abstract)—is a prevalence assertion. It is supported only by 13 purposively and snowball-recruited instructors (Section 3.1, Table 2), with no claim to representativeness: Section 3.3 states the findings 'are not meant to be representative of a broader population.' Yet the abstract and Section 4.2.2 ('surprisingly common, with 9 interviewees adopting some form of it') use the sample to make a population-level frequency statement. The authors' own literature review (Section 2.2) cites prior syllabus analyses documenting wide variability in policies [3, 16, 73], so self-selection into a study about AI policies could plausibly over-recruit instructors with strong assessment concerns. If the sample is skewed, the headline finding—that assessment AI-proofing dominates and imposes policing burden—may be a property of this sample, not of CS instruction generally. The self-report accuracy issue flagged by the reader is real but secondary: even if every interview were fully truthful, the 'primarily' claim would still lack support. This is not a disagreement with consensus; it is an internal tension between the method's stated non-representativeness and the abstract's frequency language.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This qualitative study reports a reflexive thematic analysis of 13 semi-structured interviews with U.S. undergraduate computer science instructors about their generative AI course policies. The authors distinguish two perceived types of harm: learning harms to students (cognitive offloading, illusion of competence, social isolation) and assessment harms to instructors (inability to verify student work). They argue that most instructors respond to assessment harms by AI-proofing assessments (e.g., switching to paper exams, increasing exam weight), which creates second-order costs: policing burden on instructors, strained instructor–student relationships, and a shift of responsibility for learning onto students. The paper then proposes learning-oriented alternatives, such as transparent AI instruction, modeling good AI use, formative assessments for metacognition, and motivation-focused course design. The manuscript is explicitly exploratory and disclaims representativeness (Section 3.3).","tokens_in":24738,"tokens_out":4756,"duration_ms":51778,"significance":"If the thematic claims hold, the study makes a useful contribution by shifting attention from AI tools themselves to the instructor–student relationship, a dimension often missing in tool-focused GenAI education research. The learning-harm vs. assessment-harm distinction is a clear conceptual contribution. The paper's method is transparent: it documents purposive/snowball sampling, interview protocol, reflexive thematic analysis, positionality, and extensive verbatim quotes that ground the analysis. The explicit limitations and the non-representativeness disclaimer (Section 3.3) are commendable. The recommendations in Section 5.2 are clearly framed as hypotheses requiring further empirical validation. Overall, this is a well-scoped exploratory study whose central claims are supported by the reported data, provided the language is carefully qualified.","major_comments":[{"comment":"The abstract's claim that AI policies \"primarily seek to AI-proof assessments\" is a frequency assertion that outruns the sample. Section 3.3 explicitly states the findings are \"not meant to be representative of a broader population,\" and recruitment was purposive and snowball-based (Section 3.1). Section 4.2.2 reports that \"9 interviewees\" adopted paper exams and calls this \"surprisingly common,\" which implies population-level prevalence. The qualitative analysis can support a claim about the interviewed instructors, but the current wording invites over-generalization. Please qualify the abstract and Section 4.2.2 (e.g., \"among the instructors we interviewed\") and soften the prevalence language in the recommendations.","section":"Abstract and Section 4.2.2"}],"minor_comments":[{"comment":"The Limitations section is truncated mid-sentence: \"skills like memorizing syntax might be less necessary for some students, such as\". Complete the sentence or remove the dangling phrase.","section":"Section 5.3"},{"comment":"Typo: \"asksed\" should be \"asked\".","section":"Section 3.1"},{"comment":"Reference [72] lists an author as \"mark w uci\" — likely a placeholder that should be corrected.","section":"References"},{"comment":"The phrase \"surprisingly common\" is ambiguous; as noted above, it is appropriate to describe the sample, but the paper should avoid implying a broader base rate given the non-representative sampling.","section":"Section 4.2.2"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid exploratory qualitative study, but the authors should make the abstract and any frequency-implying language consistent with their own non-representativeness disclaimer. This is a local fix and should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing worth knowing about this paper is that it's a solid, carefully done qualitative study of a real gap: how CS instructors actually design, implement, and enforce generative AI policies, and how those policies affect instructor-student relationships. The analytic split between learning harms (to students) and assessment harms (to instructors) is genuinely useful, and the finding that assessment-oriented policies like paper exams create policing burden and relationship strain is well supported by the interview data. The paper also does something rare: it gives concrete, low-cost learning-oriented alternatives, and it explicitly does not claim empirical proof for them. Credit where due: 13 interviews across nine institutions, rich quotes, reflexive thematic analysis done properly, and a limitations section that acknowledges self-report and non-generalizability. The authors know their method and stay within its interpretive frame most of the time.\n\nThe soft spots are real but not fatal. The biggest one is a mismatch between the abstract and the method. Section 3.3 says the findings 'are not meant to be representative of a broader population,' yet the abstract states that 'AI policies primarily seek to AI-proof assessments.' That's a population-level frequency claim, and 13 purposively and snowball-recruited instructors cannot support it. Same with 'surprisingly common, with 9 interviewees' — a sample count presented as if it tells us something about the field. Given self-selection into a study about AI policies, instructors with strong assessment concerns are plausibly over-recruited. The fix is easy: soften the abstract's 'primarily' and frame the pattern as emergent from this sample, not a prevalence finding. The self-report issue is secondary but worth noting; the paper already acknowledges it in 5.3, so I won't beat it. Minor editing slip: Section 5.3 has an unfinished sentence ending \"such as,\" presumably a lost clause.\n\nBottom line: this is a well-executed qualitative study that deserves to be published and should get a serious referee. It advances the CS-education AI literature by shifting focus from tool effects to instructor-student relationships, and it generates testable recommendations. If I were reviewing, I'd ask for the abstract's language to be aligned with the non-representative design, but I would not request new data. Worth a reading group slot for anyone working on AI policy or qualitative methods in computing education.","headline":"A genuinely useful qualitative study of how CS instructors design and enforce GenAI policies, but the abstract's 'primarily' frequency claim overstates what 13 self-selected interviews can support.","tokens_in":25264,"tokens_out":1229,"would_cite":true,"duration_ms":15665,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Undergraduate CS instructors, the paper argues, have mainly tried to AI-proof assessments, and this choice adds policing burden and worsens instructor-student relationships while learning harms go unaddressed.","keywords":["computer science education","generative AI policy","assessment integrity","instructor-student relationship","learning harms","semi-structured interviews","higher education","cognitive offloading"],"falsifier":"A multi-section comparison would settle it: if courses with strict AI-proofing (paper exams, exam-weighted grades) produced equal or better measures of long-term coding skill, self-efficacy, help-seeking, and instructor-student trust than courses using transparent, formative, AI-guidance policies, then the paper's central claim that assessment-oriented policies leave learning harms unaddressed and strain relationships would be wrong.","tokens_in":24364,"feed_emoji":"🎓","tokens_out":5813,"duration_ms":53636,"temperature":0.7,"pith_summary":"The paper claims that undergraduate computer science instructors have responded to generative AI mainly by trying to make assessments AI-proof—through paper exams, exam-heavy grading, and detection—rather than by changing how students learn. Drawing on 13 semi-structured interviews with US instructors, it distinguishes two sets of problems: learning harms (students offloading thinking, overestimating their skills, becoming socially isolated) and assessment harms (instructors no longer able to tell who has actually mastered the material). The paper argues that assessment-focused policies impose second-order costs: policing burden on instructors, stress and shame for students, and worsening instructor-student relationships. It recommends learning-oriented alternatives—transparent discussion of AI, modeled good use, and low-stakes formative assessments—that could guide students without curriculum overhaul.","feed_headline":"13 instructors reveal AI policies prioritize exams over learning","feed_subtitle":"CS instructor interviews: AI-proofing exams strains trust while learning harms go unaddressed.","key_machinery":"The paper's analytic engine is the distinction between learning harms and assessment harms, with AI policies treated as reified social contracts between instructors and students. Learning harms are the ways AI undermines student development: offloading cognitive work, creating illusions of competence, and eroding social learning. Assessment harms are the ways AI undermines the instructor's ability to tell whether submitted work reflects learning. The argument is carried by showing that policies responding to assessment harms—paper exams, exam-heavy grading, AI-use detection—leave learning harms in place while shifting responsibility onto students, whereas policies responding to learning harm","core_discovery":"The central finding is that AI policies in undergraduate CS are primarily reactions to assessment harms, not learning harms. Instructors report that AI tools induce cognitive offloading, illusions of competence, and reduced human help-seeking, but the policies they most readily adopt—proctored paper exams, heavier exam weighting, and detection of AI signatures in code—target the instructor's ability to trust submitted work. These policies leave the learning harms in place and create new second-order burdens: labor-intensive detection that most instructors cannot sustain, punitive grading that students experience as adversarial, and a shift of responsibility onto students to self-regulate in","pith_inferences":["If the self-report pattern generalizes, the institutional debate over 'cheating' with AI may be misframed: the harder problem is not detecting misuse but redesigning assessment so that earning a grade requires the learning the course wants to produce.","A natural next study would pair course AI policies with student learning analytics, assignment-replay logs, and relationship-quality surveys to test whether assessment-focused policies are as ineffective as these instructors report.","The paper's analogy to abstinence-only education points toward a testable extension: permissive, guidance-based AI policies that openly scaffold use may yield better long-term coding skill than restrictive policies once AI tools are part of everyday work.","Because instructors report nominally enforcing policies they cannot actually police, written AI policies may increasingly become symbolic documents, with actual norms negotiated informally in each classroom."],"forward_implications":["Assessment-oriented policies such as proctored paper exams and exam-heavy grading preserve short-term assessment integrity but do not fix the AI-driven learning harms instructors themselves describe.","Detection-based enforcement is so labor-intensive and unreliable that many instructors either stop enforcing, lower expectations, or rely on students to self-regulate, leaving vulnerable students without support.","The erosion of trust and rise of AI stigma reduce students' help-seeking from humans, compounding social isolation and making it harder for instructors to intervene.","Learning-oriented policies—clear AI-use guidance, instructor modeling, and low-stakes formative assessments—are reported by some instructors to be feasible without major curriculum change and deserving of broader adoption."],"fun_headline_variants":["AI policies guard exams, not student learning","CS instructors: AI policies sacrifice learning for exam integrity","AI-proofing exams strains trust, leaves learning harms","Policing AI in CS: policies protect tests, not learners","Instructor interviews: AI policies favor assessment, not learning"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire pattern rests on 13 instructors' self-reports accurately describing their own policies and what students do; the paper explicitly says in its limitations that no independent source of empirical data validates these accounts.","fun_headline_variants_meta":{"raw":{"variants":["AI policies guard exams, not student learning","CS instructors: AI policies sacrifice learning for exam integrity","AI-proofing exams strains trust, leaves learning harms","Policing AI in CS: policies protect tests, not learners","Instructor interviews: AI policies favor assessment, not learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000521,"raw_usage":{"total_tokens":2309,"prompt_tokens":646,"completion_tokens":1663,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":390,"completion_tokens_details":{"reasoning_tokens":1599}},"tokens_in":390,"tokens_out":1663,"duration_ms":10224,"temperature":1.0,"reasoning_tokens":1599,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T20:50:48.263942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A multi-section comparison would settle it: if courses with strict AI-proofing (paper exams, exam-weighted grades) produced equal or better measures of long-term coding skill, self-efficacy, help-seeking, and instructor-student trust than courses using transparent, formative, AI-guidance policies, then the paper's central claim that assessment-oriented policies leave learning harms unaddressed and strain relationships would be wrong.","supporting_citations":[],"review_version":1}