{"id":"81715ae2-18e4-466f-9ff8-807ca3fd2793","arxiv_id":"2412.13412","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a controlled experiment, AI plagiarism fell sharply on a higher-order 'create' task compared to recall tasks, especially among ChatGPT users.","lead":"This study ran a controlled experiment with 123 students who completed three tasks of increasing difficulty using no tools, an e-textbook, Google, or ChatGPT. AI plagiarism dropped sharply on the hardest, higher-order task, especially for ChatGPT users, suggesting assessment design can deter AI-assisted cheating.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Causal claim is not yet identified: the monotonic AI-plagiarism drop appears only in the ChatGPT group, and Task 3's local-knowledge demand plus fixed task order confound Bloom's complexity with task content.","rationale":"The reader's weakest assumption correctly identifies the Task 3 confound, and I agree that this is a serious threat to causal attribution. My stress-test adds two closely connected points that strengthen the concern: first, the fixed task order entangles complexity with practice, fatigue, and time-on-task; second, the claimed monotonic decrease is not actually present in the e-textbook and Google groups, only in the ChatGPT group. The central claim therefore goes beyond what the data support, even before the confounding is addressed. The right disposition is the same conditional acceptance the reader gave: the paper is worth publishing as a preliminary, well-designed exploratory study, but the headline claim should be restricted to ChatGPT users and explicitly caveated with the confounds. I do not see grounds for rejection because the basic direction is plausible, the randomization and control group are strengths, and the Task 1 vs. Task 3 difference in the ChatGPT group is large and worth follow-up. A counterbalanced, content-matched replication or even a focused re-analysis of within-group contrasts would materially settle whether the recommendation is sound.","tokens_in":10430,"tokens_out":3877,"duration_ms":39536,"concrete_test":"Run a counterbalanced replication in which task order is varied across participants and the three tasks are matched on topic, context, and local-knowledge demands, changing only the Bloom verb (e.g., define vs. apply vs. evaluate/create on the same scenario). Also report pairwise within-group contrasts for the e-textbook and Google groups on Tasks 2 vs. 3 in the current data; if the Task 3 AI-plagiarism drop disappears when order and content are controlled, or if these contrasts are non-significant, the effect cannot be attributed to cognitive complexity alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('AI plagiarism decreases as task complexity increases') requires that the three tasks differ only in Bloom's cognitive complexity, not in content, access to local knowledge, or order. Section 4.2 reports that two experts agreed on task taxonomy levels, but no inter-rater reliability statistic is provided, and the design is not counterbalanced: every participant completed Task 1, then Task 2, then Task 3. Task 3 asks students to propose a solution to a problem they identified, which plausibly requires knowledge of their own laboratory environment or personal experience that ChatGPT-3.5 cannot access. The observed drop in AI plagiarism could therefore be caused by task specificity or unavailability of contextual information, rather than by the 'create' level. The authors acknowledge LLM context limitations in Section 6.3 but do not treat this as a confound. Moreover, the claimed monotonic decrease is not present in the e-textbook group (6.68, 9.24, 4.68) or the Google group (11.34, 21.59, 2.28); only the ChatGPT group (71.89, 64.92, 26.74) decreases monotonically. No pairwise within-group contrasts are reported, so the significant interaction does not by itself establish the monotonic pattern that the abstract and conclusion assert. This is load-bearing because the recommendation to design higher-order assessments as a general deterrent depends on complexity, not content or task order, being the active ingredient.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a classroom experiment with 123 students randomly assigned to a control, e-textbook, Google, or ChatGPT condition. Participants completed three data-privacy tasks intended to correspond to lower-, medium-, and higher-order Bloom's taxonomy levels, always in the same order. Turnitin similarity scores and Turnitin AI-writing-detection percentages were analyzed with repeated-measures ANOVA. The paper's central claim is that AI plagiarism decreases as task complexity increases, and that higher-order assessments are therefore an effective strategy for minimizing AI-assisted plagiarism.","tokens_in":10674,"tokens_out":4869,"duration_ms":44792,"significance":"If the central claim were fully supported, the paper would provide useful empirical evidence for a widely discussed but under-tested pedagogical recommendation: that assessments requiring higher-order thinking can deter AI-assisted plagiarism. The study has notable strengths: random assignment with cluster-based balancing, a pretest MANOVA showing group equivalence, a repeated-measures design, and explicit engagement with the false-positive limitation of AI detectors. The distinction between Turnitin similarity scores and AI-writing-detection scores is practically valuable. However, the causal claim is currently stronger than the design and analysis support, and several load-bearing points need to be addressed before the conclusions can be accepted as stated.","major_comments":[{"comment":"The causal attribution that task complexity drives the change in AI plagiarism requires that the three tasks differ only in Bloom's cognitive level. The design does not ensure this: every participant completed Task 1, then Task 2, then Task 3, so task order and task content are fully confounded with complexity. In addition, Task 3 requires students to propose a solution to a problem they themselves identified, which plausibly requires local, personal, or laboratory-specific knowledge that ChatGPT-3.5 cannot access; Section 6.3 acknowledges that LLMs struggle with contextual understanding, but it does not treat this as a confound. The manuscript should either provide counterbalanced or task-matched data, or substantially soften the causal language. The expert validation in Section 4.2 should also be strengthened with an inter-rater reliability statistic, since the current statement of two experts agreeing does not quantify agreement.","section":"Section 4.2 / Figure 1"},{"comment":"The abstract and conclusion claim that AI plagiarism decreases as task complexity increases, and Section 3 hypothesizes this 'regardless of the technology available to students.' The data in Table 3 do not support a monotonic decrease for the e-textbook group (6.68, 9.24, 4.68) or the Google group (11.34, 21.59, 2.28); both increase from Task 1 to Task 2. The total Task 1 and Task 2 means are nearly identical (27.42 and 27.02), and the overall decline is driven almost entirely by Task 3 and by the ChatGPT group (71.89, 64.92, 26.74). No within-group pairwise contrasts are reported, so the significant tasks-by-group interaction does not by itself establish the monotonic pattern claimed. The authors should report simple effects and pairwise comparisons for each group and revise the central claim to match the actual pattern.","section":"Table 3 / Section 3"},{"comment":"The manuscript acknowledges that Turnitin's AI-writing-detection scores have false positives and recommends accounting for up to 20% false positives. Yet several of the substantive between-group and between-task differences used to support the conclusions fall within that range, for example the e-textbook Task 1 mean of 6.68 and Task 2 mean of 9.24, and the Google Task 2 mean of 21.59. The argument that a repeated-measures design mitigates false positives assumes that false-positive rates are stable across tasks and groups, which is not demonstrated, especially since the control group shows nonzero AI-plagiarism scores in Task 1 (4.68) despite having no access to tools. A sensitivity analysis restricted to groups and comparisons plausibly above the false-positive threshold, or an explicit modeling of detection noise, is needed before interpreting these values as genuine AI plagiarism.","section":"Sections 4.5 and 6.1"},{"comment":"There is an internal inconsistency in the reported sample size for the ChatGPT group. Table 1 lists 38 students in the ChatGPT group, while Tables 2 and 3 report n = 37 for the ChatGPT group. The manuscript should clarify whether one participant was excluded post hoc, and if so, document the reason and report the degrees of freedom of all analyses accordingly.","section":"Section 5"}],"minor_comments":[{"comment":"The sphericity notation appears incomplete: 'Mauchly's W =.973, 2(2) = 3.178' should include the chi-square symbol and value, for example χ²(2) = 3.178.","section":"Section 4.5"},{"comment":"The statement that 'the data obtained in each condition was found to be normally distributed' should be accompanied by the specific test statistics or a citation, since normality of bounded percentages is not self-evident.","section":"Section 4.5"},{"comment":"The text refers to 'Table 7' and 'Figure 5' when discussing AI plagiarism percentages, but the actual table is Table 3 and the figure is Figure 4. The cross-references should be corrected.","section":"Section 5"},{"comment":"The sentence 'The interaction between groups and repeated measures shows that the Control and ChatGPT groups had a consistent performance with lower variability' is puzzling, because the ChatGPT group has the largest standard deviations in Table 3 (32.89, 37.69, 36.74). Please clarify what 'consistent' is intended to mean.","section":"Section 5 / Figure 4"},{"comment":"The citation in the phrase 'with 'remember' being the least complex and 'create' being the most complex' appears to reference [18] only at the start of the paragraph; a direct citation at this sentence would help the reader locate the source of the complexity ordering.","section":"Section 2.1"},{"comment":"The phrase 'Participants in each completed the tasks in the following order' is grammatically incomplete; it should read 'Participants in each group completed the tasks in the following order.'","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful empirical core, but the central causal claim is currently overreach relative to the design: the fixed task order, task-specific content, and the non-monotonic pattern in two of the four groups are load-bearing issues. I would encourage the editor to require either a substantially reframed central claim or new data/analyses that address the confounding; a simple rewrite of the abstract without addressing the confounds would not be sufficient. There is no indication of any research-integrity problem in the manuscript itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is a real experiment with a plausible design, but the abstract's monotonic claim is not supported by the group means. The drop is driven almost entirely by the ChatGPT group (71.89, 64.92, 26.74). The e-textbook and Google groups actually increase from Task 1 to Task 2 before falling, so the interaction is not the same as a general monotonic decrease. Be careful about citing this as evidence that higher-order tasks broadly deter AI plagiarism.\n\nWhat is genuinely new: prior work argued the strategy theoretically, but this is a controlled test with random assignment, multiple tool groups, and three tasks mapped to Bloom's levels. The stratified assignment and MANOVA checks are solid. The paper also clearly separates similarity scores from AI plagiarism scores and openly discusses Turnitin's false-positive issue, going so far as to suggest a 0–20% false-positive tolerance. That is honest and practically useful.\n\nThe soft spots are real. Task 3 asks students to propose a solution to a problem they identified, which plausibly requires knowledge of their own lab environment or personal experience that ChatGPT-3.5 cannot access. So the drop may be due to task specificity rather than Bloom's 'create' level. The design is not counterbalanced: everyone does Task 1, then Task 2, then Task 3. The two expert raters confirmed the Bloom levels, but no inter-rater reliability statistic is reported. And no pairwise within-group contrasts are given, so the significant interaction does not by itself establish the monotonic trend the conclusion asserts. The paper acknowledges LLM context limitations in Section 6.3 but does not treat them as a confound for this specific design.\n\nStill, the design is better than most studies in this space, and the authors are transparent about limitations. This deserves a serious referee, not a desk rejection. The authors should be pushed to either soften the causal claim or strengthen the design (counterbalanced, content-matched tasks, reported rater agreement).\n\nFor your own use: worth reading for the ChatGPT-group pattern and the false-positive discussion, but not as a standalone proof that higher-order assessments solve AI plagiarism.","headline":"Useful experiment, but the headline claim that AI plagiarism falls monotonically with task complexity only holds in the ChatGPT group, and task specificity/order confound the causal story.","tokens_in":11192,"tokens_out":1357,"would_cite":true,"duration_ms":14250,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that AI plagiarism falls as task complexity increases, and that assessments designed around higher-order thinking are a viable way to reduce AI-driven plagiarism.","keywords":["AI plagiarism","Bloom's taxonomy","ChatGPT","generative AI","task complexity","higher-order thinking","academic integrity","Turnitin similarity score"],"falsifier":"Re-run the same three tasks with ChatGPT-4o and add a fourth task at the create level that relies only on general knowledge; if AI-plagiarism percentages on the general-knowledge create task stay as high as on the recall task, then the observed decline is driven by ChatGPT-3.5's lack of access to course-specific context, not by cognitive complexity per se.","tokens_in":10211,"feed_emoji":"🤖","tokens_out":6698,"duration_ms":57046,"temperature":0.7,"pith_summary":"The paper tests a practical question: can the design of an assignment make students less likely to turn to generative AI for their work? In a controlled lab study, 123 undergraduates completed three tasks of rising complexity—recalling and understanding, applying, and finally analyzing, evaluating, and creating—using either no tools, an e-textbook, Google, or ChatGPT. Turnitin's AI-plagiarism percentages dropped as the tasks became harder, most sharply in the ChatGPT group, from about 72% on the simplest task to about 27% on the hardest. The authors read this as evidence that higher-order assessments are a viable strategy for reducing AI plagiarism, and they argue that similarity scores and AI-plagiarism scores measure different things and should both be reported when checking student work.","feed_headline":"Higher-order tasks cut AI plagiarism by two-thirds","feed_subtitle":"In a 123-student experiment, ChatGPT users' AI-detected text fell from 72% on recall to 27% on create-level work.","key_machinery":"The machinery is a task-complexity gradient built on Bloom's revised taxonomy, a six-level hierarchy of cognitive processes running from remembering up to creating. The authors designed three tasks (remember/understand, apply, analyze/evaluate/create), had two subject experts independently confirm the levels, and used a repeated-measures design in which every participant completed all three tasks in one of four tool conditions. The outcomes are Turnitin's two scores for the same submission: the similarity score, which measures text matching existing sources, and the AI plagiarism score, which measures text the detector judges to be AI-generated. The within-subjects comparison across the three complexity levels is what lets the authors attribute changes in those scores to task complexity rather than to differences between students.","core_discovery":"The paper's central claim is that AI plagiarism decreases as task complexity increases, so that assessments aimed at Bloom's higher-order levels—analyzing, evaluating, and creating—are a workable way to curb AI-driven plagiarism. The evidence is a within-subjects experiment: the same 123 students produced text for three tasks of increasing complexity, and Turnitin's AI-writing score fell from a mean of 27.42% across all groups on Task 1 to 9.75% on Task 3. The ChatGPT group showed the largest decline, from 71.89% to 26.74%, while the control group stayed near zero. The authors also report that similarity scores and AI-plagiarism scores are distinct: treatment groups looked similar on traditional similarity but differed sharply on AI detection, so they recommend using both metrics and human review, allowing for up to 20% false positives.","pith_inferences":["The paper leaves implicit that its create-level Task 3 may draw on course-specific or lab-local knowledge that ChatGPT-3.5 cannot access, so the drop in AI plagiarism could reflect task specificity rather than cognitive complexity alone.","Because the study used ChatGPT-3.5, the results should be read as a lower bound for current models; newer models that reason better at higher levels could shrink the gap, a possibility the paper's data do not test.","A testable extension is to hold the Bloom level fixed while varying whether the task needs local knowledge, which would separate cognitive complexity from access-to-context effects.","The paper's suggested 20% false-positive allowance is detector- and context-specific; an extension would calibrate that threshold per institution by running control submissions known to be human-written."],"forward_implications":["If the central claim holds, redesigning assessments around analysis, evaluation, and creation should reduce AI-generated text in student submissions even when students have ChatGPT or similar tools open.","Institutions that use only traditional similarity checks will miss AI-generated work; the paper implies both similarity and AI-plagiarism scores should be reported together.","AI-plagiarism scores will carry false positives in the 0–20% range even when no generative AI was used, so automatic penalties should be set above that band or paired with human review.","The assessment-led approach to academic integrity, already recommended in the pre-AI cheating literature, remains relevant for generative AI rather than being obsolete."],"supporting_citations":[{"why":"Supplies Bloom's revised taxonomy that defines the complexity ordering of the three tasks.","marker":"[18]"},{"why":"Gives the premise that lower-order recall-oriented tests make cheating easier to commit.","marker":"[24]"},{"why":"Supports the claim that short, disconnected tests increase cheating while connected tasks reduce it.","marker":"[25]"},{"why":"Recommends case studies and open-ended tasks at higher Bloom levels to deter cheating.","marker":"[26]"},{"why":"Provides empirical evidence that critical-thinking assignments produce less plagiarism than opinion or documentation tasks.","marker":"[27]"},{"why":"Frames AI plagiarism as a distinct concern and motivates using AI detection alongside similarity checks.","marker":"[37]"},{"why":"Defines the Turnitin similarity score used as the first outcome measure.","marker":"[38]"},{"why":"Documents high false-positive rates in AI detectors, motivating the paper's caution and the 20% allowance.","marker":"[39]"},{"why":"Supplies the repeated-measures ANOVA procedure used for the within-subjects comparisons.","marker":"[40]"}],"fun_headline_variants":["Complex tasks cut AI plagiarism by two-thirds","AI plagiarism drops as task difficulty increases","Study: Harder assignments reduce AI plagiarism","Task complexity stifles AI-generated cheating","Higher-order thinking lowers AI plagiarism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three tasks differ only in their Bloom's cognitive complexity; in reality the create-level Task 3 may also require local or course-specific knowledge unavailable to ChatGPT-3.5, and its classification as high order rests on two expert reviewers with no reported inter-rater reliability statistic, so the drop in AI plagiarism could come from task content rather than complexity.","fun_headline_variants_meta":{"raw":{"variants":["Complex tasks cut AI plagiarism by two-thirds","AI plagiarism drops as task difficulty increases","Study: Harder assignments reduce AI plagiarism","Task complexity stifles AI-generated cheating","Higher-order thinking lowers AI plagiarism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00125,"raw_usage":{"total_tokens":5056,"prompt_tokens":806,"completion_tokens":4250,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":4188}},"tokens_in":422,"tokens_out":4250,"duration_ms":26505,"temperature":1.0,"reasoning_tokens":4188,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:09:05.300359+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same three tasks with ChatGPT-4o and add a fourth task at the create level that relies only on general knowledge; if AI-plagiarism percentages on the general-knowledge create task stay as high as on the recall task, then the observed decline is driven by ChatGPT-3.5's lack of access to course-specific context, not by cognitive complexity per se.","supporting_citations":[{"cited_title":"A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition","cited_arxiv_id":null,"evidence_quote":"Supplies Bloom's revised taxonomy that defines the complexity ordering of the three tasks."},{"cited_title":"Supporting academic honesty in online courses","cited_arxiv_id":null,"evidence_quote":"Gives the premise that lower-order recall-oriented tests make cheating easier to commit."},{"cited_title":"Curbing academic dishonesty in online courses","cited_arxiv_id":null,"evidence_quote":"Supports the claim that short, disconnected tests increase cheating while connected tasks reduce it."},{"cited_title":"Towards academic integrity: Using bloom’s taxonomy and technology to deter cheating in online courses","cited_arxiv_id":null,"evidence_quote":"Recommends case studies and open-ended tasks at higher Bloom levels to deter cheating."},{"cited_title":"Using writing assignment designs to mitigate plagiarism","cited_arxiv_id":null,"evidence_quote":"Provides empirical evidence that critical-thinking assignments produce less plagiarism than opinion or documentation tasks."},{"cited_title":"Understanding the Turnitin similarity report, 2021","cited_arxiv_id":null,"evidence_quote":"Defines the Turnitin similarity score used as the first outcome measure."},{"cited_title":"Evaluating the efficacy of ai content detection tools in differentiating between human and ai-generated text","cited_arxiv_id":null,"evidence_quote":"Documents high false-positive rates in AI detectors, motivating the paper's caution and the 20% allowance."}],"review_version":1}