Pith. sign in

REVIEW 4 major objections 6 minor 57 references

The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that AI plagiarism falls as task complexity increases, and that assessments designed around higher-order thinking are a viable way to reduce AI-driven plagiarism.

desk verdict Useful experiment, but the headline claim that AI plagiarism falls monotonically with task complexity only holds in the ChatGPT group, and task specificity/order confound the causal story. read the letter →

arxiv 2412.13412 v1 pith:CMKLUMP2 submitted 2024-12-18 cs.HC

classification cs.HC
keywords AIplagiarismBloom'staxonomyChatGPTgenerativetaskcomplexityhigher-orderthinkingacademicintegrityTurnitinsimilarityscore
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests a practical question: can the design of an assignment make students less likely to turn to generative AI for their work? In a controlled lab study, 123 undergraduates completed three tasks of rising complexity—recalling and understanding, applying, and finally analyzing, evaluating, and creating—using either no tools, an e-textbook, Google, or ChatGPT. Turnitin's AI-plagiarism percentages dropped as the tasks became harder, most sharply in the ChatGPT group, from about 72% on the simplest task to about 27% on the hardest. The authors read this as evidence that higher-order assessments are a viable strategy for reducing AI plagiarism, and they argue that similarity scores and AI-plagiarism scores measure different things and should both be reported when checking student work.

What carries the argument

The machinery is a task-complexity gradient built on Bloom's revised taxonomy, a six-level hierarchy of cognitive processes running from remembering up to creating. The authors designed three tasks (remember/understand, apply, analyze/evaluate/create), had two subject experts independently confirm the levels, and used a repeated-measures design in which every participant completed all three tasks in one of four tool conditions. The outcomes are Turnitin's two scores for the same submission: the similarity score, which measures text matching existing sources, and the AI plagiarism score, which measures text the detector judges to be AI-generated. The within-subjects comparison across the three complexity levels is what lets the authors attribute changes in those scores to task complexity rather than to differences between students.

What would settle it

Re-run the same three tasks with ChatGPT-4o and add a fourth task at the create level that relies only on general knowledge; if AI-plagiarism percentages on the general-knowledge create task stay as high as on the recall task, then the observed decline is driven by ChatGPT-3.5's lack of access to course-specific context, not by cognitive complexity per se.

Watch

Extended reading notes

Core claim

The paper's central claim is that AI plagiarism decreases as task complexity increases, so that assessments aimed at Bloom's higher-order levels—analyzing, evaluating, and creating—are a workable way to curb AI-driven plagiarism. The evidence is a within-subjects experiment: the same 123 students produced text for three tasks of increasing complexity, and Turnitin's AI-writing score fell from a mean of 27.42% across all groups on Task 1 to 9.75% on Task 3. The ChatGPT group showed the largest decline, from 71.89% to 26.74%, while the control group stayed near zero. The authors also report that similarity scores and AI-plagiarism scores are distinct: treatment groups looked similar on traditional similarity but differed sharply on AI detection, so they recommend using both metrics and human review, allowing for up to 20% false positives.

Load-bearing premise

The load-bearing premise is that the three tasks differ only in their Bloom's cognitive complexity; in reality the create-level Task 3 may also require local or course-specific knowledge unavailable to ChatGPT-3.5, and its classification as high order rests on two expert reviewers with no reported inter-rater reliability statistic, so the drop in AI plagiarism could come from task content rather than complexity.

Editorial extensions

If this is right

  • If the central claim holds, redesigning assessments around analysis, evaluation, and creation should reduce AI-generated text in student submissions even when students have ChatGPT or similar tools open.
  • Institutions that use only traditional similarity checks will miss AI-generated work; the paper implies both similarity and AI-plagiarism scores should be reported together.
  • AI-plagiarism scores will carry false positives in the 0–20% range even when no generative AI was used, so automatic penalties should be set above that band or paired with human review.
  • The assessment-led approach to academic integrity, already recommended in the pre-AI cheating literature, remains relevant for generative AI rather than being obsolete.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that its create-level Task 3 may draw on course-specific or lab-local knowledge that ChatGPT-3.5 cannot access, so the drop in AI plagiarism could reflect task specificity rather than cognitive complexity alone.
  • Because the study used ChatGPT-3.5, the results should be read as a lower bound for current models; newer models that reason better at higher levels could shrink the gap, a possibility the paper's data do not test.
  • A testable extension is to hold the Bloom level fixed while varying whether the task needs local knowledge, which would separate cognitive complexity from access-to-context effects.
  • The paper's suggested 20% false-positive allowance is detector- and context-specific; an extension would calibrate that threshold per institution by running control submissions known to be human-written.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a classroom experiment with 123 students randomly assigned to a control, e-textbook, Google, or ChatGPT condition. Participants completed three data-privacy tasks intended to correspond to lower-, medium-, and higher-order Bloom's taxonomy levels, always in the same order. Turnitin similarity scores and Turnitin AI-writing-detection percentages were analyzed with repeated-measures ANOVA. The paper's central claim is that AI plagiarism decreases as task complexity increases, and that higher-order assessments are therefore an effective strategy for minimizing AI-assisted plagiarism.

Significance. If the central claim were fully supported, the paper would provide useful empirical evidence for a widely discussed but under-tested pedagogical recommendation: that assessments requiring higher-order thinking can deter AI-assisted plagiarism. The study has notable strengths: random assignment with cluster-based balancing, a pretest MANOVA showing group equivalence, a repeated-measures design, and explicit engagement with the false-positive limitation of AI detectors. The distinction between Turnitin similarity scores and AI-writing-detection scores is practically valuable. However, the causal claim is currently stronger than the design and analysis support, and several load-bearing points need to be addressed before the conclusions can be accepted as stated.

major comments (4)
  1. [Section 4.2 / Figure 1] The causal attribution that task complexity drives the change in AI plagiarism requires that the three tasks differ only in Bloom's cognitive level. The design does not ensure this: every participant completed Task 1, then Task 2, then Task 3, so task order and task content are fully confounded with complexity. In addition, Task 3 requires students to propose a solution to a problem they themselves identified, which plausibly requires local, personal, or laboratory-specific knowledge that ChatGPT-3.5 cannot access; Section 6.3 acknowledges that LLMs struggle with contextual understanding, but it does not treat this as a confound. The manuscript should either provide counterbalanced or task-matched data, or substantially soften the causal language. The expert validation in Section 4.2 should also be strengthened with an inter-rater reliability statistic, since the current statement of two experts agreeing does not quantify agreement.
  2. [Table 3 / Section 3] The abstract and conclusion claim that AI plagiarism decreases as task complexity increases, and Section 3 hypothesizes this 'regardless of the technology available to students.' The data in Table 3 do not support a monotonic decrease for the e-textbook group (6.68, 9.24, 4.68) or the Google group (11.34, 21.59, 2.28); both increase from Task 1 to Task 2. The total Task 1 and Task 2 means are nearly identical (27.42 and 27.02), and the overall decline is driven almost entirely by Task 3 and by the ChatGPT group (71.89, 64.92, 26.74). No within-group pairwise contrasts are reported, so the significant tasks-by-group interaction does not by itself establish the monotonic pattern claimed. The authors should report simple effects and pairwise comparisons for each group and revise the central claim to match the actual pattern.
  3. [Sections 4.5 and 6.1] The manuscript acknowledges that Turnitin's AI-writing-detection scores have false positives and recommends accounting for up to 20% false positives. Yet several of the substantive between-group and between-task differences used to support the conclusions fall within that range, for example the e-textbook Task 1 mean of 6.68 and Task 2 mean of 9.24, and the Google Task 2 mean of 21.59. The argument that a repeated-measures design mitigates false positives assumes that false-positive rates are stable across tasks and groups, which is not demonstrated, especially since the control group shows nonzero AI-plagiarism scores in Task 1 (4.68) despite having no access to tools. A sensitivity analysis restricted to groups and comparisons plausibly above the false-positive threshold, or an explicit modeling of detection noise, is needed before interpreting these values as genuine AI plagiarism.
  4. [Section 5] There is an internal inconsistency in the reported sample size for the ChatGPT group. Table 1 lists 38 students in the ChatGPT group, while Tables 2 and 3 report n = 37 for the ChatGPT group. The manuscript should clarify whether one participant was excluded post hoc, and if so, document the reason and report the degrees of freedom of all analyses accordingly.
minor comments (6)
  1. [Section 4.5] The sphericity notation appears incomplete: 'Mauchly's W =.973, 2(2) = 3.178' should include the chi-square symbol and value, for example χ²(2) = 3.178.
  2. [Section 4.5] The statement that 'the data obtained in each condition was found to be normally distributed' should be accompanied by the specific test statistics or a citation, since normality of bounded percentages is not self-evident.
  3. [Section 5] The text refers to 'Table 7' and 'Figure 5' when discussing AI plagiarism percentages, but the actual table is Table 3 and the figure is Figure 4. The cross-references should be corrected.
  4. [Section 5 / Figure 4] The sentence 'The interaction between groups and repeated measures shows that the Control and ChatGPT groups had a consistent performance with lower variability' is puzzling, because the ChatGPT group has the largest standard deviations in Table 3 (32.89, 37.69, 36.74). Please clarify what 'consistent' is intended to mean.
  5. [Section 2.1] The citation in the phrase 'with 'remember' being the least complex and 'create' being the most complex' appears to reference [18] only at the start of the paragraph; a direct citation at this sentence would help the reader locate the source of the complexity ordering.
  6. [Section 4.2] The phrase 'Participants in each completed the tasks in the following order' is grammatically incomplete; it should read 'Participants in each group completed the tasks in the following order.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected: the paper's central claim is an empirical, measured result, not a self-referential derivation.

full rationale

The paper is an empirical experiment rather than a formal derivation, and no load-bearing step reduces to its own inputs. The independent variable (task complexity) is operationalized through Bloom's revised taxonomy with expert validation: Section 4.2 states that 'two independent subject experts confirmed the tasks' differentiation according to Bloom's taxonomy, agreeing on the levels for each task,' which is independent of the outcome measures (Turnitin similarity scores and AI plagiarism percentages). The dependent variables come from an external detection tool, and no parameter is fitted to a subset of the data and then presented as a prediction. There are no self-citations: neither author appears among the 57 references, and the load-bearing background claims (Bloom's taxonomy structure [18], AI detector unreliability [39, 45], and the assessment-deters-plagiarism literature [24-28]) all rest on external sources. The skeptical concerns about the causal claim — fixed task order (all participants completed Task 1, then Task 2, then Task 3), Task 3's possible demand for local or contextual knowledge, and the non-monotonic patterns in the e-textbook and Google groups — are threats to internal and causal validity, not circularity; the central claim was empirically falsifiable and, in fact, was not reproduced in the other treatment groups. Section 7.1 candidly acknowledges limits of sample, content area, and generalizability. The reader's-take score of 0.0 is appropriate: this is a self-contained empirical study with no circular derivation chain.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities; the paper is an empirical study using an existing plagiarism detector. The central inference rests on assumptions about the validity of Turnitin's AI score and the equivalence of the three tasks except for cognitive complexity.

assumptions (4)
  • domain assumption Turnitin's AI writing score provides a valid relative measure of AI-generated content in student submissions.
    Invoked throughout Section 4.5 and the Results; the paper acknowledges false positives but relies on repeated-measures comparisons to cancel them. If false-positive rates vary by task type, the observed declines could be artifacts.
  • domain assumption The three tasks differ only in cognitive complexity (Bloom levels) and not in other factors such as answerability by ChatGPT or local knowledge.
    Task design in Section 4.2 is validated only by two expert reviewers; no inter-rater reliability or pilot data is reported. This is the causal crux of the study.
  • domain assumption Students followed the intended tool-use protocol under continuous monitoring.
    Section 4.2 states continuous monitoring and that five deviating students were excluded. The paper assumes the remaining students used only the assigned tools, which cannot be verified from the manuscript.
  • standard math Standard statistical assumptions for repeated measures ANOVA hold.
    The paper reports Mauchly's test and states normality was checked; no additional justification is needed for a standard statistical procedure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools." pith.science (2026). https://pith.science/paper/CMKLUMP2

@misc{pith2026241213412,
  author       = {Pith},
  title        = {Pith review of: The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CMKLUMP2}},
  note         = {Machine review of arXiv:2412.13412}
}
read the original abstract

This study investigates whether assessments fostering higher-order thinking skills can reduce plagiarism involving generative AI tools. Participants completed three tasks of varying complexity in four groups: control, e-textbook, Google, and ChatGPT. Findings show that AI plagiarism decreases as task complexity increases, with higher-order tasks resulting in lower similarity scores and AI plagiarism percentages. The study also highlights the distinction between similarity scores and AI plagiarism, recommending both for effective plagiarism detection. Results suggest that assessments promoting higher-order thinking are a viable strategy for minimizing AI-driven plagiarism.

Figures

Figures reproduced from arXiv: 2412.13412 by the authors.

Figure 1
Figure 1. Tasks used in the study evenly distributed across the four experimental groups, maintaining diversity in experience and knowledge. We also ensured gender balance in this process [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Group formation [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Similarity scores were smaller in Task 3, the ChatGPT group consistently had the highest AI plagiarism percentages, while the control group had the lowest. As task complexity increased, AI plagiarism decreased in the ChatGPT group, but AI plagiarism remained in the range of 10–20% in the e-textbook and Google groups, likely due to false positives. The findings suggest a significant decrease in AI plagiarism with inc… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: AI Plagiarism scores 6 Discussion With advances in AI, humans began to use AI tools to produce creative outputs in different aspects of their lives. Although generative AI technologies have experienced quick and widespread adoption, their place in education is still up…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 50 canonical work pages

  1. [1]

    Chinese university students’ perceptions of plagiarism

    Guangwei Hu and Jun Lei. Chinese university students’ perceptions of plagiarism. Ethics & Behavior, 25(3):233– 255, 2015

  2. [2]

    An empirical analysis of differences in plagiarism among world cultures

    David C Ison. An empirical analysis of differences in plagiarism among world cultures. Journal of Higher Education Policy and Management, 40(4):291–304, 2018

  3. [3]

    Factors influencing plagiarism in higher education: A comparison of german and slovene students

    Eva Jereb, Matjaž Perc, Barbara Lämmlein, Janja Jerebic, Marko Urh, Iztok Podbregar, and Polona Šprajc. Factors influencing plagiarism in higher education: A comparison of german and slovene students. PloS one, 13(8):e0202252, 2018

  4. [4]

    Plagiarism and its treatment in higher education

    Peter J Larkham and Susan Manns. Plagiarism and its treatment in higher education. Journal of Further and Higher Education, 26(4):339–349, 2002

  5. [5]

    Factors associated with student plagiarism in a post-1992 university

    Roger Bennett*. Factors associated with student plagiarism in a post-1992 university. Assessment & Evaluation in Higher Education, 30(2):137–162, 2005

  6. [6]

    How common is commercial contract cheating in higher education and is it increasing? a systematic review

    Philip M Newton. How common is commercial contract cheating in higher education and is it increasing? a systematic review. In Frontiers in Education, volume 3, page 67. Frontiers Media SA, 2018

  7. [7]

    How prevalent is plagiarism among college students? anonymity preserving evidence from austrian undergraduates

    Christian Hopp and Alexander Speil. How prevalent is plagiarism among college students? anonymity preserving evidence from austrian undergraduates. Accountability in Research, 28(3):133–148, 2021

  8. [8]

    Is plagiarism really on the rise? results from four 5-yearly surveys

    Guy J Curtis and Kell Tremayne. Is plagiarism really on the rise? results from four 5-yearly surveys. Studies in Higher Education, 46(9):1816–1826, 2021

Show all 57 references
  1. [9]

    Anupama Prashar, Priyanka Gupta, and Yogesh K. Dwivedi. Plagiarism awareness efforts, students’ ethical judgment and behaviors: a longitudinal experiment study on ethical nuances of plagiarism in higher education. Studies in Higher Education, 49(6):929–955, 2023

  2. [10]

    Does culture influence understanding and perceived seriousness of plagiarism? International Journal for Educational Integrity, 4(2), 2008

    Amanda Maxwell, Guy J Curtis, and Lucia Vardanega. Does culture influence understanding and perceived seriousness of plagiarism? International Journal for Educational Integrity, 4(2), 2008

  3. [11]

    Why do postgraduate students commit plagiarism? an empirical study

    Apatsa Selemani, Winner Dominic Chawinga, and Gift Dube. Why do postgraduate students commit plagiarism? an empirical study. International Journal for Educational Integrity, 14:1–15, 2018

  4. [12]

    Plagiarism, international students, and the second-language writer

    Diane Pecorari. Plagiarism, international students, and the second-language writer. In Second Handbook of Academic Integrity, pages 451–465. Springer, 2024

  5. [13]

    Plagiarism prevention through pedagogy: an instructional design approach

    Amanda Dinscore. Plagiarism prevention through pedagogy: an instructional design approach. Public Services Quarterly, 18(4):271–292, 2022

  6. [14]

    Stopping plagiarism through enculturation: A practice-based approach

    Shirley Ann McDonald and Ramine Adl. Stopping plagiarism through enculturation: A practice-based approach. Revue internationale des technologies en pédagogie universitaire, 16(2):86–99, 2019. 10 A PREPRINT - D ECEMBER 19, 2024

  7. [15]

    Curriculum redesign as a faculty-centred approach to plagiarism reduction

    Sue Hrasky and David Kronenberg. Curriculum redesign as a faculty-centred approach to plagiarism reduction. International Journal for Educational Integrity, 7(2), 2011

  8. [16]

    Comparing scientific abstracts generated by chatgpt to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers

    Catherine A Gao, Frederick M Howard, Nikolay S Markov, Emma C Dyer, Siddhi Ramesh, Yuan Luo, and Alexander T Pearson. Comparing scientific abstracts generated by chatgpt to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded hu...

  9. [17]

    Generative artificial intelligence in information systems education: Challenges, consequences, and responses

    Craig Van Slyke, Richard D Johnson, and Jalal Sarabadani. Generative artificial intelligence in information systems education: Challenges, consequences, and responses. Communications of the Association for Information Systems, 53(1):1–21, 2023

  10. [18]

    A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition

    Lorin W Anderson and David R Krathwohl. A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc., 2001

  11. [19]

    Developing case-based learning activities based on the revised bloom’s taxonomy

    Mathews Nkhoma, Tri Lam, Joan Richardson, K Kam, and Kwok Hung Lau. Developing case-based learning activities based on the revised bloom’s taxonomy. In Informing science & IT education conference (In SITE), pages 85–93, 2016

  12. [20]

    Assessing problem-based learning in a software engineering curriculum using bloom’s taxonomy and the ieee software engineering body of knowledge

    Peter Dolog, Lone Leth Thomsen, and Bent Thomsen. Assessing problem-based learning in a software engineering curriculum using bloom’s taxonomy and the ieee software engineering body of knowledge. ACM Transactions on Computing Education (TOCE), 16(3):1–41, 2016

  13. [21]

    Bloom’s learning outcomes’ automatic classification using lstm and pretrained word embeddings

    Sarang Shaikh, Sher Muhammad Daudpotta, and Ali Shariq Imran. Bloom’s learning outcomes’ automatic classification using lstm and pretrained word embeddings. IEEE Access, 9:117887–117909, 2021

  14. [22]

    Multi-agent system for students cognitive assessment in e-learning environment

    Rimsha Shahzad, Muhammad Aslam, Shaha Al-Otaibi, Muhammad Saqib Javed, Amjad Rehman Khan, Saeed Ali Bahaj, and Tanzila Saba. Multi-agent system for students cognitive assessment in e-learning environment. IEEE Access, 2024

  15. [23]

    Aligning learning outcomes and assessment methods: a web tool for e-learning courses

    Inés Gil-Jaurena and Sandra Kucina Softic. Aligning learning outcomes and assessment methods: a web tool for e-learning courses. International Journal of Educational Technology in Higher Education, 13:1–16, 2016

  16. [24]

    Supporting academic honesty in online courses

    Patricia McGee. Supporting academic honesty in online courses. Journal of Educators Online, 10(1):1–31, 2013

  17. [25]

    Curbing academic dishonesty in online courses

    Anita Krsak. Curbing academic dishonesty in online courses. In TCC, pages 159–170. TCCHawaii, 2007

  18. [26]

    Towards academic integrity: Using bloom’s taxonomy and technology to deter cheating in online courses

    Kakul Agha, Xia Zhu, and Gladson Chikwa. Towards academic integrity: Using bloom’s taxonomy and technology to deter cheating in online courses. In Technologies, Artificial Intelligence and the Future of Learning Post- COVID-19: The Crucial Role of International Accreditation, ...

  19. [27]

    Using writing assignment designs to mitigate plagiarism

    Nina C Heckler, David R Forde, and C Hobson Bryan. Using writing assignment designs to mitigate plagiarism. Teaching Sociology, 41(1):94–105, 2013

  20. [28]

    An academic integrity approach to learning and assessment design

    Margaret Hamilton and Joan Richardson. An academic integrity approach to learning and assessment design. Journal of Learning Design, 2(1):37–51, 2007

  21. [29]

    Andy Field and Graham J. Hole. How to Design and Report Experiments. SAGE Publications Ltd, London, 2002

  22. [30]

    New york university press

    Henry Jenkins. New york university press. Convergence Culture: where old and new media collide. New York University, pages 307–319, 2006

  23. [31]

    Convergence culture in the creative industries

    Mark Deuze. Convergence culture in the creative industries. International journal of cultural studies, 10(2):243– 263, 2007

  24. [32]

    Software takes command

    Lev Manovich. Software takes command. Bloomsbury Academic, 2013

  25. [33]

    How does search behavior change as search becomes more difficult? In Proceedings of the SIGCHI conference on human factors in computing systems, pages 35–44, 2010

    Anne Aula, Rehan M Khan, and Zhiwei Guan. How does search behavior change as search becomes more difficult? In Proceedings of the SIGCHI conference on human factors in computing systems, pages 35–44, 2010

  26. [34]

    Cyberchondria: studies of the escalation of medical concerns in web search

    Ryen W White and Eric Horvitz. Cyberchondria: studies of the escalation of medical concerns in web search. ACM Transactions on Information Systems (TOIS), 27(4):1–37, 2009

  27. [35]

    When self is the source: Effects of media customization on message processing

    Hyunjin Kang and S Shyam Sundar. When self is the source: Effects of media customization on message processing. Media Psychology, 19(4):561–588, 2016

  28. [36]

    A concise guide to market research

    Marko Sarstedt, Erik Mooi, et al. A concise guide to market research. The Process, Data, and, 12:1–7, 2014

  29. [37]

    Chatting and cheating: Ensuring academic integrity in the era of chatgpt

    Debby RE Cotton, Peter A Cotton, and J Reuben Shipway. Chatting and cheating: Ensuring academic integrity in the era of chatgpt. Innovations in education and teaching international, 61(2):228–239, 2024

  30. [38]

    Understanding the Turnitin similarity report, 2021

    Turnitin LLC. Understanding the Turnitin similarity report, 2021. Retrieved on January 25, 2024

  31. [39]

    Evaluating the efficacy of ai content detection tools in differentiating between human and ai-generated text

    Ahmed M Elkhatat, Khaled Elsaid, and Saeed Almeer. Evaluating the efficacy of ai content detection tools in differentiating between human and ai-generated text. International Journal for Educational Integrity, 19(1):17, 2023. 11 A PREPRINT - D ECEMBER 19, 2024

  32. [40]

    Discovering Statistics Using IBM SPSS Statistics

    Andy Field. Discovering Statistics Using IBM SPSS Statistics . SAGE Publications Ltd, 5 edition, 2017. 5th Edition

  33. [41]

    Repeated-measures analysis: issues and options

    Allen T Bramwell, Alvah C Bittner Jr, and Steven J Morrissey. Repeated-measures analysis: issues and options. International Journal of Industrial Ergonomics, 10(3):185–197, 1992

  34. [42]

    Statistical power analysis for the behavioral sciences

    Jacob Cohen. Statistical power analysis for the behavioral sciences. routledge, 2013

  35. [43]

    ’a real opportunity’: how ChatGPT could help college applicants

    Emily Bobrow. ’a real opportunity’: how ChatGPT could help college applicants. The Guardian, aug 2023

  36. [44]

    Is ChatGPT a threat to education? UC Riverside News, jan 2023

    Iqbal Pittalwala. Is ChatGPT a threat to education? UC Riverside News, jan 2023

  37. [45]

    Reviewing the performance of ai detection tools in differentiating between ai-generated and human-written texts: A literature and integrative hybrid review

    Chaka Chaka. Reviewing the performance of ai detection tools in differentiating between ai-generated and human-written texts: A literature and integrative hybrid review. Journal of Applied Learning and Teaching, 7(1), 2024

  38. [46]

    The great detectives: humans versus ai detectors in catching large language model-generated medical writing

    Jae QJ Liu, Kelvin TK Hui, Fadi Al Zoubi, Zing ZX Zhou, Dino Samartzis, Curtis CH Yu, Jeremy R Chang, and Arnold YL Wong. The great detectives: humans versus ai detectors in catching large language model-generated medical writing. International Journal for Educational Integrit...

  39. [47]

    Detecting and deterring plagiarism in social work students: Implications for learning for practice

    Karen Postle. Detecting and deterring plagiarism in social work students: Implications for learning for practice. Social Work Education, 28(4):351–362, 2009

  40. [48]

    Internet plagiarism: Developing strategies to curb student academic dishonesty

    M Jill Austin and Linda D Brown. Internet plagiarism: Developing strategies to curb student academic dishonesty. The Internet and higher education, 2(1):21–33, 1999

  41. [49]

    Helen Duffy and Samuel E. Weil. ChatGPT, cheating, and the future of education. The Harvard Crimson, feb 2023

  42. [50]

    Don’t ban ChatGPT in schools

    Kevin Roose. Don’t ban ChatGPT in schools. Teach with it. The New York Times, jan 2023

  43. [51]

    Knowledge building: Advancing the state of community knowledge

    Marlene Scardamalia and Carl Bereiter. Knowledge building: Advancing the state of community knowledge. International handbook of computer-supported collaborative learning, pages 261–279, 2021

  44. [52]

    Dellma: A framework for decision making under uncertainty with large language models

    Ollie Liu, Deqing Fu, Dani Yogatama, and Willie Neiswanger. Dellma: A framework for decision making under uncertainty with large language models. arXiv preprint arXiv:2402.02392, 2024

  45. [53]

    Foundation models for decision making: Problems, methods, and opportunities

    Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, Pieter Abbeel, and Dale Schuurmans. Foundation models for decision making: Problems, methods, and opportunities. arXiv preprint arXiv:2303.04129, 2023

  46. [54]

    A complete survey on llm-based ai chatbots

    Sumit Kumar Dam, Choong Seon Hong, Yu Qiao, and Chaoning Zhang. A complete survey on llm-based ai chatbots. arXiv preprint arXiv:2406.16937, 2024

  47. [55]

    Can large language models understand context?arXiv preprint arXiv:2402.00858, 2024

    Yilun Zhu, Joel Ruben Antony Moniz, Shruti Bhargava, Jiarui Lu, Dhivya Piraviperumal, Site Li, Yuan Zhang, Hong Yu, and Bo-Hsiang Tseng. Can large language models understand context?arXiv preprint arXiv:2402.00858, 2024

  48. [56]

    Creativity and machine learning: A survey

    Giorgio Franceschelli and Mirco Musolesi. Creativity and machine learning: A survey. ACM Computing Surveys, 56(11):1–41, 2024

  49. [57]

    Exploring the potential utility of ai large language models for medical ethics: an expert panel evaluation of gpt-4

    Michael Balas, Jordan Joseph Wadden, Philip C Hébert, Eric Mathison, Marika D Warren, Victoria Seavilleklein, Daniel Wyzynski, Alison Callahan, Sean A Crawford, Parnian Arjmand, et al. Exploring the potential utility of ai large language models for medical ethics: an expert pa...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.