REVIEW 4 major objections 6 minor 57 references
The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that AI plagiarism falls as task complexity increases, and that assessments designed around higher-order thinking are a viable way to reduce AI-driven plagiarism.
desk verdict Useful experiment, but the headline claim that AI plagiarism falls monotonically with task complexity only holds in the ChatGPT group, and task specificity/order confound the causal story. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a task-complexity gradient built on Bloom's revised taxonomy, a six-level hierarchy of cognitive processes running from remembering up to creating. The authors designed three tasks (remember/understand, apply, analyze/evaluate/create), had two subject experts independently confirm the levels, and used a repeated-measures design in which every participant completed all three tasks in one of four tool conditions. The outcomes are Turnitin's two scores for the same submission: the similarity score, which measures text matching existing sources, and the AI plagiarism score, which measures text the detector judges to be AI-generated. The within-subjects comparison across the three complexity levels is what lets the authors attribute changes in those scores to task complexity rather than to differences between students.
What would settle it
Re-run the same three tasks with ChatGPT-4o and add a fourth task at the create level that relies only on general knowledge; if AI-plagiarism percentages on the general-knowledge create task stay as high as on the recall task, then the observed decline is driven by ChatGPT-3.5's lack of access to course-specific context, not by cognitive complexity per se.
Extended reading notes
Core claim
The paper's central claim is that AI plagiarism decreases as task complexity increases, so that assessments aimed at Bloom's higher-order levels—analyzing, evaluating, and creating—are a workable way to curb AI-driven plagiarism. The evidence is a within-subjects experiment: the same 123 students produced text for three tasks of increasing complexity, and Turnitin's AI-writing score fell from a mean of 27.42% across all groups on Task 1 to 9.75% on Task 3. The ChatGPT group showed the largest decline, from 71.89% to 26.74%, while the control group stayed near zero. The authors also report that similarity scores and AI-plagiarism scores are distinct: treatment groups looked similar on traditional similarity but differed sharply on AI detection, so they recommend using both metrics and human review, allowing for up to 20% false positives.
Load-bearing premise
The load-bearing premise is that the three tasks differ only in their Bloom's cognitive complexity; in reality the create-level Task 3 may also require local or course-specific knowledge unavailable to ChatGPT-3.5, and its classification as high order rests on two expert reviewers with no reported inter-rater reliability statistic, so the drop in AI plagiarism could come from task content rather than complexity.
Editorial extensions
If this is right
- If the central claim holds, redesigning assessments around analysis, evaluation, and creation should reduce AI-generated text in student submissions even when students have ChatGPT or similar tools open.
- Institutions that use only traditional similarity checks will miss AI-generated work; the paper implies both similarity and AI-plagiarism scores should be reported together.
- AI-plagiarism scores will carry false positives in the 0–20% range even when no generative AI was used, so automatic penalties should be set above that band or paired with human review.
- The assessment-led approach to academic integrity, already recommended in the pre-AI cheating literature, remains relevant for generative AI rather than being obsolete.
Reading between the lines
- The paper leaves implicit that its create-level Task 3 may draw on course-specific or lab-local knowledge that ChatGPT-3.5 cannot access, so the drop in AI plagiarism could reflect task specificity rather than cognitive complexity alone.
- Because the study used ChatGPT-3.5, the results should be read as a lower bound for current models; newer models that reason better at higher levels could shrink the gap, a possibility the paper's data do not test.
- A testable extension is to hold the Bloom level fixed while varying whether the task needs local knowledge, which would separate cognitive complexity from access-to-context effects.
- The paper's suggested 20% false-positive allowance is detector- and context-specific; an extension would calibrate that threshold per institution by running control submissions known to be human-written.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a classroom experiment with 123 students randomly assigned to a control, e-textbook, Google, or ChatGPT condition. Participants completed three data-privacy tasks intended to correspond to lower-, medium-, and higher-order Bloom's taxonomy levels, always in the same order. Turnitin similarity scores and Turnitin AI-writing-detection percentages were analyzed with repeated-measures ANOVA. The paper's central claim is that AI plagiarism decreases as task complexity increases, and that higher-order assessments are therefore an effective strategy for minimizing AI-assisted plagiarism.
Significance. If the central claim were fully supported, the paper would provide useful empirical evidence for a widely discussed but under-tested pedagogical recommendation: that assessments requiring higher-order thinking can deter AI-assisted plagiarism. The study has notable strengths: random assignment with cluster-based balancing, a pretest MANOVA showing group equivalence, a repeated-measures design, and explicit engagement with the false-positive limitation of AI detectors. The distinction between Turnitin similarity scores and AI-writing-detection scores is practically valuable. However, the causal claim is currently stronger than the design and analysis support, and several load-bearing points need to be addressed before the conclusions can be accepted as stated.
major comments (4)
- [Section 4.2 / Figure 1] The causal attribution that task complexity drives the change in AI plagiarism requires that the three tasks differ only in Bloom's cognitive level. The design does not ensure this: every participant completed Task 1, then Task 2, then Task 3, so task order and task content are fully confounded with complexity. In addition, Task 3 requires students to propose a solution to a problem they themselves identified, which plausibly requires local, personal, or laboratory-specific knowledge that ChatGPT-3.5 cannot access; Section 6.3 acknowledges that LLMs struggle with contextual understanding, but it does not treat this as a confound. The manuscript should either provide counterbalanced or task-matched data, or substantially soften the causal language. The expert validation in Section 4.2 should also be strengthened with an inter-rater reliability statistic, since the current statement of two experts agreeing does not quantify agreement.
- [Table 3 / Section 3] The abstract and conclusion claim that AI plagiarism decreases as task complexity increases, and Section 3 hypothesizes this 'regardless of the technology available to students.' The data in Table 3 do not support a monotonic decrease for the e-textbook group (6.68, 9.24, 4.68) or the Google group (11.34, 21.59, 2.28); both increase from Task 1 to Task 2. The total Task 1 and Task 2 means are nearly identical (27.42 and 27.02), and the overall decline is driven almost entirely by Task 3 and by the ChatGPT group (71.89, 64.92, 26.74). No within-group pairwise contrasts are reported, so the significant tasks-by-group interaction does not by itself establish the monotonic pattern claimed. The authors should report simple effects and pairwise comparisons for each group and revise the central claim to match the actual pattern.
- [Sections 4.5 and 6.1] The manuscript acknowledges that Turnitin's AI-writing-detection scores have false positives and recommends accounting for up to 20% false positives. Yet several of the substantive between-group and between-task differences used to support the conclusions fall within that range, for example the e-textbook Task 1 mean of 6.68 and Task 2 mean of 9.24, and the Google Task 2 mean of 21.59. The argument that a repeated-measures design mitigates false positives assumes that false-positive rates are stable across tasks and groups, which is not demonstrated, especially since the control group shows nonzero AI-plagiarism scores in Task 1 (4.68) despite having no access to tools. A sensitivity analysis restricted to groups and comparisons plausibly above the false-positive threshold, or an explicit modeling of detection noise, is needed before interpreting these values as genuine AI plagiarism.
- [Section 5] There is an internal inconsistency in the reported sample size for the ChatGPT group. Table 1 lists 38 students in the ChatGPT group, while Tables 2 and 3 report n = 37 for the ChatGPT group. The manuscript should clarify whether one participant was excluded post hoc, and if so, document the reason and report the degrees of freedom of all analyses accordingly.
minor comments (6)
- [Section 4.5] The sphericity notation appears incomplete: 'Mauchly's W =.973, 2(2) = 3.178' should include the chi-square symbol and value, for example χ²(2) = 3.178.
- [Section 4.5] The statement that 'the data obtained in each condition was found to be normally distributed' should be accompanied by the specific test statistics or a citation, since normality of bounded percentages is not self-evident.
- [Section 5] The text refers to 'Table 7' and 'Figure 5' when discussing AI plagiarism percentages, but the actual table is Table 3 and the figure is Figure 4. The cross-references should be corrected.
- [Section 5 / Figure 4] The sentence 'The interaction between groups and repeated measures shows that the Control and ChatGPT groups had a consistent performance with lower variability' is puzzling, because the ChatGPT group has the largest standard deviations in Table 3 (32.89, 37.69, 36.74). Please clarify what 'consistent' is intended to mean.
- [Section 2.1] The citation in the phrase 'with 'remember' being the least complex and 'create' being the most complex' appears to reference [18] only at the start of the paragraph; a direct citation at this sentence would help the reader locate the source of the complexity ordering.
- [Section 4.2] The phrase 'Participants in each completed the tasks in the following order' is grammatically incomplete; it should read 'Participants in each group completed the tasks in the following order.'
Circularity Check
No significant circularity detected: the paper's central claim is an empirical, measured result, not a self-referential derivation.
full rationale
The paper is an empirical experiment rather than a formal derivation, and no load-bearing step reduces to its own inputs. The independent variable (task complexity) is operationalized through Bloom's revised taxonomy with expert validation: Section 4.2 states that 'two independent subject experts confirmed the tasks' differentiation according to Bloom's taxonomy, agreeing on the levels for each task,' which is independent of the outcome measures (Turnitin similarity scores and AI plagiarism percentages). The dependent variables come from an external detection tool, and no parameter is fitted to a subset of the data and then presented as a prediction. There are no self-citations: neither author appears among the 57 references, and the load-bearing background claims (Bloom's taxonomy structure [18], AI detector unreliability [39, 45], and the assessment-deters-plagiarism literature [24-28]) all rest on external sources. The skeptical concerns about the causal claim — fixed task order (all participants completed Task 1, then Task 2, then Task 3), Task 3's possible demand for local or contextual knowledge, and the non-monotonic patterns in the e-textbook and Google groups — are threats to internal and causal validity, not circularity; the central claim was empirically falsifiable and, in fact, was not reproduced in the other treatment groups. Section 7.1 candidly acknowledges limits of sample, content area, and generalizability. The reader's-take score of 0.0 is appropriate: this is a self-contained empirical study with no circular derivation chain.
Assumptions & free parameters
assumptions (4)
- domain assumption Turnitin's AI writing score provides a valid relative measure of AI-generated content in student submissions.
- domain assumption The three tasks differ only in cognitive complexity (Bloom levels) and not in other factors such as answerability by ChatGPT or local knowledge.
- domain assumption Students followed the intended tool-use protocol under continuous monitoring.
- standard math Standard statistical assumptions for repeated measures ANOVA hold.
Cite this review
Pith. "Pith review of The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools." pith.science (2026). https://pith.science/paper/CMKLUMP2
@misc{pith2026241213412,
author = {Pith},
title = {Pith review of: The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools},
year = {2026},
howpublished = {\url{https://pith.science/paper/CMKLUMP2}},
note = {Machine review of arXiv:2412.13412}
}
read the original abstract
This study investigates whether assessments fostering higher-order thinking skills can reduce plagiarism involving generative AI tools. Participants completed three tasks of varying complexity in four groups: control, e-textbook, Google, and ChatGPT. Findings show that AI plagiarism decreases as task complexity increases, with higher-order tasks resulting in lower similarity scores and AI plagiarism percentages. The study also highlights the distinction between similarity scores and AI plagiarism, recommending both for effective plagiarism detection. Results suggest that assessments promoting higher-order thinking are a viable strategy for minimizing AI-driven plagiarism.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Chinese university students’ perceptions of plagiarism
Guangwei Hu and Jun Lei. Chinese university students’ perceptions of plagiarism. Ethics & Behavior, 25(3):233– 255, 2015
work page 2015
-
[2]
An empirical analysis of differences in plagiarism among world cultures
David C Ison. An empirical analysis of differences in plagiarism among world cultures. Journal of Higher Education Policy and Management, 40(4):291–304, 2018
work page 2018
-
[3]
Factors influencing plagiarism in higher education: A comparison of german and slovene students
Eva Jereb, Matjaž Perc, Barbara Lämmlein, Janja Jerebic, Marko Urh, Iztok Podbregar, and Polona Šprajc. Factors influencing plagiarism in higher education: A comparison of german and slovene students. PloS one, 13(8):e0202252, 2018
work page 2018
-
[4]
Plagiarism and its treatment in higher education
Peter J Larkham and Susan Manns. Plagiarism and its treatment in higher education. Journal of Further and Higher Education, 26(4):339–349, 2002
work page 2002
-
[5]
Factors associated with student plagiarism in a post-1992 university
Roger Bennett*. Factors associated with student plagiarism in a post-1992 university. Assessment & Evaluation in Higher Education, 30(2):137–162, 2005
work page 1992
-
[6]
Philip M Newton. How common is commercial contract cheating in higher education and is it increasing? a systematic review. In Frontiers in Education, volume 3, page 67. Frontiers Media SA, 2018
work page 2018
-
[7]
Christian Hopp and Alexander Speil. How prevalent is plagiarism among college students? anonymity preserving evidence from austrian undergraduates. Accountability in Research, 28(3):133–148, 2021
work page 2021
-
[8]
Is plagiarism really on the rise? results from four 5-yearly surveys
Guy J Curtis and Kell Tremayne. Is plagiarism really on the rise? results from four 5-yearly surveys. Studies in Higher Education, 46(9):1816–1826, 2021
work page 2021
Show all 57 references
-
[9]
Anupama Prashar, Priyanka Gupta, and Yogesh K. Dwivedi. Plagiarism awareness efforts, students’ ethical judgment and behaviors: a longitudinal experiment study on ethical nuances of plagiarism in higher education. Studies in Higher Education, 49(6):929–955, 2023
2023
-
[10]
Does culture influence understanding and perceived seriousness of plagiarism? International Journal for Educational Integrity, 4(2), 2008
Amanda Maxwell, Guy J Curtis, and Lucia Vardanega. Does culture influence understanding and perceived seriousness of plagiarism? International Journal for Educational Integrity, 4(2), 2008
2008
-
[11]
Why do postgraduate students commit plagiarism? an empirical study
Apatsa Selemani, Winner Dominic Chawinga, and Gift Dube. Why do postgraduate students commit plagiarism? an empirical study. International Journal for Educational Integrity, 14:1–15, 2018
2018
-
[12]
Plagiarism, international students, and the second-language writer
Diane Pecorari. Plagiarism, international students, and the second-language writer. In Second Handbook of Academic Integrity, pages 451–465. Springer, 2024
2024
-
[13]
Plagiarism prevention through pedagogy: an instructional design approach
Amanda Dinscore. Plagiarism prevention through pedagogy: an instructional design approach. Public Services Quarterly, 18(4):271–292, 2022
2022
-
[14]
Stopping plagiarism through enculturation: A practice-based approach
Shirley Ann McDonald and Ramine Adl. Stopping plagiarism through enculturation: A practice-based approach. Revue internationale des technologies en pédagogie universitaire, 16(2):86–99, 2019. 10 A PREPRINT - D ECEMBER 19, 2024
2019
-
[15]
Curriculum redesign as a faculty-centred approach to plagiarism reduction
Sue Hrasky and David Kronenberg. Curriculum redesign as a faculty-centred approach to plagiarism reduction. International Journal for Educational Integrity, 7(2), 2011
2011
-
[16]
Comparing scientific abstracts generated by chatgpt to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers
Catherine A Gao, Frederick M Howard, Nikolay S Markov, Emma C Dyer, Siddhi Ramesh, Yuan Luo, and Alexander T Pearson. Comparing scientific abstracts generated by chatgpt to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded hu...
2022
-
[17]
Generative artificial intelligence in information systems education: Challenges, consequences, and responses
Craig Van Slyke, Richard D Johnson, and Jalal Sarabadani. Generative artificial intelligence in information systems education: Challenges, consequences, and responses. Communications of the Association for Information Systems, 53(1):1–21, 2023
2023
-
[18]
A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition
Lorin W Anderson and David R Krathwohl. A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc., 2001
2001
-
[19]
Developing case-based learning activities based on the revised bloom’s taxonomy
Mathews Nkhoma, Tri Lam, Joan Richardson, K Kam, and Kwok Hung Lau. Developing case-based learning activities based on the revised bloom’s taxonomy. In Informing science & IT education conference (In SITE), pages 85–93, 2016
2016
-
[20]
Assessing problem-based learning in a software engineering curriculum using bloom’s taxonomy and the ieee software engineering body of knowledge
Peter Dolog, Lone Leth Thomsen, and Bent Thomsen. Assessing problem-based learning in a software engineering curriculum using bloom’s taxonomy and the ieee software engineering body of knowledge. ACM Transactions on Computing Education (TOCE), 16(3):1–41, 2016
2016
-
[21]
Bloom’s learning outcomes’ automatic classification using lstm and pretrained word embeddings
Sarang Shaikh, Sher Muhammad Daudpotta, and Ali Shariq Imran. Bloom’s learning outcomes’ automatic classification using lstm and pretrained word embeddings. IEEE Access, 9:117887–117909, 2021
2021
-
[22]
Multi-agent system for students cognitive assessment in e-learning environment
Rimsha Shahzad, Muhammad Aslam, Shaha Al-Otaibi, Muhammad Saqib Javed, Amjad Rehman Khan, Saeed Ali Bahaj, and Tanzila Saba. Multi-agent system for students cognitive assessment in e-learning environment. IEEE Access, 2024
2024
-
[23]
Aligning learning outcomes and assessment methods: a web tool for e-learning courses
Inés Gil-Jaurena and Sandra Kucina Softic. Aligning learning outcomes and assessment methods: a web tool for e-learning courses. International Journal of Educational Technology in Higher Education, 13:1–16, 2016
2016
-
[24]
Supporting academic honesty in online courses
Patricia McGee. Supporting academic honesty in online courses. Journal of Educators Online, 10(1):1–31, 2013
2013
-
[25]
Curbing academic dishonesty in online courses
Anita Krsak. Curbing academic dishonesty in online courses. In TCC, pages 159–170. TCCHawaii, 2007
2007
-
[26]
Towards academic integrity: Using bloom’s taxonomy and technology to deter cheating in online courses
Kakul Agha, Xia Zhu, and Gladson Chikwa. Towards academic integrity: Using bloom’s taxonomy and technology to deter cheating in online courses. In Technologies, Artificial Intelligence and the Future of Learning Post- COVID-19: The Crucial Role of International Accreditation, ...
2022
-
[27]
Using writing assignment designs to mitigate plagiarism
Nina C Heckler, David R Forde, and C Hobson Bryan. Using writing assignment designs to mitigate plagiarism. Teaching Sociology, 41(1):94–105, 2013
2013
-
[28]
An academic integrity approach to learning and assessment design
Margaret Hamilton and Joan Richardson. An academic integrity approach to learning and assessment design. Journal of Learning Design, 2(1):37–51, 2007
2007
-
[29]
Andy Field and Graham J. Hole. How to Design and Report Experiments. SAGE Publications Ltd, London, 2002
2002
-
[30]
New york university press
Henry Jenkins. New york university press. Convergence Culture: where old and new media collide. New York University, pages 307–319, 2006
2006
-
[31]
Convergence culture in the creative industries
Mark Deuze. Convergence culture in the creative industries. International journal of cultural studies, 10(2):243– 263, 2007
2007
-
[32]
Software takes command
Lev Manovich. Software takes command. Bloomsbury Academic, 2013
2013
-
[33]
How does search behavior change as search becomes more difficult? In Proceedings of the SIGCHI conference on human factors in computing systems, pages 35–44, 2010
Anne Aula, Rehan M Khan, and Zhiwei Guan. How does search behavior change as search becomes more difficult? In Proceedings of the SIGCHI conference on human factors in computing systems, pages 35–44, 2010
2010
-
[34]
Cyberchondria: studies of the escalation of medical concerns in web search
Ryen W White and Eric Horvitz. Cyberchondria: studies of the escalation of medical concerns in web search. ACM Transactions on Information Systems (TOIS), 27(4):1–37, 2009
2009
-
[35]
When self is the source: Effects of media customization on message processing
Hyunjin Kang and S Shyam Sundar. When self is the source: Effects of media customization on message processing. Media Psychology, 19(4):561–588, 2016
2016
-
[36]
A concise guide to market research
Marko Sarstedt, Erik Mooi, et al. A concise guide to market research. The Process, Data, and, 12:1–7, 2014
2014
-
[37]
Chatting and cheating: Ensuring academic integrity in the era of chatgpt
Debby RE Cotton, Peter A Cotton, and J Reuben Shipway. Chatting and cheating: Ensuring academic integrity in the era of chatgpt. Innovations in education and teaching international, 61(2):228–239, 2024
2024
-
[38]
Understanding the Turnitin similarity report, 2021
Turnitin LLC. Understanding the Turnitin similarity report, 2021. Retrieved on January 25, 2024
2021
-
[39]
Evaluating the efficacy of ai content detection tools in differentiating between human and ai-generated text
Ahmed M Elkhatat, Khaled Elsaid, and Saeed Almeer. Evaluating the efficacy of ai content detection tools in differentiating between human and ai-generated text. International Journal for Educational Integrity, 19(1):17, 2023. 11 A PREPRINT - D ECEMBER 19, 2024
2023
-
[40]
Discovering Statistics Using IBM SPSS Statistics
Andy Field. Discovering Statistics Using IBM SPSS Statistics . SAGE Publications Ltd, 5 edition, 2017. 5th Edition
2017
-
[41]
Repeated-measures analysis: issues and options
Allen T Bramwell, Alvah C Bittner Jr, and Steven J Morrissey. Repeated-measures analysis: issues and options. International Journal of Industrial Ergonomics, 10(3):185–197, 1992
1992
-
[42]
Statistical power analysis for the behavioral sciences
Jacob Cohen. Statistical power analysis for the behavioral sciences. routledge, 2013
2013
-
[43]
’a real opportunity’: how ChatGPT could help college applicants
Emily Bobrow. ’a real opportunity’: how ChatGPT could help college applicants. The Guardian, aug 2023
2023
-
[44]
Is ChatGPT a threat to education? UC Riverside News, jan 2023
Iqbal Pittalwala. Is ChatGPT a threat to education? UC Riverside News, jan 2023
2023
-
[45]
Reviewing the performance of ai detection tools in differentiating between ai-generated and human-written texts: A literature and integrative hybrid review
Chaka Chaka. Reviewing the performance of ai detection tools in differentiating between ai-generated and human-written texts: A literature and integrative hybrid review. Journal of Applied Learning and Teaching, 7(1), 2024
2024
-
[46]
The great detectives: humans versus ai detectors in catching large language model-generated medical writing
Jae QJ Liu, Kelvin TK Hui, Fadi Al Zoubi, Zing ZX Zhou, Dino Samartzis, Curtis CH Yu, Jeremy R Chang, and Arnold YL Wong. The great detectives: humans versus ai detectors in catching large language model-generated medical writing. International Journal for Educational Integrit...
2024
-
[47]
Detecting and deterring plagiarism in social work students: Implications for learning for practice
Karen Postle. Detecting and deterring plagiarism in social work students: Implications for learning for practice. Social Work Education, 28(4):351–362, 2009
2009
-
[48]
Internet plagiarism: Developing strategies to curb student academic dishonesty
M Jill Austin and Linda D Brown. Internet plagiarism: Developing strategies to curb student academic dishonesty. The Internet and higher education, 2(1):21–33, 1999
1999
-
[49]
Helen Duffy and Samuel E. Weil. ChatGPT, cheating, and the future of education. The Harvard Crimson, feb 2023
2023
-
[50]
Don’t ban ChatGPT in schools
Kevin Roose. Don’t ban ChatGPT in schools. Teach with it. The New York Times, jan 2023
2023
-
[51]
Knowledge building: Advancing the state of community knowledge
Marlene Scardamalia and Carl Bereiter. Knowledge building: Advancing the state of community knowledge. International handbook of computer-supported collaborative learning, pages 261–279, 2021
2021
-
[52]
Dellma: A framework for decision making under uncertainty with large language models
Ollie Liu, Deqing Fu, Dani Yogatama, and Willie Neiswanger. Dellma: A framework for decision making under uncertainty with large language models. arXiv preprint arXiv:2402.02392, 2024
2024 arXiv
-
[53]
Foundation models for decision making: Problems, methods, and opportunities
Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, Pieter Abbeel, and Dale Schuurmans. Foundation models for decision making: Problems, methods, and opportunities. arXiv preprint arXiv:2303.04129, 2023
2023 arXiv
-
[54]
A complete survey on llm-based ai chatbots
Sumit Kumar Dam, Choong Seon Hong, Yu Qiao, and Chaoning Zhang. A complete survey on llm-based ai chatbots. arXiv preprint arXiv:2406.16937, 2024
2024 arXiv
-
[55]
Can large language models understand context?arXiv preprint arXiv:2402.00858, 2024
Yilun Zhu, Joel Ruben Antony Moniz, Shruti Bhargava, Jiarui Lu, Dhivya Piraviperumal, Site Li, Yuan Zhang, Hong Yu, and Bo-Hsiang Tseng. Can large language models understand context?arXiv preprint arXiv:2402.00858, 2024
2024 arXiv
-
[56]
Creativity and machine learning: A survey
Giorgio Franceschelli and Mirco Musolesi. Creativity and machine learning: A survey. ACM Computing Surveys, 56(11):1–41, 2024
2024
-
[57]
Exploring the potential utility of ai large language models for medical ethics: an expert panel evaluation of gpt-4
Michael Balas, Jordan Joseph Wadden, Philip C Hébert, Eric Mathison, Marika D Warren, Victoria Seavilleklein, Daniel Wyzynski, Alison Callahan, Sean A Crawford, Parnian Arjmand, et al. Exploring the potential utility of ai large language models for medical ethics: an expert pa...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.