Pith. sign in

REVIEW 3 cited by

Detecting LLM-Generated Text in Computing Education: A Comparative Study for ChatGPT Cases

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.07411 v1 pith:LHR4J3CC submitted 2023-07-10 cs.CL cs.CY

classification cs.CLcs.CY
keywords llm-generatedtextdetectorsacademicchatgptdetectorfalseintegrity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to the recent improvements and wide availability of Large Language Models (LLMs), they have posed a serious threat to academic integrity in education. Modern LLM-generated text detectors attempt to combat the problem by offering educators with services to assess whether some text is LLM-generated. In this work, we have collected 124 submissions from computer science students before the creation of ChatGPT. We then generated 40 ChatGPT submissions. We used this data to evaluate eight publicly-available LLM-generated text detectors through the measures of accuracy, false positives, and resilience. The purpose of this work is to inform the community of what LLM-generated text detectors work and which do not, but also to provide insights for educators to better maintain academic integrity in their courses. Our results find that CopyLeaks is the most accurate LLM-generated text detector, GPTKit is the best LLM-generated text detector to reduce false positives, and GLTR is the most resilient LLM-generated text detector. We also express concerns over 52 false positives (of 114 human written submissions) generated by GPTZero. Finally, we note that all LLM-generated text detectors are less accurate with code, other languages (aside from English), and after the use of paraphrasing tools (like QuillBot). Modern detectors are still in need of improvements so that they can offer a full-proof solution to help maintain academic integrity. Further, their usability can be improved by facilitating a smooth API integration, providing clear documentation of their features and the understandability of their model(s), and supporting more commonly used languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Using Machine Learning to Distinguish Human-written from Machine-generated Creative Fiction

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Naive Bayes and MLP classifiers distinguish about 100-word excerpts of human-written detective fiction from ChatGPT-3.5 output with roughly 96% accuracy on in-domain tests.

  2. Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text

    cs.CL 2025-01 reject novelty 3.0 of 10

    On a 600-essay, two-topic dataset, TF-IDF features let a Random Forest identify which of five LLMs or a human wrote a text with about 97% accuracy, though the comparison to GPTZero is inconsistent.

  3. From Automation to Cognition: Redefining the Roles of Educators and Generative AI in Computing Education

    cs.CY 2024-12 conditional novelty 3.0 of 10

    Computing educators propose redesigning take-home assignments to include and assess student use of generative AI, while shifting educator focus to metacognitive skill development.

Pith tools