Pith. sign in

A Practical Examination of AI-Generated Text Detectors for Large Language Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The proliferation of large language models has raised growing concerns about their misuse, particularly in cases where AI-generated text is falsely attributed to human authors. Machine-generated content detectors claim to effectively identify such text under various conditions and from any language model. This paper critically evaluates these claims by assessing several popular detectors (RADAR, Wild, T5Sentinel, Fast-DetectGPT, PHD, LogRank, Binoculars) on a range of domains, datasets, and models that these detectors have not previously encountered. We employ various prompting strategies to simulate practical adversarial attacks, demonstrating that even moderate efforts can significantly evade detection. We emphasize the importance of the true positive rate at a specific false positive rate (TPR@FPR) metric and demonstrate that these detectors perform poorly in certain settings, with TPR@.01 as low as 0%. Our findings suggest that both trained and zero-shot detectors struggle to maintain high sensitivity while achieving a reasonable true positive rate.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

A Mathematical Theory of Discursive Networks

cs.CL · 2025-07-09 · reject · novelty 3.0

A two-state Markov model of error propagation suggests that small amounts of cross-agent peer review can flip a network of fallible language models from a falsehood-dominant to a truth-dominant state.

citing papers explorer

Showing 1 of 1 citing paper.

  • A Mathematical Theory of Discursive Networks cs.CL · 2025-07-09 · reject · none · ref 38 · internal anchor

    A two-state Markov model of error propagation suggests that small amounts of cross-agent peer review can flip a network of fallible language models from a falsehood-dominant to a truth-dominant state.