Pith. sign in

REVIEW 2 cited by

Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15666 v2 pith:JWUDRXY6 submitted 2025-02-21 cs.CL cs.AIcs.HCcs.LG

Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing

classification cs.CL cs.AIcs.HCcs.LG
keywords textai-generatedcontentai-polishedalmostchallengedetectiondetectors
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The growing use of large language models (LLMs) for text generation has led to widespread concerns about AI-generated content detection. However, an overlooked challenge is AI-polished text, where human-written content undergoes subtle refinements using AI tools. This raises a critical question: should minimally polished text be classified as AI-generated? Such classification can lead to false plagiarism accusations and misleading claims about AI prevalence in online content. In this study, we systematically evaluate twelve state-of-the-art AI-text detectors using our AI-Polished-Text Evaluation (APT-Eval) dataset, which contains 14.7K samples refined at varying AI-involvement levels. Our findings reveal that detectors frequently flag even minimally polished text as AI-generated, struggle to differentiate between degrees of AI involvement, and exhibit biases against older and smaller models. These limitations highlight the urgent need for more nuanced detection methodologies.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

    cs.CL 2026-03 accept novelty 6.5

    State-of-the-art AI detectors misclassify a non-trivial fraction of LLM-polished peer reviews as fully AI-generated, rendering polishing-only policies currently unenforceable.

  2. Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift

    cs.CL 2026-06 unverdicted novelty 6.0

    Test-time adaptation with semi-supervised learning leverages inference-time homogeneity to maintain AI text detection performance under adversarial humanization, new LLMs, and temporal drift.