Pith. sign in

REVIEW 5 cited by

Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15264 v1 pith:2QTTMRA7 submitted 2023-10-23 cs.CL cs.AI

Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey

classification cs.CL cs.AI
keywords detectiontextai-generatedconcernsframeworksresearchaddresscommunity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) have revolutionized the domain of natural language processing (NLP) with remarkable capabilities of generating human-like text responses. However, despite these advancements, several works in the existing literature have raised serious concerns about the potential misuse of LLMs such as spreading misinformation, generating fake news, plagiarism in academia, and contaminating the web. To address these concerns, a consensus among the research community is to develop algorithmic solutions to detect AI-generated text. The basic idea is that whenever we can tell if the given text is either written by a human or an AI, we can utilize this information to address the above-mentioned concerns. To that end, a plethora of detection frameworks have been proposed, highlighting the possibilities of AI-generated text detection. But in parallel to the development of detection frameworks, researchers have also concentrated on designing strategies to elude detection, i.e., focusing on the impossibilities of AI-generated text detection. This is a crucial step in order to make sure the detection frameworks are robust enough and it is not too easy to fool a detector. Despite the huge interest and the flurry of research in this domain, the community currently lacks a comprehensive analysis of recent developments. In this survey, we aim to provide a concise categorization and overview of current work encompassing both the prospects and the limitations of AI-generated text detection. To enrich the collective knowledge, we engage in an exhaustive discussion on critical and challenging open questions related to ongoing research on AI-generated text detection.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability

    cs.CL 2026-07 accept novelty 7.0

    Telescope Perplexity, the average negative log probability a reference LM assigns to each token immediately after seeing it, yields strong zero-shot LLM-text detection by probing an early-training aversion to repetition.

  2. Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

    cs.CL 2026-06 unverdicted novelty 6.0

    DEW creates a robust watermark for LLM text by applying vector-space operations to dual embeddings and hiding the signal via key-seeded random projections, showing improved detection after paraphrasing and translation.

  3. Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy

    cs.CL 2026-03 conditional novelty 6.0

    AI-generated text detectors achieve high benchmark accuracy by exploiting unstable dataset-specific linguistic features, as evidenced by cross-domain degradation and differing SHAP explanations across corpora.

  4. Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

    cs.CL 2026-06 unverdicted novelty 5.0

    DEW is a semantic watermarking method for LLMs that derives a robust signal from dual embeddings via vector-space algebra and pseudo-random projections, remaining detectable after paraphrasing and translation.

  5. Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption

    cs.CR 2025-10 unverdicted novelty 4.0

    LLM watermarking adoption is limited by misaligned stakeholder incentives; incentive-aligned approaches such as in-context watermarking can enable practical use in targeted domains like education and peer review.