ExaGPT uses span-level similarity retrieval from human and LLM datastores to detect machine-generated text while supplying the matching spans as human-interpretable evidence, achieving up to 37-point accuracy gains over prior interpretable detectors at 1% FPR.
Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense
7 Pith papers cite this work, alongside 89 external citations. Polarity classification is still indexing.
verdicts
UNVERDICTED 7representative citing papers
Steering LLM residual streams with random sparse vectors creates detectable self-recognition fingerprints that enable over 98% accurate attribution of generated text to specific models without degrading output quality.
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.
Users in r/isthisAI and r/RealOrAI employ 12 evolving strategies for AI detection that shift with model capabilities and online trends.
Binoculars-inclusive ensembles detect AI text best overall but suffer the largest performance drops under paraphrasing attacks.
A 1.5B LLM fine-tuned on a curated rationale dataset (READ) detects AI text with explanations and reportedly outperforms much larger prompted LLMs.
Shared task findings show near-perfect binary detection of AI-generated text but greater difficulty in attributing outputs to particular language models.
citing papers explorer
-
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
ExaGPT uses span-level similarity retrieval from human and LLM datastores to detect machine-generated text while supplying the matching spans as human-interpretable evidence, achieving up to 37-point accuracy gains over prior interpretable detectors at 1% FPR.
-
LLM Self-Recognition: Steering and Retrieving Activation Signatures
Steering LLM residual streams with random sparse vectors creates detectable self-recognition fingerprints that enable over 98% accurate attribution of generated text to specific models without degrading output quality.
-
Asking Back: Interaction-Layer Antidistillation Watermarks
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.
-
Is This AI? Longitudinal Analysis of Strategies Used for AI Detection on Two Subreddits
Users in r/isthisAI and r/RealOrAI employ 12 evolving strategies for AI detection that shift with model capabilities and online trends.
-
Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods
Binoculars-inclusive ensembles detect AI text best overall but suffer the largest performance drops under paraphrasing attacks.
-
READER: Reasoning-Enhanced AI-Generated Text Detection
A 1.5B LLM fine-tuned on a curated rationale dataset (READ) detects AI text with explanations and reportedly outperforms much larger prompted LLMs.
-
Findings of the Counter Turing Test: AI-Generated Text Detection
Shared task findings show near-perfect binary detection of AI-generated text but greater difficulty in attributing outputs to particular language models.