REVIEW 6 cited by
RADAR: Robust AI-Text Detection via Adversarial Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advances in large language models (LLMs) and the intensifying popularity of ChatGPT-like applications have blurred the boundary of high-quality text generation between humans and machines. However, in addition to the anticipated revolutionary changes to our technology and society, the difficulty of distinguishing LLM-generated texts (AI-text) from human-generated texts poses new challenges of misuse and fairness, such as fake content generation, plagiarism, and false accusations of innocent writers. While existing works show that current AI-text detectors are not robust to LLM-based paraphrasing, this paper aims to bridge this gap by proposing a new framework called RADAR, which jointly trains a robust AI-text detector via adversarial learning. RADAR is based on adversarial training of a paraphraser and a detector. The paraphraser's goal is to generate realistic content to evade AI-text detection. RADAR uses the feedback from the detector to update the paraphraser, and vice versa. Evaluated with 8 different LLMs (Pythia, Dolly 2.0, Palmyra, Camel, GPT-J, Dolly 1.0, LLaMA, and Vicuna) across 4 datasets, experimental results show that RADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in place. We also identify the strong transferability of RADAR from instruction-tuned LLMs to other LLMs, and evaluate the improved capability of RADAR via GPT-3.5-Turbo.
Forward citations
Cited by 6 Pith papers
-
Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
A style-aware paraphrasing attack evades all nine tested AI-text detectors at the single-document level, but multi-document analysis makes the attack detectable again.
-
Performance Analysis and Optimization for Laser-Phase-Noise based Quantum Random Number Generation
A validated physical model predicts power spectrum and raw-data distributions for laser-phase-noise QRNGs, enabling quantitative rate optimization and proactive photonic-integrated design.
-
DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
DEER, a disentangled mixture-of-experts detector with RL-based instance routing, reports F1 gains of about 1.4 in-domain and 5.3 points out-of-domain over prior MGT detectors.
-
Domain Gating Ensemble Networks for AI-Generated Text Detection
DoGEN routes each input to the two most likely domain-specialist detectors and blends their scores, beating internally trained single models on in-domain MAGE and out-of-domain RAID AUROC.
-
DAMAGE: Detecting Adversarially Modified AI Generated Text
Adding humanizer-processed text to training data yields a detector that catches 98.26% of humanized AI essays at a 5% false-positive rate and stays robust to a detector-targeted attack.
-
Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems
A compilation of three research programs showing that GPT detectors are biased against non-native writers, that population-level estimates place AI-modified text at up to 16.9% of AI-conference reviews and up to 24% i...
Discussion (0). Continue with ORCID to comment.