REVIEW 4 cited by
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The powerful ability to understand, follow, and generate complex language emerging from large language models (LLMs) makes LLM-generated text flood many areas of our daily lives at an incredible speed and is widely accepted by humans. As LLMs continue to expand, there is an imperative need to develop detectors that can detect LLM-generated text. This is crucial to mitigate potential misuse of LLMs and safeguard realms like artistic expression and social networks from harmful influence of LLM-generated content. The LLM-generated text detection aims to discern if a piece of text was produced by an LLM, which is essentially a binary classification task. The detector techniques have witnessed notable advancements recently, propelled by innovations in watermarking techniques, statistics-based detectors, neural-base detectors, and human-assisted methods. In this survey, we collate recent research breakthroughs in this area and underscore the pressing need to bolster detector research. We also delve into prevalent datasets, elucidating their limitations and developmental requirements. Furthermore, we analyze various LLM-generated text detection paradigms, shedding light on challenges like out-of-distribution problems, potential attacks, real-world data issues and the lack of effective evaluation framework. Conclusively, we highlight interesting directions for future research in LLM-generated text detection to advance the implementation of responsible artificial intelligence (AI). Our aim with this survey is to provide a clear and comprehensive introduction for newcomers while also offering seasoned researchers a valuable update in the field of LLM-generated text detection. The useful resources are publicly available at: https://github.com/NLP2CT/LLM-generated-Text-Detection.
Forward citations
Cited by 4 Pith papers
-
Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation
In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.
-
The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text
Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.
-
How does Misinformation Affect Large Language Model Behaviors and Preferences?
MisBench provides a 10.3M-example benchmark of styled, conflict-based misinformation and shows LLMs' detection accuracy depends strongly on conflict type and textual style.
-
AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models
Fine-tuned GPT-4o-mini detects AI versus human text at 95.5% F1 on the DeFactify test set, but identifying the specific generator LLM reaches only 47% F1 with BERT.
Discussion (0). Continue with ORCID to comment.