REVIEW 2 cited by
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text anomaly detection is crucial for identifying spam, misinformation, and offensive language in natural language processing tasks. Despite the growing adoption of embedding-based methods, their effectiveness and generalizability across diverse application scenarios remain under-explored. To address this, we present TAD-Bench, a comprehensive benchmark designed to systematically evaluate embedding-based approaches for text anomaly detection. TAD-Bench integrates multiple datasets spanning different domains, combining state-of-the-art embeddings from large language models with a variety of anomaly detection algorithms. Through extensive experiments, we analyze the interplay between embeddings and detection methods, uncovering their strengths, weaknesses, and applicability to different tasks. These findings offer new perspectives on building more robust, efficient, and generalizable anomaly detection systems for real-world applications.
Forward citations
Cited by 2 Pith papers
-
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
LLM embeddings improve text anomaly detection, shallow detectors match deep ones only under oracle embedding selection, and AUROC matrices are low-rank enough to support fast model evaluation.
-
Anomaly Detection in Human Language via Meta-Learning: A Few-Shot Approach
A meta-learning framework with cross-domain episode sampling achieves higher anomaly detection AUC and F1 than fine-tuned BERT and one-class SVM on SMS spam, COVID-19 fake news, and hate speech tasks.
Discussion (0). Sign in to comment.