Pith. sign in

REVIEW 2 cited by

TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11960 v2 pith:YZXPGCLV submitted 2025-01-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords detectionanomalyembedding-basedlanguagetad-benchtextbenchmarkcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text anomaly detection is crucial for identifying spam, misinformation, and offensive language in natural language processing tasks. Despite the growing adoption of embedding-based methods, their effectiveness and generalizability across diverse application scenarios remain under-explored. To address this, we present TAD-Bench, a comprehensive benchmark designed to systematically evaluate embedding-based approaches for text anomaly detection. TAD-Bench integrates multiple datasets spanning different domains, combining state-of-the-art embeddings from large language models with a variety of anomaly detection algorithms. Through extensive experiments, we analyze the interplay between embeddings and detection methods, uncovering their strengths, weaknesses, and applicability to different tasks. These findings offer new perspectives on building more robust, efficient, and generalizable anomaly detection systems for real-world applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLM embeddings improve text anomaly detection, shallow detectors match deep ones only under oracle embedding selection, and AUROC matrices are low-rank enough to support fast model evaluation.

  2. Anomaly Detection in Human Language via Meta-Learning: A Few-Shot Approach

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A meta-learning framework with cross-domain episode sampling achieves higher anomaly detection AUC and F1 than fine-tuned BERT and one-class SVM on SMS spam, COVID-19 fake news, and hate speech tasks.

Pith tools