Pith. sign in

REVIEW 1 cited by

A Robust Autoencoder Ensemble-Based Approach for Anomaly Detection in Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13031 v2 pith:3VYOQGYK submitted 2024-05-16 cs.CL cs.LG

classification cs.CLcs.LG
keywords anomalycontextualanomaliesapproachdetectionrobustapplicationsautoencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Anomaly detection (AD) is a fast growing and popular domain among established applications like vision and time series. We observe a rich literature for these applications, but anomaly detection in text is only starting to blossom. Recently, self-supervised methods with self-attention mechanism have been the most popular choice. While recent works have proposed a working ground for building and benchmarking state of the art approaches, we propose two principal contributions in this paper: contextual anomaly contamination and a novel ensemble-based approach. Our method, Textual Anomaly Contamination (TAC), allows to contaminate inlier classes with either independent or contextual anomalies. In the literature, it appears that this distinction is not performed. For finding contextual anomalies, we propose RoSAE, a Robust Subspace Local Recovery Autoencoder Ensemble. All autoencoders of the ensemble present a different latent representation through local manifold learning. Benchmark shows that our approach outperforms recent works on both independent and contextual anomalies, while being more robust. We also provide 8 dataset comparison instead of only relying to Reuters and 20 Newsgroups corpora.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLM embeddings improve text anomaly detection, shallow detectors match deep ones only under oracle embedding selection, and AUROC matrices are low-rank enough to support fast model evaluation.

Pith tools