pith. machine review for the scientific record. sign in

arxiv: 1907.03813 · v1 · submitted 2019-07-08 · 📊 stat.ML · cs.LG· math.ST· stat.TH

Recognition: unknown

Statistical Analysis of Nearest Neighbor Methods for Anomaly Detection

Authors on Pith no claims yet
classification 📊 stat.ML cs.LGmath.STstat.TH
keywords anomalydetectionmethodsanalysisdatasetsgeometricperformancealgorithms
0
0 comments X
read the original abstract

Nearest-neighbor (NN) procedures are well studied and widely used in both supervised and unsupervised learning problems. In this paper we are concerned with investigating the performance of NN-based methods for anomaly detection. We first show through extensive simulations that NN methods compare favorably to some of the other state-of-the-art algorithms for anomaly detection based on a set of benchmark synthetic datasets. We further consider the performance of NN methods on real datasets, and relate it to the dimensionality of the problem. Next, we analyze the theoretical properties of NN-methods for anomaly detection by studying a more general quantity called distance-to-measure (DTM), originally developed in the literature on robust geometric and topological inference. We provide finite-sample uniform guarantees for the empirical DTM and use them to derive misclassification rates for anomalous observations under various settings. In our analysis we rely on Huber's contamination model and formulate mild geometric regularity assumptions on the underlying distribution of the data.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPC

    cs.PF 2026-04 conditional novelty 6.0

    COMPASS formalizes HPC configuration questions as ML tasks on traces, quantifies recommendation trustworthiness, and delivers 65.93% lower average job turnaround time plus 80.93% lower node usage versus prior methods ...