Pith. sign in

REVIEW 2 cited by

Automating Outlier Detection via Meta-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.10606 v2 pith:KLU3LU4Q submitted 2020-09-22 cs.LG cs.IRstat.ML

classification cs.LGcs.IRstat.ML
keywords detectionmodeloutlierdatasetmeta-learningmetaodautomaticallybenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Given an unsupervised outlier detection (OD) task on a new dataset, how can we automatically select a good outlier detection method and its hyperparameter(s) (collectively called a model)? Thus far, model selection for OD has been a "black art"; as any model evaluation is infeasible due to the lack of (i) hold-out data with labels, and (ii) a universal objective function. In this work, we develop the first principled data-driven approach to model selection for OD, called MetaOD, based on meta-learning. MetaOD capitalizes on the past performances of a large body of detection models on existing outlier detection benchmark datasets, and carries over this prior experience to automatically select an effective model to be employed on a new dataset without using any labels. To capture task similarity, we introduce specialized meta-features that quantify outlying characteristics of a dataset. Through comprehensive experiments, we show the effectiveness of MetaOD in selecting a detection model that significantly outperforms the most popular outlier detectors (e.g., LOF and iForest) as well as various state-of-the-art unsupervised meta-learners while being extremely fast. To foster reproducibility and further research on this new problem, we open-source our entire meta-learning system, benchmark environment, and testbed datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. We Need to Rethink Benchmarking in Anomaly Detection

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Evaluating anomaly detection by averaging over diverse datasets is misleading; the paper proposes scenario-based benchmarking organized by shared structural properties.

  2. Coincident Learning for Beam-based RF Station Fault Identification Using Phase Information at the SLAC Linac Coherent Light Source

    physics.acc-ph 2025-05 conditional novelty 5.0 of 10

    Using RF phase data with the CoAD framework detects about three times as many RF station anomalies at LCLS as using amplitude data, and clusters them by fault type.

Pith tools