Pith. sign in

REVIEW 1 major objections 1 minor 180 references

Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version

T0 review · 1 major / 1 minor · reviewed 2026-05-20 · grok-4.3

Pith's one-line read MetaEns automatically selects compact high-quality outlier detection ensembles without labels by learning to predict marginal gains from meta-datasets.

desk verdict MetaEns meta-learns marginal gains for outlier ensemble selection and pairs them with a diversity-aware proxy, delivering compact high-AP ensembles on 39 datasets, but the zero-shot transfer step remains the weakest link. read the letter →

arxiv 2605.16567 v1 pith:LJYXT3E6 submitted 2026-05-15 cs.LG cs.AIcs.DB

classification cs.LGcs.AIcs.DB
keywords unsupervisedoutlierdetectionensemblemodelselectionmeta-learningmarginalgainssubmodulardiversityregularizationensembles
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes MetaEns as a way to form ensembles of outlier detectors when no ground-truth labels exist for the target data. It trains a predictor on labeled meta-datasets to estimate the expected improvement from adding each candidate model to a growing ensemble. This signal is combined with a proxy objective that encourages diversity and penalizes risk at the model-family level, allowing greedy selection to stop early when further additions yield little benefit. A sympathetic reader would care because outlier detection is typically unsupervised and single models can be unreliable, yet naively combining many models leads to redundancy and wasted computation. If the approach works, practitioners could obtain more accurate detection with smaller, more efficient model sets across varied real-world data.

What carries the argument

A meta-learned predictor of marginal ensemble gains combined with a submodular proxy objective enforcing diminishing returns through diversity discounting and family risk regularization.

What would settle it

Testing the selected ensembles on additional unlabeled real-world datasets and finding that they fail to achieve higher average precision than state-of-the-art unsupervised selectors while also using more models would falsify the performance claims.

Watch

Extended reading notes

Core claim

MetaEns learns a model on labeled meta-datasets to predict marginal ensemble gains and then, at test time on unlabeled data, uses this signal together with a submodular-inspired proxy objective that applies diversity-aware discounting and family-level risk regularization to drive greedy sequential selection with adaptive early stopping, thereby constructing compact high-quality ensembles.

Load-bearing premise

A model trained to predict marginal ensemble gains on labeled meta-datasets will produce accurate and useful signals when applied to new, unlabeled target datasets.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper proposes MetaEns, an automatic unsupervised framework for selecting compact ensembles of outlier detection models. It trains a meta-predictor on labeled meta-datasets to estimate marginal ensemble gains (improvement in average precision when adding a candidate model to a partial ensemble), then at test time combines this signal with a submodular-inspired proxy objective incorporating diversity-aware discounting and family-level risk regularization to enable greedy sequential selection with adaptive early stopping on unlabeled target datasets. Experiments on 39 real-world datasets report that MetaEns outperforms state-of-the-art unsupervised selectors and ensemble baselines in average precision while using fewer models.

Significance. If the meta-predictor transfers reliably, the approach could meaningfully advance unsupervised outlier detection by automating the construction of robust, computationally efficient ensembles without requiring labels on the target data. The submodular proxy provides a structured way to enforce diminishing returns and diversity, which is a strength relative to purely heuristic selection methods.

major comments (1)
  1. [Abstract and methodology description of meta-predictor training and test-time application] The central claim depends on zero-shot transfer of the meta-predictor (trained on labeled meta-datasets) to new unlabeled target datasets, yet no section isolates or validates the predictor's accuracy or ranking quality on held-out distributions whose statistics may differ from the meta-training data. The 39-dataset end-to-end results therefore do not rule out that reported gains arise from favorable meta-dataset selection or post-hoc choices rather than robust generalization of the marginal-gain estimates.
minor comments (1)
  1. [Methodology] Clarify the exact input features to the meta-predictor (model-family statistics, data characteristics, etc.) and whether any normalization or invariance properties are assumed or enforced.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive feedback and for recognizing the potential of MetaEns to advance unsupervised outlier detection. We address the major comment below.

read point-by-point responses
  1. Referee: The central claim depends on zero-shot transfer of the meta-predictor (trained on labeled meta-datasets) to new unlabeled target datasets, yet no section isolates or validates the predictor's accuracy or ranking quality on held-out distributions whose statistics may differ from the meta-training data. The 39-dataset end-to-end results therefore do not rule out that reported gains arise from favorable meta-dataset selection or post-hoc choices rather than robust generalization of the marginal-gain estimates.

    Authors: We agree that isolating the meta-predictor's performance on held-out distributions would provide stronger direct evidence for zero-shot transfer. The current manuscript emphasizes end-to-end results on 39 diverse real-world datasets to demonstrate practical utility, but these do not separately quantify the predictor's ranking quality or accuracy under distribution shift. In the revised version we will add a new subsection that holds out a subset of meta-datasets, evaluates the meta-predictor on those held-out sets (reporting correlation between predicted and observed marginal gains as well as top-k selection accuracy), and discusses how the meta-training distribution was constructed to promote generalization. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; meta-learning framework is empirically grounded

full rationale

The paper describes a meta-learning method that trains a predictor on separate labeled meta-datasets to estimate marginal ensemble gains, then applies the predictor plus a submodular proxy objective to select ensembles on new unlabeled target datasets. This structure does not reduce any claimed result to its inputs by construction: the predictor is explicitly fitted on distinct meta-data rather than self-defined, the test-time selection uses an independent proxy, and performance claims rest on end-to-end experiments across 39 real-world datasets rather than tautological renaming or self-citation chains. No equations or steps in the provided description exhibit fitted inputs relabeled as independent predictions or uniqueness imported from prior self-work. The transfer assumption is a methodological risk but does not constitute circularity under the specified patterns.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Only the abstract is available, so the ledger is necessarily incomplete; the method rests on the representativeness of the meta-datasets and on the assumption that the learned predictor transfers to unseen data.

assumptions (1)
  • domain assumption Labeled meta-datasets are sufficiently representative of the distribution of real-world outlier detection tasks encountered at test time.
    The entire meta-learning step depends on this transfer assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version." pith.science (2026). https://pith.science/paper/LJYXT3E6

@misc{pith2026260516567,
  author       = {Pith},
  title        = {Pith review of: Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LJYXT3E6}},
  note         = {Machine review of arXiv:2605.16567}
}
read the original abstract

Unsupervised outlier detection is attractive because it eliminates the need for labeled data. Moreover, forming multi-model ensembles can improve detection robustness. However, composing an ensemble without labeled data is challenging. Naively composed ensembles can suffer from ensemble saturation, where redundant or unreliable detection models degrade performance and incur unnecessary computation. We propose MetaEns, an automatic unsupervised framework for selecting ensembles of outlier detection models. Using labeled meta-datasets, MetaEns learns a model that predicts marginal ensemble gains, estimating the expected improvement from adding a candidate model to a partially constructed ensemble. At test time, this learned signal is combined with a submodular-inspired proxy objective that enforces diminishing returns through diversity-aware discounting and family-level risk regularization, thereby enabling greedy sequential selection with adaptive early stopping. As a result, MetaEns constructs compact, high-quality ensembles without access to ground-truth labels. Experiments on 39 real-world datasets show that MetaEns consistently outperforms state-of-the-art unsupervised selectors and ensemble baselines, achieving higher average precision while using fewer models.

Figures

Figures reproduced from arXiv: 2605.16567 by the authors.

Figure 1
Figure 1. Overview of MetaEns. (A) Offline Meta-Training: Oracle-greedy rollouts on labeled meta-datasets M generate state–gain pairs for partial ensembles. These pairs supervise a two-part gain predictor consisting of classifier fcls, which estimates whether a candidate will improve the ensemble, and regressor freg, which estimates the positive gain magnitude. Family-risk priors πF are computed from lower-tail oracle gains t… view at source ↗
Figure 2
Figure 2. Across all selectors, MetaEns consistently improves over the starting model, demonstrating that its partner selection mechanism is selector-agnostic and not tied to a particular initialization strategy. Improvements are especially pro￾nounced in challenging cases where the primary model underperforms. Rather than propagating initial errors, MetaEns effectively recovers performance by selecting complementary detector… view at source ↗
Figure 2
Figure 2. Robustness analysis across four different primary selectors: ELECT, LOF, IForest, and Random Selection. Each panel compares the primary model’s performance (x-axis) against the final MetaEns ensemble (y-axis). Points above the diagonal indicate improvement. The shaded red “Rescue Zone” highlights where the primary model fails (AP < 0.4). MetaEns consistently rescues performance in these failure modes across all prim… view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: Model diversity visualization using t-SNE projection across four datasets. ELECT-10 selections tend to cluster within a single family, whereas MetaEns selects models spanning multiple families, indicating greater ensemble diversity. ods learn expressive representations…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

180 extracted references · 180 canonical work pages

  1. [1]

    Arthur Zimek and Ricardo J. G. B. Campello and J. Ensembles for unsupervised outlier detection: challenges and research questions a position paper , journal =

  2. [2]

    Deep One-Class Classification , booktitle =

    Lukas Ruff and Nico G. Deep One-Class Classification , booktitle =

  3. [3]

    Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection , booktitle =

    Bo Zong and Qi Song and Martin Renqiang Min and Wei Cheng and Cristian Lumezanu and Dae. Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection , booktitle =

  4. [4]

    Breunig and Hans

    Markus M. Breunig and Hans. Proceedings of the

  5. [5]

    Isolation Forest , booktitle =

    Fei Tony Liu and Kai Ming Ting and Zhi. Isolation Forest , booktitle =

  6. [6]

    Zheng Li and Yue Zhao and Nicola Botta and Cezar Ionescu and Xiyang Hu , year = 2020, booktitle =

  7. [7]

    Yue Zhao and Zain Nasrullah and Zheng Li , year = 2019, journal =. PyOD:

  8. [8]

    CoRR , volume =

    A Large-scale Study on Unsupervised Outlier Model Selection: Do Internal Strategies Suffice? , author =. CoRR , volume =

Show all 180 references
  1. [9]

    Proceedings of the

    Shebuti Rayana and Wen Zhong and Leman Akoglu , title =. Proceedings of the

  2. [10]

    Marques and Ricardo J

    Henrique O. Marques and Ricardo J. G. B. Campello and J. Internal Evaluation of Unsupervised Outlier Detection , journal =

  3. [11]

    Varun Chandola and Arindam Banerjee and Vipin Kumar , title =

  4. [12]

    ADBench: Anomaly Detection Benchmark , author =

  5. [13]

    CoRR , volume =

    Automating Outlier Detection via Meta-Learning , author =. CoRR , volume =

  6. [14]

    Nemhauser and Laurence A

    George L. Nemhauser and Laurence A. Wolsey and Marshall L. Fisher , title =. Math. Program. , volume =

  7. [15]

    Alex Kulesza and Ben Taskar , title =. Found. Trends Mach. Learn. , volume =

  8. [16]

    Rossi and Leman Akoglu , title =

    Yue Zhao and Ryan A. Rossi and Leman Akoglu , title =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  9. [17]

    Proceedings of the

    Toward Unsupervised Outlier Model Selection , author =. Proceedings of the

  10. [18]

    Proceedings of the

    Li Cheng and Yijie Wang and Xinwang Liu and Bin Li , title =. Proceedings of the

  11. [19]

    Hospedales and Antreas Antoniou and Paul Micaelli and Amos Storkey , title =

    Timothy M. Hospedales and Antreas Antoniou and Paul Micaelli and Amos Storkey , title =

  12. [20]

    Theoretical Foundations and Algorithms for Outlier Ensembles , author =

  13. [21]

    CoRR , volume =

    How to Evaluate the Quality of Unsupervised Anomaly Detection Algorithms? , author =. CoRR , volume =

  14. [22]

    Kauffmann and Robert A

    Lukas Ruff and Jacob R. Kauffmann and Robert A. Vandermeulen and Gr. A Unifying Review of Deep and Shallow Anomaly Detection , journal =

  15. [23]

    Proceedings of the

    AutoOD: Neural Architecture Search for Outlier Detection , author =. Proceedings of the

  16. [24]

    Proceedings of the Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining (PAKDD) , pages =

    An Unsupervised Boosting Strategy for Outlier Detection Ensembles , author =. Proceedings of the Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining (PAKDD) , pages =

  17. [25]

    Hryniewicki and Zheng Li , year = 2019, booktitle =

    Yue Zhao and Zain Nasrullah and Maciej K. Hryniewicki and Zheng Li , year = 2019, booktitle =

  18. [26]

    Less is More: Building Selective Anomaly Ensembles , author =

  19. [27]

    Proceedings of the Conference on Annual Meeting of the Association for Computational Linguistics (ACL) , pages =

    A Class of Submodular Functions for Document Summarization , author =. Proceedings of the Conference on Annual Meeting of the Association for Computational Linguistics (ACL) , pages =

  20. [28]

    Near-Optimal Sensor Placements in Gaussian Processes: Theory, Efficient Algorithms and Empirical Studies , author =. J. Mach. Learn. Res. , volume = 9, pages =

  21. [29]

    Proceedings of the International Conference on Machine Learning (ICML) , pages =

    Near-optimal Batch Mode Active Learning and Adaptive Submodular Optimization , author =. Proceedings of the International Conference on Machine Learning (ICML) , pages =

  22. [30]

    Data Min

    On the Evaluation of Unsupervised Outlier Detection: Measures, Datasets, and an Empirical Study , author =. Data Min. Knowl. Discov. , volume = 30, pages =

  23. [31]

    LightGBM:

    Guolin Ke and Qi Meng and Thomas Finley and Taifeng Wang and Wei Chen and Weidong Ma and Qiwei Ye and Tie. LightGBM:. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  24. [32]

    Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy , author =

  25. [33]

    Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

    PyTorch: An Imperative Style, High-Performance Deep Learning Library , author =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  26. [34]

    Sunil Aryal and Arbind Agrahari Baniya and Imran Razzak and KC Santosh , year = 2021, booktitle =

  27. [35]

    Chen , year = 2023, journal =

    Zheng Li and Yue Zhao and Xiyang Hu and Nicola Botta and Cezar Ionescu and George H. Chen , year = 2023, journal =

  28. [36]

    Proceedings of the International Conference on Pattern Recognition (ICPR) , pages =

    Outlier Detection Using k-Nearest Neighbour Graph , author =. Proceedings of the International Conference on Pattern Recognition (ICPR) , pages =

  29. [37]

    Proceedings of the International Joint Conference on Neural Networks (IJCNN) , pages =

    An Outlier Detection Algorithm based on KNN-kernel Density Estimation , author =. Proceedings of the International Joint Conference on Neural Networks (IJCNN) , pages =

  30. [38]

    Proceedings of the

    K-Nearest Neighbor Search and Outlier Detection via Minimax Distances , author =. Proceedings of the

  31. [39]

    Anomaly Detection with Score Functions Based on the Reconstruction Error of the Kernel

    Laetitia Chapel and Chlo. Anomaly Detection with Score Functions Based on the Reconstruction Error of the Kernel. Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD) , pages =

  32. [40]

    Proceedings of the International Conference on World Wide Web (WWW) , pages =

    Unsupervised Anomaly Detection via Variational Auto-Encoder for Seasonal KPIs in Web Applications , author =. Proceedings of the International Conference on World Wide Web (WWW) , pages =

  33. [41]

    Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI) , pages =

    Robustness of Autoencoders for Anomaly Detection Under Adversarial Impact , author =. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI) , pages =

  34. [42]

    Proceedings of the

    Swee Kiat Lim and Yi Loo and Ngoc. Proceedings of the

  35. [43]

    Proceedings of the

    Transformer for Point Anomaly Detection , author =. Proceedings of the

  36. [44]

    Proceedings of the

    Anomaly Detection with Robust Deep Autoencoders , author =. Proceedings of the

  37. [45]

    Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

    Support Vector Method for Novelty Detection , author =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  38. [46]

    Support Vector Data Description , author =. Mach. Learn. , volume = 54, number = 1, pages =

  39. [47]

    Proceedings of the

    Feature Bagging for Outlier Detection , author =. Proceedings of the

  40. [48]

    Hryniewicki , year = 2018, booktitle =

    Yue Zhao and Maciej K. Hryniewicki , year = 2018, booktitle =

  41. [49]

    Proceedings of the

    Outlier Detection with Autoencoder Ensembles , author =. Proceedings of the

  42. [50]

    Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI) , pages =

    Outlier Detection for Time Series with Recurrent Autoencoder Ensembles , author =. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI) , pages =

  43. [51]

    Proceedings of the

    Anomaly Detection Using an Ensemble of Feature Models , author =. Proceedings of the

  44. [52]

    Proceedings of the

    Xu Han and Xiaohui Chen and Li. Proceedings of the

  45. [53]

    Goadrich , year = 2006, booktitle =

    Jesse Davis and Mark H. Goadrich , year = 2006, booktitle =. The relationship between Precision-Recall and

  46. [54]

    Scikit-learn: Machine Learning in Python , author =. J. Mach. Learn. Res. , volume = 12, pages =

  47. [55]

    Kybernetika , volume = 17, number = 6, pages =

    Estimating the Dimension of a Linear Model , author =. Kybernetika , volume = 17, number = 6, pages =

  48. [56]

    Estimating the Dimension of a Model , author =. Ann. Stat. , volume = 6, number = 2, pages =

  49. [57]

    The Determination of the Order of an Autoregression , author =. J. R. Stat. Soc., B: Stat. Methodol. , volume = 41, number = 2, pages =

  50. [58]

    Cross-Validatory Choice and Assessment of Statistical Predictions , author =. J. R. Stat. Soc., B: Stat. Methodol. , volume = 36, number = 1, pages =

  51. [59]

    Estimating the Error Rate of a Prediction Rule: Improvement on Cross-validation , author =. J. Am. Stat. Assoc. , volume = 78, number = 382, pages =

  52. [60]

    Statistical Learning Theory , author =

  53. [61]

    Rademacher and Gaussian Complexities: Risk Bounds and Structural Results , author =. J. Mach. Learn. Res. , volume = 3, pages =

  54. [62]

    Random Search for Hyper-Parameter Optimization , author =. J. Mach. Learn. Res. , volume = 13, pages =

  55. [63]

    Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

    Practical Bayesian Optimization of Machine Learning Algorithms , author =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  56. [64]

    Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

    Efficient and Robust Automated Machine Learning , author =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , pages =

  57. [65]

    Covariate Shift Adaptation by Importance Weighted Cross Validation , author =. J. Mach. Learn. Res. , volume = 8, pages =

  58. [66]

    In Search of Lost Domain Generalization , author =

  59. [67]

    A Perspective View and Survey of Meta-Learning , author =. Artif. Intell. Rev. , volume = 18, number = 2, pages =

  60. [68]

    Extremely Randomized Trees , author =. Mach. Learn. , volume = 63, number = 1, pages =

  61. [69]

    Visualizing Data using t-SNE , author =. J. Mach. Learn. Res. , volume = 9, number = 86, pages =

  62. [70]

    Random Forests , author =. Mach. Learn. , volume = 45, number = 1, pages =

  63. [71]

    2016 , journal =

    Loda: Lightweight on-line detector of anomalies , author =. 2016 , journal =

  64. [72]

    2012 , journal=

    Histogram-based outlier score (hbos): A fast unsupervised anomaly detection algorithm , author=. 2012 , journal=

  65. [73]

    2008 , booktitle =

    Angle-based outlier detection in high-dimensional data , author =. 2008 , booktitle =

  66. [74]

    2002 , booktitle =

    Enhancing Effectiveness of Outlier Detections for Low Density Patterns , author =. 2002 , booktitle =

  67. [75]

    2020 , journal =

    Generative Adversarial Active Learning for Unsupervised Outlier Detection , author =. 2020 , journal =

  68. [76]

    Proceedings of the

    XGBoost: A Scalable Tree Boosting System , author =. Proceedings of the

  69. [77]

    Deep Learning , author =

  70. [78]

    An Introduction to Statistical Learning--with Applications in R , author =

  71. [79]

    Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , year =

    Why do tree-based models still outperform deep learning on typical tabular data? , author =. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) , year =

  72. [80]

    2022 , journal =

    Tabular data: Deep learning is not all you need , author =. 2022 , journal =

  73. [81]

    2022 , booktitle =

    Adam Goodge and Bryan Hooi and See. 2022 , booktitle =

  74. [82]

    2022 , booktitle =

    Hyperparameter Sensitivity in Deep Outlier Detection: Analysis and a Scalable Hyper-Ensemble Solution , author =. 2022 , booktitle =

  75. [83]

    2024 , booktitle =

    On Diffusion Modeling for Anomaly Detection , author =. 2024 , booktitle =

  76. [84]

    Jensen , title =

    David Campos and Tung Kieu and Chenjuan Guo and Feiteng Huang and Kai Zheng and Bin Yang and Christian S. Jensen , title =. Proc

  77. [85]

    Jensen and Bin Yang , title =

    Buang Zhang and Tung Kieu and Xiangfei Qiu and Chenjuan Guo and Jilin Hu and Aoying Zhou and Christian S. Jensen and Bin Yang , title =. Proceedings of the

  78. [86]

    2025 , journal =

    Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching , author =. 2025 , journal =

  79. [87]

    M. M. Breunig, H.-P. Kriegel, R. T. Ng, and J. Sander, ``LOF: Identifying density-based local outliers,'' in Proc. ACM SIGMOD, 2000, pp. 93--104

  80. [88]

    Z. Li, Y. Zhao, X. Hu, N. Botta, C. Ionescu, and G. Chen, ``COPOD: Copula-based outlier detection,'' in Proc. IEEE ICDM, 2020, pp. 1118--1123

  81. [89]

    F. T. Liu, K. M. Ting, and Z.-H. Zhou, ``Isolation forest,'' in Proc. IEEE ICDM, 2008, pp. 413--422

  82. [90]

    Y. Zhao, Z. Nasrullah, and Z. Li, ``PyOD: A Python toolbox for scalable outlier detection,'' Journal of Machine Learning Research, vol. 20, no. 96, pp. 1--7, 2019

  83. [91]

    M. Q. Ma, Y. Zhao, X. Zhang, and L. Akoglu, ``A large-scale study on unsupervised outlier model selection: Do internal strategies suffice?'' arXiv preprint arXiv:2104.01422, 2021

  84. [92]

    S. Han, X. Hu, H. Huang, M. Jiang, and Y. Zhao, ``ADBench: Anomaly detection benchmark,'' in Proc. NeurIPS Datasets and Benchmarks Track, 2022

  85. [93]

    Y. Zhao, R. A. Rossi, and L. Akoglu, ``Automating outlier detection via meta-learning,'' in Proc. NeurIPS, 2021, pp. 4607--4619

  86. [94]

    Y. Zhao, S. Zhang, and L. Akoglu, ``Toward unsupervised outlier model selection,'' in Proc. NeurIPS, 2022, pp. 2205--2218

  87. [95]

    C. C. Aggarwal and S. Sathe, ``Theoretical foundations and algorithms for outlier ensembles,'' ACM SIGKDD Explorations Newsletter, vol. 17, no. 1, pp. 24--47, 2015

  88. [96]

    Goldstein and A

    M. Goldstein and A. Dengel, ``Histogram-based outlier score (HBOS): A fast unsupervised anomaly detection algorithm,'' in Proc. KI, 2012, pp. 59--63

  89. [98]

    H. O. Marques, R. J. Campello, J. Sander, and A. Zimek, ``Internal evaluation of unsupervised outlier detection,'' ACM Trans. on Knowledge Discovery from Data, vol. 14, no. 4, pp. 1--42, 2020

  90. [99]

    Y. Li, Z. Chen, D. Zha, K. Zhou, H. Jin, H. Chen, and X. Hu, ``AutoOD: Neural architecture search for outlier detection,'' in Proc. IEEE ICDE, 2021, pp. 2117--2122

  91. [100]

    G. O. Campos, A. Zimek, and W. Meira Jr, ``An unsupervised boosting strategy for outlier detection ensembles,'' in Proc. PAKDD, 2018, pp. 564--576

  92. [101]

    Y. Zhao, Z. Nasrullah, M. K. Hryniewicki, and Z. Li, ``LSCP: Locally selective combination in parallel outlier ensembles,'' in Proc. SIAM SDM, 2019, pp. 585--593

  93. [102]

    Rayana and L

    S. Rayana and L. Akoglu, ``Less is more: Building selective anomaly ensembles,'' ACM Trans. on Knowledge Discovery from Data, vol. 10, no. 4, pp. 1--33, 2016

  94. [103]

    G. L. Nemhauser, L. A. Wolsey, and M. L. Fisher, ``An analysis of approximations for maximizing submodular set functions---I,'' Mathematical Programming, vol. 14, no. 1, pp. 265--294, 1978

  95. [104]

    Minoux, ``Accelerated greedy algorithms for maximizing submodular set functions,'' in Optimization Techniques

    M. Minoux, ``Accelerated greedy algorithms for maximizing submodular set functions,'' in Optimization Techniques. Springer, 1978, pp. 234--243

  96. [105]

    Lin and J

    H. Lin and J. Bilmes, ``A class of submodular functions for document summarization,'' in Proc. ACL, 2011, pp. 510--520

  97. [106]

    Krause, A

    A. Krause, A. Singh, and C. Guestrin, ``Near-optimal sensor placements in Gaussian processes: Theory, efficient algorithms and empirical studies,'' Journal of Machine Learning Research, vol. 9, pp. 235--284, 2008

  98. [107]

    Chen and A

    Y. Chen and A. Krause, ``Near-optimal batch mode active learning and adaptive submodular optimization,'' in Proc. ICML, 2013, pp. 160--168

  99. [108]

    Rayana, ``ODDS Library,'' 2016

    S. Rayana, ``ODDS Library,'' 2016. [Online]. Available: http://odds.cs.stonybrook.edu

  100. [109]

    G. O. Campos, A. Zimek, J. Sander, R. J. G. B. Campello, B. Micenková, E. Schubert, I. Assent, and M. E. Houle, ``On the evaluation of unsupervised outlier detection,'' Data Mining and Knowledge Discovery, vol. 30, no. 4, pp. 891--927, 2016

  101. [110]

    G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, ``LightGBM: A highly efficient gradient boosting decision tree,'' in Proc. NIPS, 2017, pp. 3146--3154

  102. [111]

    L. Ruff, R. Vandermeulen, N. Goernitz, L. Deecke, S. A. Siddiqui, A. Binder, E. Müller, and M. Kloft, ``Deep one-class classification,'' in Proc. ICML, 2018, pp. 4393--4402

  103. [112]

    J. Xu, H. Wu, J. Wang, and M. Long, ``Anomaly transformer: Time series anomaly detection with association discrepancy,'' in Proc. ICLR, 2022

  104. [113]

    Aggarwal, C. C. and Sathe, S. Theoretical foundations and algorithms for outlier ensembles. SIGKDD Explor. , 17: 0 24--47, 2015

  105. [114]

    A., Razzak, I., and Santosh, K

    Aryal, S., Baniya, A. A., Razzak, I., and Santosh, K. SPAD+: an improved probabilistic anomaly detector based on one-dimensional histograms. In Proceedings of the International Joint Conference on Neural Networks (IJCNN), pp.\ 1--7, 2021

  106. [115]

    Random forests

    Breiman, L. Random forests. Mach. Learn., 45 0 (1): 0 5--32, 2001

  107. [116]

    M., Kriegel, H., Ng, R

    Breunig, M. M., Kriegel, H., Ng, R. T., and Sander, J. LOF: identifying density-based local outliers. In Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD) , pp.\ 93--104, 2000

  108. [117]

    Campos, D., Kieu, T., Guo, C., Huang, F., Zheng, K., Yang, B., and Jensen, C. S. Unsupervised time series outlier detection with diversity-driven convolutional ensembles. Proc. VLDB Endow. , 15 0 (3): 0 611--623, 2021

  109. [118]

    O., Zimek, A., Sander, J., Campello, R

    Campos, G. O., Zimek, A., Sander, J., Campello, R. J. G. B., Micenkov \' a , B., Schubert, E., Assent, I., and Houle, M. E. On the evaluation of unsupervised outlier detection: Measures, datasets, and an empirical study. Data Min. Knowl. Discov., 30: 0 891--927, 2016

  110. [119]

    Anomaly detection: A survey

    Chandola, V., Banerjee, A., and Kumar, V. Anomaly detection: A survey. ACM Comput. Surv. , 41 0 (3): 0 15:1--15:58, 2009

  111. [120]

    and Friguet, C

    Chapel, L. and Friguet, C. Anomaly detection with score functions based on the reconstruction error of the kernel PCA . In Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD), pp.\ 227--241, 2014

  112. [121]

    Chehreghani, M. H. K-nearest neighbor search and outlier detection via minimax distances. In Proceedings of the SIAM International Conference on Data Mining (SDM) , pp.\ 405--413, 2016

  113. [122]

    C., and Turaga, D

    Chen, J., Sathe, S., Aggarwal, C. C., and Turaga, D. S. Outlier detection with autoencoder ensembles. In Proceedings of the SIAM International Conference on Data Mining (SDM) , pp.\ 90--98, 2017

  114. [123]

    and Guestrin, C

    Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD) , pp.\ 785--794, 2016

  115. [124]

    Outlier detection ensemble with embedded feature selection

    Cheng, L., Wang, Y., Liu, X., and Li, B. Outlier detection ensemble with embedded feature selection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , pp.\ 3503--3512, 2020

  116. [125]

    Hyperparameter sensitivity in deep outlier detection: Analysis and a scalable hyper-ensemble solution

    Ding, X., Zhao, L., and Akoglu, L. Hyperparameter sensitivity in deep outlier detection: Analysis and a scalable hyper-ensemble solution. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2022

  117. [126]

    T., Blum, M., and Hutter, F

    Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., and Hutter, F. Efficient and robust automated machine learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pp.\ 2962--2970, 2015

  118. [127]

    Extremely randomized trees

    Geurts, P., Ernst, D., and Wehenkel, L. Extremely randomized trees. Mach. Learn., 63 0 (1): 0 3--42, 2006

  119. [128]

    How to evaluate the quality of unsupervised anomaly detection algorithms? CoRR, abs/1607.01152, 2016

    Goix, N. How to evaluate the quality of unsupervised anomaly detection algorithms? CoRR, abs/1607.01152, 2016

  120. [129]

    and Dengel, A

    Goldstein, M. and Dengel, A. Histogram-based outlier score (hbos): A fast unsupervised anomaly detection algorithm. KI-2012: poster and demo track, 1: 0 59--63, 2012

  121. [130]

    Deep Learning

    Goodfellow, I., Bengio, Y., and Courville, A. Deep Learning. MIT Press, 2016. ISBN 978-0-2620-3561-3

  122. [131]

    Goodge, A., Hooi, B., Ng, S., and Ng, W. S. Robustness of autoencoders for anomaly detection under adversarial impact. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pp.\ 1244--1250, 2020

  123. [132]

    Goodge, A., Hooi, B., Ng, S., and Ng, W. S. LUNAR: unifying local outlier detection methods via graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , pp.\ 6737--6745, 2022

  124. [133]

    Why do tree-based models still outperform deep learning on typical tabular data? In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2022

    Grinsztajn, L., Oyallon, E., and Varoquaux, G. Why do tree-based models still outperform deep learning on typical tabular data? In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2022

  125. [134]

    Adbench: Anomaly detection benchmark

    Han, S., Hu, X., Huang, H., Jiang, M., and Zhao, Y. Adbench: Anomaly detection benchmark. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2022

  126. [135]

    GAN ensemble for anomaly detection

    Han, X., Chen, X., and Liu, L. GAN ensemble for anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , pp.\ 4090--4097, 2021

  127. [136]

    a ki, V., K \

    Hautam \" a ki, V., K \" a rkk \" a inen, I., and Fr \" a nti, P. Outlier detection using k-nearest neighbour graph. In Proceedings of the International Conference on Pattern Recognition (ICPR), pp.\ 430--433, 2004

  128. [137]

    M., Antoniou, A., Micaelli, P., and Storkey, A

    Hospedales, T. M., Antoniou, A., Micaelli, P., and Storkey, A. Meta-learning in neural networks: A survey. IEEE Trans. Pattern Anal. Mach. Intell. , 44 0 (9): 0 5149--5169, 2022

  129. [138]

    An Introduction to Statistical Learning--with Applications in R

    James, G., Witten, D., Hastie, T., and Tibshirani, R. An Introduction to Statistical Learning--with Applications in R. Springer, 2013. ISBN 978-1-4614-7137-0

  130. [139]

    Lightgbm: A highly efficient gradient boosting decision tree

    Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T. Lightgbm: A highly efficient gradient boosting decision tree. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pp.\ 3146--3154, 2017

  131. [140]

    Kieu, T., Yang, B., Guo, C., and Jensen, C. S. Outlier detection for time series with recurrent autoencoder ensembles. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pp.\ 2725--2732, 2019

  132. [141]

    H., and Hong, C

    Kim, H., Lee, C. H., and Hong, C. Transformer for point anomaly detection. In Proceedings of the ACM International Conference on Information and Knowledge Management (CIKM) , pp.\ 1080--1088, 2024

  133. [142]

    Angle-based outlier detection in high-dimensional data

    Kriegel, H., Schubert, M., and Zimek, A. Angle-based outlier detection in high-dimensional data. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD) , pp.\ 444--452, 2008

  134. [143]

    and Taskar, B

    Kulesza, A. and Taskar, B. Determinantal point processes for machine learning. Found. Trends Mach. Learn., 5 0 (2-3): 0 123--286, 2012

  135. [144]

    and Kumar, V

    Lazarevic, A. and Kumar, V. Feature bagging for outlier detection. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD) , pp.\ 157--166, 2005

  136. [145]

    COPOD: copula-based outlier detection

    Li, Z., Zhao, Y., Botta, N., Ionescu, C., and Hu, X. COPOD: copula-based outlier detection. In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 1118--1123, 2020

  137. [146]

    Li, Z., Zhao, Y., Hu, X., Botta, N., Ionescu, C., and Chen, G. H. ECOD: unsupervised outlier detection using empirical cumulative distribution functions. IEEE Trans. Knowl. Data Eng. , 35 0 (12): 0 12181--12193, 2023

  138. [147]

    M., van Stein, N., and van Leeuwen, M

    Li, Z., Huang, Q., Zhu, Y., Yang, L., Amiri, M. M., van Stein, N., and van Leeuwen, M. Scalable, explainable and provably robust anomaly detection with one-step flow matching. CoRR, abs/2510.18328, 2025

  139. [148]

    K., Loo, Y., Tran, N., Cheung, N., Roig, G., and Elovici, Y

    Lim, S. K., Loo, Y., Tran, N., Cheung, N., Roig, G., and Elovici, Y. DOPING: generative data augmentation for unsupervised anomaly detection with GAN . In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 1122--1127, 2018

  140. [149]

    T., Ting, K

    Liu, F. T., Ting, K. M., and Zhou, Z. Isolation forest. In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 413--422, 2008

  141. [150]

    Generative adversarial active learning for unsupervised outlier detection

    Liu, Y., Li, Z., Zhou, C., Jiang, Y., Sun, J., Wang, M., and He, X. Generative adversarial active learning for unsupervised outlier detection. IEEE Trans. Knowl. Data Eng. , 32: 0 1517--1528, 2020

  142. [151]

    On diffusion modeling for anomaly detection

    Livernoche, V., Jain, V., Hezaveh, Y., and Ravanbakhsh, S. On diffusion modeling for anomaly detection. In Proceedings of the International Conference on Learning Representations (ICLR), 2024

  143. [152]

    O., Campello, R

    Marques, H. O., Campello, R. J. G. B., Sander, J., and Zimek, A. Internal evaluation of unsupervised outlier detection. ACM Trans. Knowl. Discov. Data , 14 0 (4): 0 47:1--47:42, 2020

  144. [153]

    L., Wolsey, L

    Nemhauser, G. L., Wolsey, L. A., and Fisher, M. L. An analysis of approximations for maximizing submodular set functions- I . Math. Program., 14 0 (1): 0 265--294, 1978

  145. [154]

    E., and Slonim, D

    Noto, K., Brodley, C. E., and Slonim, D. K. Anomaly detection using an ensemble of feature models. In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 953--958, 2010

  146. [155]

    Z., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S

    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., K \" o pf, A., Yang, E. Z., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: A...

  147. [156]

    Scikit-learn: Machine learning in python

    Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in python. J. Mach. Lea...

  148. [157]

    Loda: Lightweight on-line detector of anomalies

    Pevn \' y , T. Loda: Lightweight on-line detector of anomalies. Mach. Learn., 102 0 (2): 0 275--304, 2016

  149. [158]

    Sequential ensemble learning for outlier detection: A bias-variance perspective

    Rayana, S., Zhong, W., and Akoglu, L. Sequential ensemble learning for outlier detection: A bias-variance perspective. In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 1167--1172, 2016

  150. [159]

    o rnitz, N., Deecke, L., Siddiqui, S. A., Vandermeulen, R. A., Binder, A., M \

    Ruff, L., G \" o rnitz, N., Deecke, L., Siddiqui, S. A., Vandermeulen, R. A., Binder, A., M \" u ller, E., and Kloft, M. Deep one-class classification. In Proceedings of the International Conference on Machine Learning (ICML), pp.\ 4390--4399, 2018

  151. [160]

    R., Vandermeulen, R

    Ruff, L., Kauffmann, J. R., Vandermeulen, R. A., Montavon, G., Samek, W., Kloft, M., Dietterich, T. G., and M \" u ller, K. A unifying review of deep and shallow anomaly detection. Proc. IEEE , 109 0 (5): 0 756--795, 2021

  152. [161]

    C., Smola, A

    Sch \" o lkopf, B., Williamson, R. C., Smola, A. J., Shawe - Taylor, J., and Platt, J. C. Support vector method for novelty detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pp.\ 582--588, 1999

  153. [162]

    Estimating the dimension of a model

    Schwarz, G. Estimating the dimension of a model. Ann. Stat., 6 0 (2): 0 461--464, 1978

  154. [163]

    and Armon, A

    Shwartz - Ziv, R. and Armon, A. Tabular data: Deep learning is not all you need. Inf. Fusion, 81: 0 84--90, 2022

  155. [164]

    Snoek, J., Larochelle, H., and Adams, R. P. Practical bayesian optimization of machine learning algorithms. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pp.\ 2960--2968, 2012

  156. [165]

    Cross-validatory choice and assessment of statistical predictions

    Stone, M. Cross-validatory choice and assessment of statistical predictions. J. R. Stat. Soc., B: Stat. Methodol., 36 0 (1): 0 111--147, 1974

  157. [166]

    W., and Cheung, D

    Tang, J., Chen, Z., Fu, A. W., and Cheung, D. W. Enhancing effectiveness of outlier detections for low density patterns. In Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp.\ 535--548, 2002

  158. [167]

    Tax, D. M. J. and Duin, R. P. W. Support vector data description. Mach. Learn., 54 0 (1): 0 45--66, 2004

  159. [168]

    and Hinton, G

    van der Maaten, L. and Hinton, G. Visualizing data using t-sne. J. Mach. Learn. Res., 9 0 (86): 0 2579--2605, 2008

  160. [169]

    Statistical Learning Theory

    Vapnik, V. Statistical Learning Theory. Wiley, 1998. ISBN 978-0-471-03003-4

  161. [170]

    and Drissi, Y

    Vilalta, R. and Drissi, Y. A perspective view and survey of meta-learning. Artif. Intell. Rev., 18 0 (2): 0 77--95, 2002

  162. [171]

    Unsupervised anomaly detection via variational auto-encoder for seasonal kpis in web applications

    Xu, H., Chen, W., Zhao, N., Li, Z., Bu, J., Li, Z., Liu, Y., Zhao, Y., Pei, D., Feng, Y., Chen, J., Wang, Z., and Qiao, H. Unsupervised anomaly detection via variational auto-encoder for seasonal kpis in web applications. In Proceedings of the International Conference on World...

  163. [172]

    S., and Yang, B

    Zhang, B., Kieu, T., Qiu, X., Guo, C., Hu, J., Zhou, A., Jensen, C. S., and Yang, B. An encode-then-decompose approach to unsupervised time series anomaly detection on contaminated training data. In Proceedings of the IEEE International Conference on Data Engineering (ICDE) , 2026

  164. [173]

    and Hryniewicki, M

    Zhao, Y. and Hryniewicki, M. K. XGBOD: improving supervised outlier detection with unsupervised representation learning. In Proceedings of the International Joint Conference on Neural Networks (IJCNN), pp.\ 1--8, 2018

  165. [174]

    K., and Li, Z

    Zhao, Y., Nasrullah, Z., Hryniewicki, M. K., and Li, Z. LSCP: locally selective combination in parallel outlier ensembles. In Proceedings of the SIAM International Conference on Data Mining (SDM) , pp.\ 585--593, 2019 a

  166. [175]

    Pyod: A python toolbox for scalable outlier detection

    Zhao, Y., Nasrullah, Z., and Li, Z. Pyod: A python toolbox for scalable outlier detection. J. Mach. Learn. Res., 20: 0 96:1--96:7, 2019 b

  167. [176]

    A., and Akoglu, L

    Zhao, Y., Rossi, R. A., and Akoglu, L. Automating outlier detection via meta-learning. CoRR, abs/2009.10606, 2020

  168. [177]

    A., and Akoglu, L

    Zhao, Y., Rossi, R. A., and Akoglu, L. Automatic unsupervised outlier model selection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), pp.\ 4489--4502, 2021

  169. [178]

    Toward unsupervised outlier model selection

    Zhao, Y., Zhang, S., and Akoglu, L. Toward unsupervised outlier model selection. In Proceedings of the IEEE International Conference on Data Mining (ICDM) , pp.\ 773--782, 2022

  170. [179]

    and Paffenroth, R

    Zhou, C. and Paffenroth, R. C. Anomaly detection with robust deep autoencoders. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (SIGKDD) , pp.\ 665--674, 2017

  171. [180]

    Zimek, A., Campello, R. J. G. B., and Sander, J. Ensembles for unsupervised outlier detection: challenges and research questions a position paper. SIGKDD Explor. , 15 0 (1): 0 11--22, 2013

  172. [181]

    R., Cheng, W., Lumezanu, C., Cho, D., and Chen, H

    Zong, B., Song, Q., Min, M. R., Cheng, W., Lumezanu, C., Cho, D., and Chen, H. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In Proceedings of the International Conference on Learning Representations (ICLR), 2018

Pith tools

Reviewed May 20, 2026 · model on record in the stance chip above.