REVIEW 3 cited by
Towards a Rigorous Evaluation of Time-series Anomaly Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In recent years, proposed studies on time-series anomaly detection (TAD) report high F1 scores on benchmark TAD datasets, giving the impression of clear improvements in TAD. However, most studies apply a peculiar evaluation protocol called point adjustment (PA) before scoring. In this paper, we theoretically and experimentally reveal that the PA protocol has a great possibility of overestimating the detection performance; that is, even a random anomaly score can easily turn into a state-of-the-art TAD method. Therefore, the comparison of TAD methods after applying the PA protocol can lead to misguided rankings. Furthermore, we question the potential of existing TAD methods by showing that an untrained model obtains comparable detection performance to the existing methods even when PA is forbidden. Based on our findings, we propose a new baseline and an evaluation protocol. We expect that our study will help a rigorous evaluation of TAD and lead to further improvement in future researches.
Forward citations
Cited by 3 Pith papers
-
Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems
A VaDE-based latent clustering detector that drops reconstruction wins a fair, difficulty-stratified anomaly detection protocol on three CPS datasets, with margins tracking dataset multimodality.
-
What Do Machine Learning Researchers Mean by "Reproducible"?
A survey-based taxonomy that splits reproducibility in AI/ML into eight rigor aspects (repeatability, reproducibility, replicability, adaptability, model selection, label/data quality, meta/incentive, maintainability)...
-
Evaluation of Stress Detection as Time Series Events -- A Novel Window-Based F1-Metric
A window-based F1 metric that rewards detections within a time tolerance reveals statistically significant stress-event prediction where pointwise F1 scores are all zero.
Discussion (0). Continue with ORCID to comment.