REVIEW 2 cited by
TimeSeriesBench: An Industrial-Grade Benchmark for Time Series Anomaly Detection Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Time series anomaly detection (TSAD) has gained significant attention due to its real-world applications to improve the stability of modern software systems. However, there is no effective way to verify whether they can meet the requirements for real-world deployment. Firstly, current algorithms typically train a specific model for each time series. Maintaining such many models is impractical in a large-scale system with tens of thousands of curves. The performance of using merely one unified model to detect anomalies remains unknown. Secondly, most TSAD models are trained on the historical part of a time series and are tested on its future segment. In distributed systems, however, there are frequent system deployments and upgrades, with new, previously unseen time series emerging daily. The performance of testing newly incoming unseen time series on current TSAD algorithms remains unknown. Lastly, the assumptions of the evaluation metrics in existing benchmarks are far from practical demands. To solve the above-mentioned problems, we propose an industrial-grade benchmark TimeSeriesBench. We assess the performance of existing algorithms across more than 168 evaluation settings and provide comprehensive analysis for the future design of anomaly detection algorithms. An industrial dataset is also released along with TimeSeriesBench.
Forward citations
Cited by 2 Pith papers
-
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
ARGOS uses LLM agents to generate explainable, reproducible anomaly detection rules and fuses them with a base detector, reporting higher F1 than deep-learning and LLM baselines on KPI, Yahoo, and a Microsoft internal...
-
TAB: Unified Benchmarking of Time Series Anomaly Detection Methods
TAB is a new time series anomaly detection benchmark that unifies datasets, methods, and evaluation protocols, and its results show classical methods remain highly competitive against deep learning and foundation models.
Discussion (0). Continue with ORCID to comment.