Pith. sign in

REVIEW 2 cited by

TimeSeriesBench: An Industrial-Grade Benchmark for Time Series Anomaly Detection Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10802 v3 pith:4KBY6UGZ submitted 2024-02-16 cs.LG

classification cs.LG
keywords seriestimealgorithmsanomalydetectionmodelsperformancetimeseriesbench
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Time series anomaly detection (TSAD) has gained significant attention due to its real-world applications to improve the stability of modern software systems. However, there is no effective way to verify whether they can meet the requirements for real-world deployment. Firstly, current algorithms typically train a specific model for each time series. Maintaining such many models is impractical in a large-scale system with tens of thousands of curves. The performance of using merely one unified model to detect anomalies remains unknown. Secondly, most TSAD models are trained on the historical part of a time series and are tested on its future segment. In distributed systems, however, there are frequent system deployments and upgrades, with new, previously unseen time series emerging daily. The performance of testing newly incoming unseen time series on current TSAD algorithms remains unknown. Lastly, the assumptions of the evaluation metrics in existing benchmarks are far from practical demands. To solve the above-mentioned problems, we propose an industrial-grade benchmark TimeSeriesBench. We assess the performance of existing algorithms across more than 168 evaluation settings and provide comprehensive analysis for the future design of anomaly detection algorithms. An industrial dataset is also released along with TimeSeriesBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models

    cs.LG 2025-01 reject novelty 7.0 of 10

    ARGOS uses LLM agents to generate explainable, reproducible anomaly detection rules and fuses them with a base detector, reporting higher F1 than deep-learning and LLM baselines on KPI, Yahoo, and a Microsoft internal...

  2. TAB: Unified Benchmarking of Time Series Anomaly Detection Methods

    cs.LG 2025-06 conditional novelty 5.0 of 10

    TAB is a new time series anomaly detection benchmark that unifies datasets, methods, and evaluation protocols, and its results show classical methods remain highly competitive against deep learning and foundation models.

Pith tools