Pith. sign in

Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress. Existing datasets suffer from issues ranging from small-scale, low-frequency, pre-training data contamination in unimodal designs to the temporal and description leakage prevalent in early multimodal designs. To address this, we formalize the core principles of high-fidelity benchmarking, focusing on data sourcing integrity, leak-free design, and structural clarity. We introduce Fidel-TS, a new large-scale benchmark built from these principles. Our experiments reveal the limitations of prior benchmarks and the potential discrepancies in model evaluation, providing new insights into multiple existing unimodal and multimodal forecasting models and LLMs across various evaluation tasks.

fields

cs.LG 3

years

2026 2 2025 1

representative citing papers

Toto 2.0: Time Series Forecasting Enters the Scaling Era

cs.LG · 2026-05-19 · unverdicted · novelty 5.0 · 2 refs

Time series foundation models scale under a single training recipe, with forecast quality improving from 4M to 2.5B parameters and new SOTA results on BOOM, GIFT-Eval, and TIME benchmarks.

citing papers explorer

Showing 3 of 3 citing papers.