Pith. sign in

REVIEW 2 cited by

Scaling-laws for Large Time-series Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13867 v2 pith:BDGTJBZ3 submitted 2024-05-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelstimelargeserieslanguagellmsscalingtraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling laws for large language models (LLMs) have provided useful guidance in training ever larger models for predictable performance gains. Time series forecasting shares a similar sequential structure to language, and is amenable to large-scale transformer architectures. Here we show that foundational decoder-only time series transformer models exhibit analogous scaling-behavior to LLMs, with architectural details (aspect ratio and number of heads) having a minimal effect over broad ranges. We assemble a large corpus of heterogenous time series data on which to train, and establish for the first time power-law scaling with parameter count, dataset size, and training compute, spanning five orders of magnitude.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Time-series foundation models are unconditionally better than classical methods on 15/30 datasets, lose early on 6, and a n_train<700 + seasonality rule resolves 10 deployment decisions without training.

  2. Learning Spatio-Temporal Foundation Models from Pure Synthetic Data

    cs.LG 2026-06 conditional novelty 6.0 of 10

    A spatio-temporal foundation model pre-trained exclusively on synthetic stochastic graph dynamics outperforms real-data-pretrained STFMs in zero-shot traffic forecasting, according to the paper's benchmarks.

Pith tools