Pith. sign in

REVIEW 12 cited by

Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.03955 v8 pith:MERUZSDQ submitted 2024-01-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords forecastingmodelsfew-shottimezerohttpshuggingfacemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large pre-trained models excel in zero/few-shot learning for language and vision tasks but face challenges in multivariate time series (TS) forecasting due to diverse data characteristics. Consequently, recent research efforts have focused on developing pre-trained TS forecasting models. These models, whether built from scratch or adapted from large language models (LLMs), excel in zero/few-shot forecasting tasks. However, they are limited by slow performance, high computational demands, and neglect of cross-channel and exogenous correlations. To address this, we introduce Tiny Time Mixers (TTM), a compact model (starting from 1M parameters) with effective transfer learning capabilities, trained exclusively on public TS datasets. TTM, based on the light-weight TSMixer architecture, incorporates innovations like adaptive patching, diverse resolution sampling, and resolution prefix tuning to handle pre-training on varied dataset resolutions with minimal model capacity. Additionally, it employs multi-level modeling to capture channel correlations and infuse exogenous signals during fine-tuning. TTM outperforms existing popular benchmarks in zero/few-shot forecasting by (4-40%), while reducing computational requirements significantly. Moreover, TTMs are lightweight and can be executed even on CPU-only machines, enhancing usability and fostering wider adoption in resource-constrained environments. The model weights for reproducibility and research use are available at https://huggingface.co/ibm/ttm-research-r2/, while enterprise-use weights under the Apache license can be accessed as follows: the initial TTM-Q variant at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r1, and the latest variants (TTM-B, TTM-E, TTM-A) weights are available at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It Drives

    cs.NI 2026-08 conditional novelty 6.0 of 10

    One pretrained telemetry model, CENTILE, improves both HPC backfilling and ISP capacity provisioning decisions under replay, with zero-shot transfer across months and domains.

  2. Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

    stat.AP 2026-07 conditional novelty 6.0 of 10

    Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.

  3. Time Series Foundation Models for Multivariate Financial Time Series Forecasting

    q-fin.GN 2025-07 reject novelty 6.0 of 10

    Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...

  4. EPBench: A Benchmark for Short-term Earthquake Prediction with Neural Networks

    physics.geo-ph 2025-05 conditional novelty 6.0 of 10

    A new global regional-scale benchmark provides data, splits, evaluation metrics, and neural network plus ETAS baselines for short-term earthquake prediction.

  5. MoTime: A Dataset Suite for Multimodal Time Series Forecasting

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.

  6. Investigating Compositional Reasoning in Time Series Foundation Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    On a benchmark where models train on Fourier components and test on their sums, patch-based Transformers and residual MLP architectures show the strongest compositional generalization, while most standard transformers...

  7. RhyMix: A Lightweight Adaptive Multi-Rhythm Network for Long-Term Time Series Forecasting

    cs.LG 2026-07 conditional novelty 5.5 of 10

    RhyMix reaches state-of-the-art long-term multivariate forecasting on 10 of 12 public benchmarks with a ~40K-parameter dual-path adaptive architecture of linear complexity.

  8. Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting

    stat.ML 2026-07 reject novelty 5.0 of 10

    A two-stage forecaster (SPA trend extraction + LoRA-fine-tuned residual Transformer) that the paper claims beats prior models by 6.56% MASE, though the claim is not robust to its own extended baseline tables.

  9. Causal Graph Fuzzy LLMs: A First Introduction and Applications in Time Series Forecasting

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CGF-LLM combines fuzzy time series and PCMCI causal graphs into text input for fine-tuned GPT-2, reporting improved one-step-ahead forecast NRMSE and a reduction in token count on four datasets.

  10. Beyond Data Scarcity: A Frequency-Driven Framework for Zero-Shot Forecasting

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Freq-Synth trains zero-shot forecasters on harmonic sine waves generated from the target sampling rate and beats real-data training on 6 of 8 benchmarks.

  11. Causal Time-Series Synchronization for Multi-Dimensional Forecasting

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Aligning cause-effect pairs by their estimated Granger lag improves channel-dependent forecasting accuracy and transfer learning on synthetic time-series data.

  12. A Survey of AIOps in the Era of Large Language Models

    cs.SE 2025-06 conditional novelty 3.0 of 10

    A systematic survey that categorizes LLM-based AIOps research into four dimensions: data sources, tasks, methods, and evaluation, claiming to be the first comprehensive such overview.

Pith tools