Pith. sign in

REVIEW 5 cited by

Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05244 v1 pith:S56DRMWI submitted 2024-12-06 cs.LG cs.AI

Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization

classification cs.LG cs.AI
keywords modelstimeseriesforecastingbestbettercoefficientscomplex
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

How to best develop foundational models for time series forecasting remains an important open question. Tokenization is a crucial consideration in this effort: what is an effective discrete vocabulary for a real-valued sequential input? To address this question, we develop WaveToken, a wavelet-based tokenizer that allows models to learn complex representations directly in the space of time-localized frequencies. Our method first scales and decomposes the input time series, then thresholds and quantizes the wavelet coefficients, and finally pre-trains an autoregressive model to forecast coefficients for the forecast horizon. By decomposing coarse and fine structures in the inputs, wavelets provide an eloquent and compact language for time series forecasting that simplifies learning. Empirical results on a comprehensive benchmark, including 42 datasets for both in-domain and zero-shot settings, show that WaveToken: i) provides better accuracy than recently proposed foundation models for forecasting while using a much smaller vocabulary (1024 tokens), and performs on par or better than modern deep learning models trained specifically on each dataset; and ii) exhibits superior generalization capabilities, achieving the best average rank across all datasets for three complementary metrics. In addition, we show that our method can easily capture complex temporal patterns of practical relevance that are challenging for other recent pre-trained models, including trends, sparse spikes, and non-stationary time series with varying frequencies evolving over time.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Data-Driven Forecasting of three-Component Seismograms Using Transformer Architectures

    astro-ph.IM 2026-06 unverdicted novelty 6.0

    SeismoGPT is a transformer autoregressive model achieving median normalized cross-correlation above 0.93 when forecasting synthetic three-component seismograms up to 240 s ahead from P- and S-wave context.

  2. Data-Driven Forecasting of three-Component Seismograms Using Transformer Architectures

    astro-ph.IM 2026-06 conditional novelty 6.0

    A causal transformer trained on synthetic seismograms autoregressively continues three-component waveforms past the S-wave with median NCC of 0.93 or higher across controlled geometries.

  3. Spectral Vision Transformer for Efficient Tokenization with Limited Data

    cs.CV 2026-05 unverdicted novelty 6.0

    A spectral vision transformer achieves equitable or superior performance with fewer parameters than standard ViTs, CNNs, and other models by using spectral projections for tokenization in limited-data medical imaging.

  4. What If We Let Forecasting Forget? A Sparse Bottleneck for Cross-Variable Dependencies

    cs.LG 2026-05 unverdicted novelty 6.0

    MS-FLOW uses a capacity-limited sparse routing mechanism to model only critical inter-variable dependencies in time series data, achieving state-of-the-art accuracy on 12 benchmarks with fewer but more reliable connections.

  5. JEPA for AI-Native 6G: Predictive Representations and Open Challenges

    cs.LG 2026-07 conditional novelty 5.5

    Wireless-aware JEPA pretraining with an auxiliary future beam-energy target improves label-efficient beam ranking and OOD robustness on synthetic mmWave data relative to generic JEPA, MAE, SimCLR and supervised scratch.