Pith. sign in

REVIEW 26 cited by

Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16040 v4 pith:XGRNUMWN submitted 2024-09-24 cs.LG cs.AI

Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts

classification cs.LG cs.AI
keywords modelstimeforecastingseriestime-moefoundationmodeladvancements
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep learning for time series forecasting has seen significant advancements over the past decades. However, despite the success of large-scale pre-training in language and vision domains, pre-trained time series models remain limited in scale and operate at a high cost, hindering the development of larger capable forecasting models in real-world applications. In response, we introduce Time-MoE, a scalable and unified architecture designed to pre-train larger, more capable forecasting foundation models while reducing inference costs. By leveraging a sparse mixture-of-experts (MoE) design, Time-MoE enhances computational efficiency by activating only a subset of networks for each prediction, reducing computational load while maintaining high model capacity. This allows Time-MoE to scale effectively without a corresponding increase in inference costs. Time-MoE comprises a family of decoder-only transformer models that operate in an auto-regressive manner and support flexible forecasting horizons with varying input context lengths. We pre-trained these models on our newly introduced large-scale data Time-300B, which spans over 9 domains and encompassing over 300 billion time points. For the first time, we scaled a time series foundation model up to 2.4 billion parameters, achieving significantly improved forecasting precision. Our results validate the applicability of scaling laws for training tokens and model size in the context of time series forecasting. Compared to dense models with the same number of activated parameters or equivalent computation budgets, our models consistently outperform them by large margin. These advancements position Time-MoE as a state-of-the-art solution for tackling real-world time series forecasting challenges with superior capability, efficiency, and flexibility.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density

    cs.LG 2026-05 unverdicted novelty 7.0

    Olivia harmonizes time series datasets via normalized power spectral density using a Harmonizer module and resonator-based HarmonicAttention, achieving state-of-the-art zero-shot, few-shot, and full-shot forecasting o...

  2. Beyond Static Forecasting: Unleashing the Power of World Models for Mobile Traffic Extrapolation

    cs.NI 2026-04 unverdicted novelty 7.0

    MobiWM is a multimodal world model for mobile networks that learns state-action dynamics to enable unlimited-horizon counterfactual traffic simulations and optimization.

  3. TS-Arena -- A Live Forecast Pre-Registration Platform

    cs.LG 2025-12 conditional novelty 7.0

    TS-Arena is a live pre-registration platform that evaluates time series forecasts on future data streams to eliminate information leakage.

  4. Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting

    cs.LG 2025-09 unverdicted novelty 7.0

    Super-Linear introduces a pretrained MoE architecture using frequency-specialized linear experts and spectral gating for efficient general time series forecasting.

  5. STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction

    cs.LG 2025-08 unverdicted novelty 7.0

    STM3 is a spatio-temporal mixture-of-multiscale-Mamba model that uses disentangled experts, stable routing, and causal contrastive learning to achieve state-of-the-art long-term forecasting on ten real-world benchmarks.

  6. Crossing-Free Probabilistic K-Line Forecasts Without Retraining

    stat.ML 2026-07 conditional novelty 6.0

    KQSP eliminates quantile and K-line crossings in probabilistic OHLC forecasts via sequential minimum-distance projections, without retraining and with smaller corrections than standard alternatives.

  7. Probabilistic Low-Voltage Peak Load Forecasting with Time Series Foundation Models Evaluated on Application-Oriented Metrics

    cs.LG 2026-07 unverdicted novelty 6.0

    Compares foundation models for probabilistic low-voltage load forecasting on 200 real feeders and introduces a grid-planning metric that scores peak prediction by its effect on asset cost-risk decisions.

  8. Learning Spatio-Temporal Foundation Models from Pure Synthetic Data

    cs.LG 2026-06 conditional novelty 6.0

    A spatio-temporal foundation model pre-trained exclusively on synthetic stochastic graph dynamics outperforms real-data-pretrained STFMs in zero-shot traffic forecasting, according to the paper's benchmarks.

  9. CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

    cs.LG 2026-06 unverdicted novelty 6.0

    CausalMoE is a multimodal foundation model with pattern-routed heterogeneous experts and LLM/VLM integration that claims new SOTA performance on supervised and few-shot Granger causal discovery benchmarks.

  10. LakeFM: Toward a Foundation Model for Aquatic Ecosystems Using Irregular Multivariate Multi-depth Time Series Data

    cs.LG 2026-06 unverdicted novelty 6.0

    LakeFM pre-trains on large ecological datasets to forecast irregular lake time series and reports competitive or superior performance with physically plausible outputs.

  11. Tyan-WP: A Wind Power Foundation Model for Ultra-Short-Term Probabilistic Forecasting

    cs.LG 2026-06 unverdicted novelty 6.0

    Tyan-WP is a pretrained wind power foundation model that outperforms site-specific TSMs and generic LTSMs in zero-shot ultra-short-term probabilistic forecasting on U.S. and U.K. sites via static embeddings and PAMF module.

  12. MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling

    cs.LG 2026-05 unverdicted novelty 6.0

    MILM fine-tunes LLMs on XML-encoded multimodal irregular time series via a two-stage process that exploits informative sampling patterns to achieve top performance on EHR classification datasets.

  13. Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration

    stat.ML 2026-05 unverdicted novelty 6.0

    A new MoE training method integrates expert-level losses and partial online updates to improve forecasting accuracy and efficiency over standard statistical and neural models.

  14. Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

    cs.LG 2026-03 conditional novelty 6.0

    Uncertainty-guided label rebalancing (uLNR) lifts UAV safety-prediction F1 to 0.806 under 46:1 imbalance by probabilistically flipping high-uncertainty safe windows to unsafe.

  15. Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling

    cs.AI 2026-03 unverdicted novelty 6.0

    Timer-S1 is a released 8.3B-parameter MoE time series model that achieves state-of-the-art MASE and CRPS scores on GIFT-Eval using serial scaling and Serial-Token Prediction.

  16. Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting

    cs.LG 2026-01 conditional novelty 6.0

    A model-agnostic module that retrieves common and rare prototype patterns improves forecasting error on many standard benchmarks, but not on all reported cases.

  17. Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models

    cs.LG 2025-09 unverdicted novelty 6.0

    Kairos is a parameter-efficient time series foundation model using dynamic patching tokenizer, mixture-of-size encoding, and spectral-conditioned positional embeddings to improve zero-shot forecasting on heterogeneous data.

  18. MoveFM-R: Advancing Mobility Foundation Models via Language-driven Semantic Reasoning

    cs.LG 2025-09 unverdicted novelty 6.0

    MoveFM-R is a framework that bridges mobility foundation models and LLMs using semantically enhanced location encoding, progressive curriculum alignment, and interactive self-reflection to generate plausible trajector...

  19. From Time Series Analysis to Question Answering: A Survey in the LLM Era

    cs.LG 2025-06 accept novelty 6.0

    A survey proposing a taxonomy of Injective, Bridging, and Internal Alignment paradigms to evolve TSA into user-driven Time Series Question Answering with LLMs.

  20. Feature to Dynamics: Feature-space to Autoregression strategy for Zero-shot Time Series Forecasting

    cs.LG 2026-05 unverdicted novelty 5.0

    FSA learns a mapping from feature space to autoregressive strategy space to improve zero-shot univariate time series forecasting over Transformer baselines under matched pretraining conditions.

  21. Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs

    cs.LG 2026-05 unverdicted novelty 5.0

    StockR1 unifies LLM-based financial reasoning and time-series forecasting by emitting verifiable forecast actions that condition a decoder, optimized via consistency-grounded RL to improve accuracy on QA and prediction tasks.

  22. TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning

    cs.CL 2025-10 conditional novelty 5.0

    A frozen time-series encoder aligned to an LLM through a small adapter and two-stage training beats same-scale open-source models on time-series reasoning benchmarks.

  23. STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction

    cs.LG 2025-08 unverdicted novelty 5.0

    STM3 is a new multiscale Mamba mixture-of-experts model with graph causal networks and contrastive routing that reports state-of-the-art results on 10 long-term spatio-temporal forecasting benchmarks.

  24. PRISM: Prioritized Channel Importance with Semi-supervised Domain Adaptation for Cross-Subject EEG Emotion Recognition

    cs.LG 2026-07 unverdicted novelty 4.0

    PRISM combines data-dependent channel weighting via expert ensemble and confidence-filtered pseudo-label domain adaptation to outperform prior methods on cross-subject EEG emotion tasks in DEAP, DREAMER, and SEED.

  25. Does Normalization Choice Matter for Causal Large Time-Series Models?

    cs.LG 2026-06 unverdicted novelty 4.0

    Normalization choice significantly influences training convergence and forecasting performance in causal large time-series models.

  26. Assessing the Operational Viability of Foundation Models for Time Series Forecasting

    cs.LG 2026-05 unverdicted novelty 4.0

    Foundation models match or approach supervised performance in periodic and cold-start domains but lag in physically constrained systems, while a feature-based router improves accuracy and cuts inference cost versus al...