Pith. sign in

REVIEW 45 cited by

MOMENT: A Family of Open Time-series Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03885 v3 pith:22PYNDU6 submitted 2024-02-06 cs.LG cs.AI

classification cs.LGcs.AI
keywords timeseriesmodelslargediversefoundationpre-trainedautonlab
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce MOMENT, a family of open-source foundation models for general-purpose time series analysis. Pre-training large models on time series data is challenging due to (1) the absence of a large and cohesive public time series repository, and (2) diverse time series characteristics which make multi-dataset training onerous. Additionally, (3) experimental benchmarks to evaluate these models, especially in scenarios with limited resources, time, and supervision, are still in their nascent stages. To address these challenges, we compile a large and diverse collection of public time series, called the Time series Pile, and systematically tackle time series-specific challenges to unlock large-scale multi-dataset pre-training. Finally, we build on recent work to design a benchmark to evaluate time series foundation models on diverse tasks and datasets in limited supervision settings. Experiments on this benchmark demonstrate the effectiveness of our pre-trained models with minimal data and task-specific fine-tuning. Finally, we present several interesting empirical observations about large pre-trained time series models. Pre-trained models (AutonLab/MOMENT-1-large) and Time Series Pile (AutonLab/Timeseries-PILE) are available on Huggingface.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 45 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A shared cardiac encoder pretrained with delay-aware cross-modal JEPA on ECG, PPG, and PCG beats modality-specific self-supervised baselines on 25 downstream tasks.

  2. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A multi-vehicle naturalistic benchmark finds learned driving embeddings retain driver identity under condition matching, while descriptors collapse and video re-ID is mostly route leakage.

  3. AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection

    cs.LG 2026-02 conditional novelty 7.0 of 10

    New RL approach (TimerPO) with ground-truth-generated expert reasoning traces lets 3B-7B multimodal LLMs outperform GPT-4o on time-series anomaly detection and explanation.

  4. LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A whole-log transformer with a geology-aware loss predicts stratigraphic zones and marker depths, beating sliding-window baselines on three well-log datasets.

  5. Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

    stat.AP 2026-07 conditional novelty 6.0 of 10

    Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.

  6. Learning Spatio-Temporal Foundation Models from Pure Synthetic Data

    cs.LG 2026-06 conditional novelty 6.0 of 10

    A spatio-temporal foundation model pre-trained exclusively on synthetic stochastic graph dynamics outperforms real-data-pretrained STFMs in zero-shot traffic forecasting, according to the paper's benchmarks.

  7. TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models

    cs.LG 2026-01 conditional novelty 6.0 of 10

    TimeSAE trains a sparse autoencoder with counterfactual and consistency losses to explain black-box time series predictions, claiming better faithfulness and out-of-distribution robustness than eight baselines.

  8. Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TimeRCD, a transformer pre-trained on 2.5 billion synthetic data points with context-relative anomaly labels, beats reconstruction-based foundation models on most zero-shot TSAD benchmarks, while the evaluation tunes ...

  9. WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A wind-specific foundation model, WindFM, uses hierarchical tokenization and autoregressive pre-training on the NREL WIND Toolkit to achieve state-of-the-art zero-shot wind power forecasts.

  10. CALM: A Framework for Continuous, Adaptive, and LLM-Mediated Anomaly Detection in Time-Series Streams

    cs.LG 2025-08 reject novelty 6.0 of 10

    CALM uses an LLM-as-a-Judge to curate anomalies for continuous fine-tuning of a time-series foundation model, improving anomaly detection on held-out stream segments.

  11. Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Pre-trained multivariate time-series imputation models frequently return values that violate known relations between variables, and a diffusion-based score can detect and filter these errors.

  12. Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs

    cs.LG 2025-06 unverdicted novelty 6.0 of 10

    Time-R1 trains LLMs via supervised fine-tuning followed by reinforcement learning with a time-series-specific reward and non-uniform GRIP sampling to enable multi-step reasoning that improves forecasting accuracy.

  13. Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions

    cs.LG 2025-05 reject novelty 6.0 of 10

    CHARM is a 7M-parameter self-supervised embedding model for multivariate time series that uses channel descriptions to beat specialized baselines on forecasting, classification, and anomaly detection.

  14. Decoding Latent Spaces: Assessing the Interpretability of Time Series Foundation Models for Visual Analytics

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Fine-tuning MOMENT time series foundation models reduces reconstruction loss but does not visually improve the interpretability of their latent space projections in the DeepVATS visual analytics environment.

  15. Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications Across Lab and Field Settings

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A field-trained, open-source PPG foundation model outperforms a clinical-data-trained model on 10 of 11 downstream health tasks across wearable and clinical settings.

  16. TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Retrieving similar time-series segments from a multi-domain knowledge base and injecting them through a learned Channel Prompting module improves zero-shot forecasting of a frozen TSFM, though gains are small and leak...

  17. Federated Foundation Models on Heterogeneous Time Series

    cs.LG 2024-12 conditional novelty 6.0 of 10

    FFTS is a federated pretraining framework with a timescale-aware mixture-of-experts module that trains a time series foundation model from scratch across heterogeneous, non-shared datasets.

  18. Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A wavelet-based tokenizer that lets an autoregressive transformer forecast quantized wavelet coefficients instead of raw values, improving accuracy and generalization on time series benchmarks.

  19. The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data

    astro-ph.IM 2024-12 conditional novelty 6.0 of 10

    The Multimodal Universe compiles hundreds of millions of astronomical observations from surveys such as DESI, Gaia and JWST into a unified 100 TB multimodal dataset for machine learning.

  20. RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data

    eess.SP 2024-11 conditional novelty 6.0 of 10

    RelCon pretrains an accelerometry foundation model with relative contrastive learning over a learned motif distance and achieves state-of-the-art activity recognition and gait regression.

  21. Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models

    cs.LG 2026-08 conditional novelty 5.0 of 10

    Reward-based fine-tuning of time series foundation models can collapse predictions away from the true future; steering probability mass into a ground-truth neighborhood reduces that collapse and improves forecasts.

  22. Benchmarking Foundation Models with Multimodal Public Electronic Health Records

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A standardized MIMIC-IV benchmark comparing eight unimodal and multimodal foundation models shows multimodal inputs improve predictive performance without adding bias, while medical LVLMs underperform on length-of-sta...

  23. Towards Interpretable Time Series Foundation Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    After fine-tuning on 180 synthetic mean-reverting series annotated by a large multimodal model, small Qwen models can describe trend direction, noise intensity, and extremum location in natural language.

  24. Causal Foundation Models: Disentangling Physics from Instrument Properties

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A dual-encoder contrastive model trained on star-instrument triplets learns separate stellar and instrumental latent spaces, improving few-shot prediction of stellar parameters in simulated TESS-like light curves.

  25. Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives

    cs.LG 2025-06 reject novelty 5.0 of 10

    TimesCLIP aligns image-based and text-based views of the same time series via contrastive learning to improve forecasting accuracy on several benchmarks, but the full multimodal model is not used on two of the six lon...

  26. Cross-Modal Epileptic Signal Harmonization: Frequency Domain Mapping Quantization for Pre-training a Unified Neurophysiological Transformer

    q-bio.NC 2025-06 conditional novelty 5.0 of 10

    EpiNT, a transformer pretrained on over 2,700 hours of EEG and iEEG from 1,199 patients with masked autoencoding and a frequency-domain quantizer, matches or beats other pretrained models on six epilepsy classificatio...

  27. Frame-Level Real-Time Assessment of Stroke Rehabilitation Exercises from Video-Level Labeled Data: Task-Specific vs. Foundation Models

    eess.IV 2025-06 conditional novelty 5.0 of 10

    Using gradient saliency and pretrained video/time-series models, frame-level stroke exercise quality can be learned from video-level labels, with the best configuration reaching 72% AUC versus 69% for a ground-truth-t...

  28. DELPHYNE: A Pre-Trained Model for General and Financial Time Series

    q-fin.ST 2025-05 conditional novelty 5.0 of 10

    The paper reports that a time-series transformer pretrained on public and proprietary financial data becomes competitive on financial tasks after fine-tuning, while zero-shot general forecasting remains behind MOIRAI.

  29. How Effective are Large Time Series Models in Hydrology? A Study on Water Level Forecasting in Everglades

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Chronos, a pretrained foundation model used zero-shot, beat 16 other models at forecasting Everglades water levels across five stations and horizons from 7 to 28 days.

  30. Gateformer: Advancing Multivariate Time Series Forecasting through Temporal and Variate-Wise Attention with Gated Representations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Gateformer combines temporal patching attention, variate-wise attention, and two gating mechanisms to achieve the best average forecasting error on most of 13 multivariate benchmarks, and it reports improved accuracy ...

  31. Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Time2Lang learns a lightweight adapter that maps time-series foundation model embeddings into a frozen LLM's input space, enabling mental health classification from wearable data without text prompting.

  32. Assessing Foundation Models' Transferability to Physiological Signals in Precision Medicine

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A simulation-based evaluation pipeline shows that Moirai embeddings distort physiological signal structure, with feature entanglement, lost temporal dynamics, and reduced scenario discrimination.

  33. Towards Foundation Models for Critical Care Time Series

    cs.LG 2024-11 conditional novelty 5.0 of 10

    The paper introduces a harmonized multi-center critical care time series dataset and transfer benchmark covering nine datasets from three continents, with treatment variables, and compares seven models on early event ...

  34. Lightweight Test-Time Adaptation for EMG-Based Gesture Recognition

    cs.LG 2026-01 conditional novelty 4.0 of 10

    Test-time adaptation—adaptive batch norm, replay-regularized feature alignment, and few-shot meta-learning—improves NinaPro DB6 inter-session EMG gesture accuracy from ~57% baseline to ~69–70% unsupervised and ~80% wi...

  35. On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating

    cs.LG 2025-08 conditional novelty 4.0 of 10

    On four public datasets, Gradient Boosting with hand-built features beat Chronos, Llama, and ARIMA on most accuracy metrics, while Chronos only led on financial sMAPE.

  36. Enhancing Transformer-Based Foundation Models for Time Series Forecasting via Bagging, Boosting and Statistical Ensembles

    cs.LG 2025-08 conditional novelty 4.0 of 10

    On one Belgian electricity load series, bagging, regression stacking, and residual correction reduce MSE relative to standalone Lag-Llama and AutoGluon forecasts, though the reported numbers are inconsistent and lack ...

  37. MoFE-Time: Mixture of Frequency Domain Experts for Time-Series Forecasting Models

    cs.LG 2025-07 conditional novelty 4.0 of 10

    MoFE-Time reports average MSE 0.2755 and MAE 0.3226 across six public benchmarks, about 7% lower than Time-MoE, by adding frequency-domain experts to a Mixture of Experts transformer.

  38. Grounding Intelligence in Movement

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.

  39. Delayformer: spatiotemporal transformation for predicting high-dimensional dynamics

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Delayformer embeds each time series variable into a Hankel matrix, processes the matrices as images with a shared ViT, and predicts all variables in parallel.

  40. Goal-Oriented Time-Series Forecasting: Foundation Framework Design

    cs.LG 2025-04 conditional novelty 4.0 of 10

    A single forecasting model can be trained to adapt its predictions to any target value interval at inference time by discretizing the output range during training and patching predictions at test time.

  41. Early Risk Prediction of Pediatric Cardiac Arrest from Electronic Health Records via Multimodal Fused Transformer

    cs.LG 2025-02 conditional novelty 4.0 of 10

    A fused tabular-textual transformer modestly improves early pediatric cardiac arrest prediction on four of five metrics in a private CICU cohort, but not on AUROC.

  42. Foundation Models for CPS-IoT: Opportunities and Challenges

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Current foundation models fall short on CPS-IoT needs in resource efficiency, spatial generalization, long-term context, and knowledge integration; the paper proposes desiderata and a community roadmap.

  43. Enhancing Masked Time-Series Modeling via Dropping Patches

    stat.ML 2024-12 reject novelty 4.0 of 10

    Randomly dropping 60% of patches before masked pre-training improves PatchTST time-series forecasting accuracy and training speed, though the paper's theoretical explanation is not sound.

  44. Comparative Analysis of Zero-Shot Capability of Time-Series Foundation Models in Short-Term Load Prediction

    eess.SY 2024-12 conditional novelty 4.0 of 10

    Zero-shot time-series foundation models, especially Chronos, outperform trained GP and SVR baselines on short-term load prediction across 11 UK, German, and Dutch datasets.

  45. Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?

    cs.LG 2025-06 reject novelty 3.0 of 10

    LLM4TS_FS achieves the best MSE on four of seven long-term datasets, but the claimed broad advantage of pre-trained large models over small transformers is not consistent across all benchmarks.

Pith tools