Pith. sign in

REVIEW 4 cited by

SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10198 v3 pith:OVJ6ZYCY submitted 2024-02-15 cs.LG stat.ML

classification cs.LGstat.ML
keywords forecastingsamformertransformersattentionlinearmodelmultivariateseries
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting problem for which we show that transformers are incapable of converging to their true solution despite their high expressive power. We further identify the attention of transformers as being responsible for this low generalization capacity. Building upon this insight, we propose a shallow lightweight transformer model that successfully escapes bad local minima when optimized with sharpness-aware optimization. We empirically demonstrate that this result extends to all commonly used real-world multivariate time series datasets. In particular, SAMformer surpasses current state-of-the-art methods and is on par with the biggest foundation model MOIRAI while having significantly fewer parameters. The code is available at https://github.com/romilbert/samformer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NILMFormer: Non-Intrusive Load Monitoring that Accounts for Non-Stationarity

    cs.LG 2025-06 conditional novelty 6.0 of 10

    NILMFormer, a Transformer that normalizes each smart-meter window and feeds the window statistics back into the model, improves appliance-level power disaggregation accuracy over prior NILM models on four datasets.

  2. Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions

    cs.LG 2025-05 reject novelty 6.0 of 10

    CHARM is a 7M-parameter self-supervised embedding model for multivariate time series that uses channel descriptions to beat specialized baselines on forecasting, classification, and anomaly detection.

  3. Evaluation of a Foundational Model and Stochastic Models for Forecasting Sporadic or Spiky Production Outages of High-Performance Machine Learning Services

    cs.LG 2025-06 conditional novelty 5.0 of 10

    On seven years of monthly production outage counts from a large ML service, a fine-tuned TimesFM foundation model beats moving-average and autoregressive baselines for total outages, but per root cause the best model varies.

  4. Human in the Loop Adaptive Optimization for Improved Time Series Forecasting

    cs.LG 2025-05 reject novelty 3.0 of 10

    The core idea is standard forecast recalibration, and the reported experiments show mixed, sometimes negative, results with internal table errors, so the claim of consistent improvement fails.

Pith tools