Pith. sign in

REVIEW 31 cited by

Transformers in Time Series: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07125 v5 pith:O77F3C7E submitted 2022-02-15 cs.LG cs.AIeess.SPstat.ML

Transformers in Time Series: A Survey

classification cs.LG cs.AIeess.SPstat.ML
keywords seriestimetransformersanalysismodelingapplicationsperformperspective
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Transformers have achieved superior performances in many tasks in natural language processing and computer vision, which also triggered great interest in the time series community. Among multiple advantages of Transformers, the ability to capture long-range dependencies and interactions is especially attractive for time series modeling, leading to exciting progress in various time series applications. In this paper, we systematically review Transformer schemes for time series modeling by highlighting their strengths as well as limitations. In particular, we examine the development of time series Transformers in two perspectives. From the perspective of network structure, we summarize the adaptations and modifications that have been made to Transformers in order to accommodate the challenges in time series analysis. From the perspective of applications, we categorize time series Transformers based on common tasks including forecasting, anomaly detection, and classification. Empirically, we perform robust analysis, model size analysis, and seasonal-trend decomposition analysis to study how Transformers perform in time series. Finally, we discuss and suggest future directions to provide useful research guidance. To the best of our knowledge, this paper is the first work to comprehensively and systematically summarize the recent advances of Transformers for modeling time series data. We hope this survey will ignite further research interests in time series Transformers.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Adaptive Oscillatory-State Alignment for Time Series Forecasting

    cs.LG 2026-06 unverdicted novelty 7.0

    AOSNET is a Hilbert-guided framework that performs adaptive oscillatory-state alignment to a learnable prior for improved long-term forecasting under non-stationary conditions.

  2. HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation

    cs.HC 2026-05 unverdicted novelty 7.0

    HapticLDM is the first latent diffusion model that generates vibrotactile signals directly from text, using dynamic text curation and global denoising to improve realism and semantic alignment over autoregressive baselines.

  3. FactoryBench: Evaluating Industrial Machine Understanding

    cs.AI 2026-05 unverdicted novelty 7.0

    FactoryBench reveals that frontier LLMs achieve under 50% on structured causal questions and under 18% on decision-making in industrial robotic telemetry.

  4. Deep Time Series Models: A Comprehensive Survey and Benchmark

    cs.LG 2024-07 unverdicted novelty 7.0

    This survey and benchmark of deep time series models using the released TSLib library finds that models with specific structures perform well only on distinct analysis tasks.

  5. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers

    cs.LG 2022-11 conditional novelty 7.0

    PatchTST uses subseries patching and channel-independent Transformers to deliver significantly better long-term multivariate time series forecasting and strong self-supervised transfer performance.

  6. MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation

    cs.LG 2026-07 conditional novelty 6.0

    A multi-view behavior-aware conditional diffusion model for imputing missing utility-meter data is claimed to beat ten baselines on a Florida utility dataset, though the paper's own tables conflict with parts of the claim.

  7. HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data

    cs.AI 2026-07 conditional novelty 6.0

    A transformer with AttentiveCAT yields class-specific, time-step importance scores for wearable health data, beats deep-learning baselines, and beats random time-step selection in masking tests.

  8. Emergent Latent-State Computation under Stochastic Volatility

    cs.LG 2026-07 conditional novelty 6.0

    Volatility forecasters develop linearly decodable representations of the next hidden log-volatility state; in long cycles this appears immediately after the input projection and ℓ2 normalization.

  9. Transformers with Physics-Informed Encodings and Simulation-Based Inference for Robust Detection of Eccentric Binary Black Holes in Pulsar Timing Array Data

    cs.LG 2026-07 conditional novelty 6.0

    Physics-informed Transformer encodings plus conditional normalizing flows yield sharper, better-calibrated posteriors for eccentric BBHs in white-noise PTA data than physics-agnostic SBI baselines.

  10. TopoCast: A Topological Fidelity Framework for Evaluating Transformer-Based Time Series Forecasting

    cs.LG 2026-06 unverdicted novelty 6.0

    A topological evaluation framework for transformer time series forecasts that derives fidelity scores from persistence diagrams of delay embeddings, including a phase-aware localized metric.

  11. VESTA: Visual Exploration with Statistical Tool Agents

    cs.AI 2026-05 unverdicted novelty 6.0

    VESTA introduces dynamic tool creation for VLMs that outperforms static-tool and no-tool baselines on distribution fitting, time series, and astronomy tasks in the new DAWN benchmark.

  12. Frequency-Guided Deformable Networks for Continuous Phase Alignment

    eess.SP 2026-03 conditional novelty 6.0

    RFFT-derived periods guide deformable convolutions with Gaussian RBF interpolation and asymmetric routing to improve multi-task time-series modeling over rigid grids and bilinear sampling.

  13. Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling

    cs.AI 2026-03 unverdicted novelty 6.0

    Timer-S1 is a released 8.3B-parameter MoE time series model that achieves state-of-the-art MASE and CRPS scores on GIFT-Eval using serial scaling and Serial-Token Prediction.

  14. Day-Ahead Forecasting of Largest Single Infeed/Outfeed on the Irish Power Grid: A Generative Artificial Intelligence Approach

    eess.SY 2026-07 conditional novelty 5.0

    A transformer-based day-ahead forecaster for Ireland's largest single infeed matches an 8-hour operational model within 1.1% MAPE for infeed, but outfeed errors and the claimed 15% cost savings are not supported.

  15. Modular Foundation Models for Time-Series Perception in Digital Twins

    cs.LG 2026-07 conditional novelty 5.0

    A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.

  16. ASTEROID: A Spatiotemporal Information Transformer for Forecasting Multi-Step Time Series of Molecular Dynamics

    cs.LG 2026-06 unverdicted novelty 5.0

    ASTEROID is a spatiotemporal Transformer that predicts multi-step MD atomic coordinates with claimed higher accuracy and lower cost than iterative simulation on quantum-derived datasets.

  17. Disjoint or Overlapping? Inference Windowing for Reconstruction-Based Time Series Anomaly Detection

    cs.LG 2026-06 unverdicted novelty 5.0

    Overlapping inference windows improve reconstruction-based time series anomaly detection by up to 28% relative gain across models on TSB-AD and UCR benchmarks and can alter rankings.

  18. Kalimati Vegetable Price Index Forecasting with a Momentum Corrected Online Stacking Ensemble

    cs.LG 2026-05 unverdicted novelty 5.0

    A momentum-corrected online stacking ensemble forecasts the new Kalimati Vegetable Price Index with RMSE 1.771, MAPE 0.68 percent, and R-squared 0.845 at the 90-day horizon.

  19. Attention-based graph neural networks: a survey

    cs.SI 2026-05 unverdicted novelty 5.0

    The survey groups attention-based GNNs into three stages—graph recurrent attention networks, graph attention networks, and graph transformers—while reviewing architectures and future directions.

  20. climt-paraformer: Stable Emulation of Convective Parameterization using a Temporal Memory-aware Transformer

    physics.ao-ph 2026-04 unverdicted novelty 5.0

    A temporal memory-aware Transformer emulator for the Emanuel convective parameterization shows lower offline errors and 10-year stability in single-column model tests compared to memory-less MLP and LSTM baselines.

  21. ASTRAFier: A Novel and Scalable Transformer-based Stellar Variability Classifier

    astro-ph.IM 2026-04 unverdicted novelty 5.0

    ASTRAFier is a Transformer-BiLSTM-CNN model that classifies stellar variability from light curves, reporting 94.26% accuracy on Kepler data and 88.22% on TESS, then applied to 2.8 million TESS curves to release a catalog.

  22. TinyD\'ej\`aVu: Smaller RAM and Faster Inference with Neural Networks on MCUs for Sensor Data Streams

    cs.LG 2025-12 conditional novelty 5.0

    TinyDéjàVu turns time-series neural-network layers into streaming buffers (SSMs), cutting peak RAM by up to 99% and redundant compute on overlapping windows for microcontroller inference.

  23. TwinTac: A Wide-Range, Highly Sensitive Tactile Sensor with Real-to-Sim Digital Twin Sensor Model

    cs.RO 2025-09 conditional novelty 5.0

    A tactile sensor made from eight barometer chips reads forces from 0.01 N to over 200 N, and a learned FEM-to-signal model generates simulated tactile data that lifts shape classification accuracy from 33.6% to 95%.

  24. AR-KAN: Autoregressive-Weight-Enhanced Kolmogorov-Arnold Network for Time Series Forecasting

    cs.LG 2025-09 unverdicted novelty 5.0

    AR-KAN combines a pre-trained AR module with KAN to reduce redundancy while preserving temporal features, delivering lower probabilistic approximation error and stronger forecasting results on synthetic almost-periodi...

  25. FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series

    cs.LG 2026-07 conditional novelty 4.0

    FMMVCC combines Mamba-based encoders with multi-view contrastive learning and fuzzy clustering to achieve state-of-the-art univariate time series clustering with linear computational complexity.

  26. A Simple but Efficient Transformer-Based Physics-Informed Neural Network for Incompressible Navier--Stokes Equations

    physics.flu-dyn 2026-01 unverdicted novelty 4.0

    PhysicsFormer applies a lightweight Transformer PINN with pseudo-sequential representations to convection, Burgers, lid-driven cavity, and inverse Navier-Stokes problems, reporting near-zero error in parameter identif...

  27. Fourier-KAN-Mamba: A Novel State-Space Equation Approach for Time-Series Anomaly Detection

    cs.LG 2025-11 unverdicted novelty 4.0

    Fourier-KAN-Mamba combines Fourier features, KAN nonlinearities, and Mamba state-space modeling with a gating mechanism and reports better anomaly detection performance than prior methods on the MSL, SMAP, and SWaT be...

  28. Transformer-Guided Deep Reinforcement Learning for Optimal Takeoff Trajectory Design of an eVTOL Drone

    cs.LG 2025-11 unverdicted novelty 4.0

    Transformer-guided DRL cut training time steps to 25% of vanilla DRL while reaching 97.2% of optimal energy consumption for eVTOL takeoff versus 96.1%.

  29. Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction

    cs.LG 2026-06 unverdicted novelty 3.0

    Proposes Adaptive Financial Transformer with regime-gated attention and a composite loss to predict stock returns while claiming to fix backtesting issues and reduce complexity by 15.2%.

  30. Positional Encoding in Transformer-Based Time Series Models: A Survey

    cs.LG 2025-02 unverdicted novelty 3.0

    A survey of positional encoding methods in transformer-based time series models that evaluates fixed, learnable, relative, and hybrid approaches on classification tasks and links effectiveness to data characteristics.

  31. Federated Weather Modeling on Sensor Data

    cs.LG 2026-05 unverdicted novelty 2.0

    A federated learning framework lets distributed weather sensors train shared deep learning models for forecasting and anomaly detection while keeping raw data private.