Pith. sign in

REVIEW 53 cited by

Transformers in Time Series: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07125 v5 pith:O77F3C7E submitted 2022-02-15 cs.LG cs.AIeess.SPstat.ML

classification cs.LGcs.AIeess.SPstat.ML
keywords seriestimetransformersanalysismodelingapplicationsperformperspective
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformers have achieved superior performances in many tasks in natural language processing and computer vision, which also triggered great interest in the time series community. Among multiple advantages of Transformers, the ability to capture long-range dependencies and interactions is especially attractive for time series modeling, leading to exciting progress in various time series applications. In this paper, we systematically review Transformer schemes for time series modeling by highlighting their strengths as well as limitations. In particular, we examine the development of time series Transformers in two perspectives. From the perspective of network structure, we summarize the adaptations and modifications that have been made to Transformers in order to accommodate the challenges in time series analysis. From the perspective of applications, we categorize time series Transformers based on common tasks including forecasting, anomaly detection, and classification. Empirically, we perform robust analysis, model size analysis, and seasonal-trend decomposition analysis to study how Transformers perform in time series. Finally, we discuss and suggest future directions to provide useful research guidance. To the best of our knowledge, this paper is the first work to comprehensively and systematically summarize the recent advances of Transformers for modeling time series data. We hope this survey will ignite further research interests in time series Transformers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 53 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 67 citations worldwide. Full citation record

  1. Saliency Methods are Encoders: Analysing Logical Relations Towards Interpretation

    cs.LG 2024-12 conditional novelty 7.0 of 10

    On synthetic AND/OR/XOR tasks, saliency methods frequently rank irrelevant baseline inputs above logically relevant ones, and retrained models can recover class information from masked inputs, indicating that score or...

  2. FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification

    cs.LG 2026-08 conditional novelty 6.0 of 10

    FreSH, a frequency-segmented multi-expert network, reports the highest average accuracy (76.1%) across 30 UEA multivariate time series classification benchmarks with a compact 54k-parameter model.

  3. MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A multi-view behavior-aware conditional diffusion model for imputing missing utility-meter data is claimed to beat ten baselines on a Florida utility dataset, though the paper's own tables conflict with parts of the claim.

  4. HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A transformer with AttentiveCAT yields class-specific, time-step importance scores for wearable health data, beats deep-learning baselines, and beats random time-step selection in masking tests.

  5. Emergent Latent-State Computation under Stochastic Volatility

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Volatility forecasters develop linearly decodable representations of the next hidden log-volatility state; in long cycles this appears immediately after the input projection and ℓ2 normalization.

  6. Transformers with Physics-Informed Encodings and Simulation-Based Inference for Robust Detection of Eccentric Binary Black Holes in Pulsar Timing Array Data

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Physics-informed Transformer encodings plus conditional normalizing flows yield sharper, better-calibrated posteriors for eccentric BBHs in white-noise PTA data than physics-agnostic SBI baselines.

  7. Frequency-Guided Deformable Networks for Continuous Phase Alignment

    eess.SP 2026-03 conditional novelty 6.0 of 10

    RFFT-derived periods guide deformable convolutions with Gaussian RBF interpolation and asymmetric routing to improve multi-task time-series modeling over rigid grids and bilinear sampling.

  8. BiND: A Neural Discriminator-Decoder for Accurate Bimanual Trajectory Prediction in Brain-Computer Interfaces

    q-bio.NC 2025-08 conditional novelty 6.0 of 10

    BiND, a discriminator-decoder architecture with motion-type routing and trial-relative time indexing, improves unimanual and bimanual hand velocity decoding from intracortical signals by about 2% over a plain GRU on o...

  9. Predicting Large-scale Urban Network Dynamics with Energy-informed Graph Neural Diffusion

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A scalable spatiotemporal Transformer, ScaleSTF, matches the accuracy of much larger models on city-scale forecasting tasks at a fraction of the compute and memory cost.

  10. The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Bidirectional joint-attention, complete forecasting aggregation, and direct mapping form the most effective Transformer design for long-term time series forecasting.

  11. A foundation model with multi-variate parallel attention to generate neuronal activity

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MVPFormer, a transformer with disentangled content, time, and channel attention, achieves expert-level zero-shot seizure detection on 50 unseen patients and near-SOTA results on speech decoding, alongside the largest ...

  12. Neural Functions for Learning Periodic Signal

    cs.LG 2025-06 conditional novelty 6.0 of 10

    NeRT factorizes periodic signals into a sine-based periodic factor and an unbounded scale factor, enabling extrapolation beyond the training range on several periodic benchmarks.

  13. Estimating Perceptual Attributes of Haptic Textures Using Visuo-Tactile Data

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A visuo-tactile deep network using a CNN autoencoder and ConvLSTM predicts four haptic attribute ratings from images and tool vibrations, beating single-modality baselines in leave-one-out tests.

  14. Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators

    cs.AR 2025-05 conditional novelty 6.0 of 10

    A fused exponential-multiplication hardware unit using logarithmic quantization and exponent adjustment reduces FlashAttention accelerator area by about 29% and power by about 18% without visible accuracy loss on GLUE.

  15. TimeCapsule: Solving the Jigsaw Puzzle of Long-Term Time Series Forecasting with Compressed Predictive Representations

    cs.LG 2025-04 conditional novelty 6.0 of 10

    TimeCapsule compresses multivariate time series into a small 3D tensor using learned mode products, forecasts inside that compressed space, and reports state-of-the-art results on ten LTSF benchmarks.

  16. Battling the Non-stationarity in Time Series Forecasting via Test-time Adaptation

    cs.LG 2025-01 conditional novelty 6.0 of 10

    TAFAS improves frozen time series forecasters at test time by calibrating inputs and outputs with periodicity-scheduled partial ground truth.

  17. Enhancing Fluorescence Lifetime Parameter Estimation Accuracy with Differential Transformer Based Deep Learning Model Incorporating Pixelwise Instrument Response Function

    eess.IV 2024-11 conditional novelty 6.0 of 10

    MFliNet, a differential-transformer network that takes pixelwise IRF as input, estimates fluorescence lifetime parameters robustly under IRF shifts from surface height variations.

  18. Day-Ahead Forecasting of Largest Single Infeed/Outfeed on the Irish Power Grid: A Generative Artificial Intelligence Approach

    eess.SY 2026-07 conditional novelty 5.0 of 10

    A transformer-based day-ahead forecaster for Ireland's largest single infeed matches an 8-hour operational model within 1.1% MAPE for infeed, but outfeed errors and the claimed 15% cost savings are not supported.

  19. Modular Foundation Models for Time-Series Perception in Digital Twins

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.

  20. TinyD\'ej\`aVu: Smaller RAM and Faster Inference with Neural Networks on MCUs for Sensor Data Streams

    cs.LG 2025-12 conditional novelty 5.0 of 10

    TinyDéjàVu turns time-series neural-network layers into streaming buffers (SSMs), cutting peak RAM by up to 99% and redundant compute on overlapping windows for microcontroller inference.

  21. TwinTac: A Wide-Range, Highly Sensitive Tactile Sensor with Real-to-Sim Digital Twin Sensor Model

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A tactile sensor made from eight barometer chips reads forces from 0.01 N to over 200 N, and a learned FEM-to-signal model generates simulated tactile data that lifts shape classification accuracy from 33.6% to 95%.

  22. Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A human-AI framework uses horizontal and vertical segmentation with LLMs, expert co-scoring, and LSTM anomaly detection to extract behavioral patterns from eye-tracking data.

  23. Scalable Unsupervised Segmentation via Random Fourier Feature-based Gaussian Process

    cs.LG 2025-07 conditional novelty 5.0 of 10

    RFF-GP-HSMM speeds up unsupervised time-series segmentation by approximating Gaussian processes with random Fourier features, cutting computation time by up to 278 times on motion capture data with similar accuracy.

  24. Evaluation of a Foundational Model and Stochastic Models for Forecasting Sporadic or Spiky Production Outages of High-Performance Machine Learning Services

    cs.LG 2025-06 conditional novelty 5.0 of 10

    On seven years of monthly production outage counts from a large ML service, a fine-tuned TimesFM foundation model beats moving-average and autoregressive baselines for total outages, but per root cause the best model varies.

  25. Benchmarking Unsupervised Strategies for Anomaly Detection in Multivariate Time Series

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Across ten public datasets, a reconstruction-based inverted transformer with per-variate anomaly labelling achieves the best or tied best MCC on most datasets, but the comparison is weakened by test-set-based configur...

  26. Representation Learning of Limit Order Book: A Comprehensive Study and Benchmarking

    cs.CE 2025-05 conditional novelty 5.0 of 10

    LOBench provides a standardized China A-share limit order book benchmark and claims that reconstruction-trained representations support downstream tasks and transfer better than end-to-end models.

  27. Gateformer: Advancing Multivariate Time Series Forecasting through Temporal and Variate-Wise Attention with Gated Representations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Gateformer combines temporal patching attention, variate-wise attention, and two gating mechanisms to achieve the best average forecasting error on most of 13 multivariate benchmarks, and it reports improved accuracy ...

  28. Unifying Prediction and Explanation in Time-Series Transformers via Shapley-based Pretraining

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A time-series transformer trained with Shapley-based and contrastive pretraining outputs predictions and Shapley explanations in one forward pass, at a fraction of post-hoc explanation cost.

  29. Gaze Prediction as a Function of Eye Movement Type and Individual Differences

    cs.HC 2024-12 conditional novelty 5.0 of 10

    Gaze prediction accuracy varies strongly across subjects, and per-subject fixation noise and saccade velocity correlate with prediction error across three different model architectures.

  30. Revisiting PCA for time series reduction in temporal dimension

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Applying PCA to the time axis of series windows before deep model training keeps average task accuracy while cutting compute and memory, but gains and losses vary strongly by model and dataset.

  31. xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition

    cs.LG 2024-12 conditional novelty 5.0 of 10

    xPatch, a dual-stream CNN/MLP forecaster with exponential moving average decomposition, patches, and channel independence, reports lower MSE/MAE than CARD and PatchTST on most of nine LTSF benchmarks.

  32. A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A benchmark shows most time-series Transformers tolerate about 50% unstructured pruning without clear accuracy loss, while structured pruning rarely delivers meaningful inference speedups.

  33. Multimodal Integration of Longitudinal Noninvasive Diagnostics for Survival Prediction in Immunotherapy Using Deep Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A transformer-based temporal attention network that integrates longitudinal blood, imaging, and medication data predicts mortality in immunotherapy patients with AUCs around 0.81 to 0.84, slightly beating blood-only models.

  34. QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting

    cs.AI 2026-08 reject novelty 4.0 of 10

    QFCQT claims up to 43.9% MSE improvement over HAT on ETTh2 via Lee-oscillator gated activation, but internal table inconsistencies and undefined ablation components undermine the claim.

  35. FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series

    cs.LG 2026-07 conditional novelty 4.0 of 10

    FMMVCC combines Mamba-based encoders with multi-view contrastive learning and fuzzy clustering to achieve state-of-the-art univariate time series clustering with linear computational complexity.

  36. TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization

    cs.SD 2025-08 reject novelty 4.0 of 10

    TinyMusician distills MusicGen and applies hand-picked mixed-precision quantization to make a 1.04 GB on-device music generator, but the headline '93% quality, 55% smaller' claims conflict with the paper's own tables.

  37. Transformer with Koopman-Enhanced Graph Convolutional Network for Spatiotemporal Dynamics Forecasting

    cs.LG 2025-07 reject novelty 4.0 of 10

    TK-GCN, a Koopman-enhanced graph convolution plus Transformer model, is proposed for spatiotemporal forecasting; its claimed consistent superiority is contradicted by its own ablation results.

  38. Time Series Transformer-Based Modeling of Pavement Skid and Texture Deterioration

    stat.AP 2025-07 reject novelty 4.0 of 10

    A time series transformer is reported to predict post-milling skid number with R2=0.981, but the evaluation may leak temporal information and the underlying data counts are inconsistent.

  39. Scalable Machine Learning Algorithms using Path Signatures

    stat.ML 2025-06 conditional novelty 4.0 of 10

    Path signatures can be embedded in Gaussian process, deep learning, kernel, and graph diffusion models to match or beat established baselines on time series and graph benchmarks.

  40. Synthetic Time Series Forecasting with Transformer Architectures: Extensive Simulation Benchmarks

    cs.LG 2025-05 reject novelty 4.0 of 10

    Autoformer and PatchTST outperform Informer across synthetic forecasting benchmarks, while a proposed Koopman-Transformer hybrid is illustrated on Van der Pol and Lorenz systems.

  41. Dynamic Perturbed Adaptive Method for Infinite Task-Conflicting Time Series

    cs.LG 2025-05 reject novelty 4.0 of 10

    A trunk-branch method for adapting to conflicting time series tasks reports large error reductions on a synthetic benchmark, but the comparison is confounded by unequal adaptation budgets and the theory overclaims rel...

  42. Non-Stationary Time Series Forecasting Based on Fourier Analysis and Cross Attention Mechanism

    cs.LG 2025-05 reject novelty 4.0 of 10

    AEFIN combines frequency-domain decomposition, cross-attention, and a Fourier feature network to forecast non-stationary time series, with partial improvements over some baselines.

  43. TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data

    q-fin.ST 2025-02 conditional novelty 4.0 of 10

    A dual-attention transformer and a simple MLP both outperform prior limit order book trend prediction models across FI-2010, Tesla/Intel, and Bitcoin datasets, with apparent decline in predictability between 2012 and 2015.

  44. Leveraging Video Vision Transformer for Alzheimer's Disease Diagnosis from 3D Brain MRI

    eess.IV 2025-01 reject novelty 4.0 of 10

    Applying a video vision transformer (ViViT) to 3D brain MRI slices treated as video frames yields 98.6% three-class accuracy on ADNI1 Complete 3Yr 3T, outperforming the paper's CNN-BiLSTM and ViT-BiLSTM baselines.

  45. Saliency Maps are Ambiguous: Analysis of Logical Relations on First and Second Order Attributions

    cs.LG 2025-01 conditional novelty 4.0 of 10

    On synthetic AND/OR/XOR datasets with perfectly accurate models, every tested saliency method sometimes ranks a truly irrelevant input above a necessary one, so the scores cannot be trusted as relevance rankings.

  46. Deep Learning and Foundation Models for Weather Prediction: A Survey

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A survey that organizes deep learning weather prediction models into three training paradigms: deterministic, generative, and pre-train-fine-tune.

  47. SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults

    cs.AI 2024-12 reject novelty 4.0 of 10

    SubstationAI, a fine-tuned LLaVA-1.5-7B model augmented with a fault knowledge base, receives higher expert ratings than GPT-4 for substation fault reports, but suspected train/test overlap makes the result unreliable.

  48. Paraformer: Parameterization of Sub-grid Scale Processes Using Transformers

    cs.LG 2024-12 conditional novelty 4.0 of 10

    Paraformer, an encoder-only Transformer with a five-step context window, achieves lower mean absolute error than an MLP on ClimSim sub-grid scale parameterization for two variable sets.

  49. Auto-Regressive Moving Diffusion Models for Time Series Forecasting

    cs.LG 2024-12 conditional novelty 4.0 of 10

    ARMD replaces noise in diffusion models with a deterministic sliding of the series window, turning denoising into iterative forecasting, and reports SOTA results on 12 of 14 diffusion-baseline settings.

  50. Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?

    cs.LG 2025-06 reject novelty 3.0 of 10

    LLM4TS_FS achieves the best MSE on four of seven long-term datasets, but the claimed broad advantage of pre-trained large models over small transformers is not consistent across all benchmarks.

  51. W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling

    cs.LG 2025-06 conditional novelty 3.0 of 10

    W4S4 initializes S4 state space models with WaLRUS wavelet frames and reports better delay reconstruction and classification accuracy than HiPPO-based S4, with frozen (A,B).

  52. Data Mining in Transportation Networks with Graph Neural Networks: A Review and Outlook

    cs.LG 2025-01 conditional novelty 3.0 of 10

    A review of graph neural network applications in transportation networks, covering traffic prediction, operations, industry practice, future directions, and public datasets and code.

  53. EDformer: Embedded Decomposition Transformer for Interpretable Multivariate Time Series Predictions

    cs.LG 2024-12 reject novelty 3.0 of 10

    EDformer combines moving-average decomposition with an iTransformer-style variate-token encoder and claims state-of-the-art forecasting, but its reported benchmark results do not consistently support that claim.

Pith tools