REVIEW 4 cited by
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-horizon forecasting problems often contain a complex mix of inputs -- including static (i.e. time-invariant) covariates, known future inputs, and other exogenous time series that are only observed historically -- without any prior information on how they interact with the target. While several deep learning models have been proposed for multi-step prediction, they typically comprise black-box models which do not account for the full range of inputs present in common scenarios. In this paper, we introduce the Temporal Fusion Transformer (TFT) -- a novel attention-based architecture which combines high-performance multi-horizon forecasting with interpretable insights into temporal dynamics. To learn temporal relationships at different scales, the TFT utilizes recurrent layers for local processing and interpretable self-attention layers for learning long-term dependencies. The TFT also uses specialized components for the judicious selection of relevant features and a series of gating layers to suppress unnecessary components, enabling high performance in a wide range of regimes. On a variety of real-world datasets, we demonstrate significant performance improvements over existing benchmarks, and showcase three practical interpretability use-cases of TFT.
Forward citations
Cited by 4 Pith papers
-
Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos
MMPM uses PIM for gaze/head/hand interactions and MTP (CVAE with query decoder) to model separate crossing/non-crossing trajectory distributions, outperforming baselines on PIE and JAAD with a new validation protocol.
-
Scaling Sequential Recommendation Models with Transformers
Transformer-based sequential recommenders exhibit power-law and saturating NDCG scaling with model size and training interactions, enabling compute-aware model selection and effective pre-train/fine-tune transfer.
-
Forecasting Anonymized Electricity Load Profiles
Microaggregating smart meter data before forecasting leaves aggregated load forecasts accurate or improves them, with volatility dropping as group size k increases.
-
Causal Time-Series Synchronization for Multi-Dimensional Forecasting
Aligning cause-effect pairs by their estimated Granger lag improves channel-dependent forecasting accuracy and transfer learning on synthetic time-series data.
Discussion (0). Continue with ORCID to comment.