REVIEW 56 cited by
N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We focus on solving the univariate times series point forecasting problem using deep learning. We propose a deep neural architecture based on backward and forward residual links and a very deep stack of fully-connected layers. The architecture has a number of desirable properties, being interpretable, applicable without modification to a wide array of target domains, and fast to train. We test the proposed architecture on several well-known datasets, including M3, M4 and TOURISM competition datasets containing time series from diverse domains. We demonstrate state-of-the-art performance for two configurations of N-BEATS for all the datasets, improving forecast accuracy by 11% over a statistical benchmark and by 3% over last year's winner of the M4 competition, a domain-adjusted hand-crafted hybrid between neural network and statistical time series models. The first configuration of our model does not employ any time-series-specific components and its performance on heterogeneous datasets strongly suggests that, contrarily to received wisdom, deep learning primitives such as residual blocks are by themselves sufficient to solve a wide range of forecasting problems. Finally, we demonstrate how the proposed architecture can be augmented to provide outputs that are interpretable without considerable loss in accuracy.
Forward citations
Cited by 56 Pith papers
-
Compositional Covariate Importance Testing via Partial Conjunction of Bivariate Hypotheses
Under compositionality, the unique nontrivial Markov boundary of a response equals the set of covariates for which every bivariate conditional independence with another covariate fails, enabling valid partial-conjunct...
-
Pretraining Large Brain Language Model for Active BCI: Silent Speech
Autoregressive spectro-temporal pretraining of an EEG transformer improves silent speech decoding accuracy in cross-session tests, with a new 120-hour dataset.
-
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
FM-LLM uses Fourier token embeddings and a mixture-of-experts decoder to adapt frozen LLMs to time series forecasting, reporting gains over AutoTimes across many benchmarks.
-
Multi-Source Dynamic Graph Learning for Compound-Flood Forecasting in Managed Coastal Systems
An anchored forecaster that adds bounded, regime-gated corrections from a dynamic multi-station graph to a local temporal forecast improves sustained high-water plateau prediction in South Florida without harming rout...
-
Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting
A model-agnostic module that retrieves common and rare prototype patterns improves forecasting error on many standard benchmarks, but not on all reported cases.
-
Fremer: Lightweight and Effective Frequency Transformer for Workload Forecasting in Cloud Services
Fremer forecasts cloud workloads by aligning frequency spectra via linear padding, filtering noise, and attending over frequency combinations.
-
Leveraging External Factors in Household-Level Electrical Consumption Forecasting using Hypernetworks
A hypernetwork that generates per-household linear forecast weights is the only global model in the test that improves its error when weather, holiday, and football-event data are added, beating other global models on...
-
Does Scaling Law Apply in Time Series Forecasting?
A parameter-light adaptive linear model (ALinear) outperforms larger baselines on long-horizon univariate forecasting benchmarks while using under 1% of their parameters, but the efficiency comparison rests on questio...
-
TimeCapsule: Solving the Jigsaw Puzzle of Long-Term Time Series Forecasting with Compressed Predictive Representations
TimeCapsule compresses multivariate time series into a small 3D tensor using learned mode products, forecasts inside that compressed space, and reports state-of-the-art results on ten LTSF benchmarks.
-
Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting
Sensorformer uses a two-stage cross-patch attention mechanism with global-patch compression to improve multivariate time series forecasting accuracy while reducing attention complexity.
-
TimeRAG: BOOSTING LLM Time Series Forecasting via Retrieval-Augmented Generation
Using K-means and DTW to retrieve similar historical sequences and add them to an LLM prompt improves M4 forecast accuracy by 2.97% on average over Time-LLM.
-
Federated Foundation Models on Heterogeneous Time Series
FFTS is a federated pretraining framework with a timescale-aware mixture-of-experts module that trains a time series foundation model from scratch across heterogeneous, non-shared datasets.
-
APS-LSTM: Exploiting Multi-Periodicity and Diverse Spatial Dependencies for Flood Forecasting
APS-LSTM combines FFT-based multi-period division, periodic and spatial self-attention, and LSTM encoding to improve flood flow forecasts on two real-world watershed datasets.
-
E-STGCN: Extreme Spatiotemporal Graph Convolutional Networks for Air Quality Forecasting
A spatiotemporal graph network with a generalized Pareto loss is introduced and shown to outperform most benchmarks for Delhi PM2.5, PM10, and NO2 forecasting.
-
GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting
Adaptive soft routing among three complementary linear bases via a channel-horizon-phase gate yields competitive multivariate forecasting accuracy with a small, interpretable model.
-
Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework
WrapFlow combines continuous-time event/gap tokenization with simulation-free residual flow matching on a Transformer to improve irregular multivariate time-series forecasting.
-
Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting
A top-down hierarchical Transformer with a coherence loss jointly forecasts Portuguese ED demand at 81 hospitals, 5 regions, and national level, cutting mean WAPE ~32% versus the best non-hierarchical deep baseline wh...
-
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
A contextual-bandit correction layer with few-shot masked updates improves ML demand forecasts by 3.7–14.9% and cuts inventory costs in two retail datasets.
-
OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.
-
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
TimesCLIP aligns image-based and text-based views of the same time series via contrastive learning to improve forecasting accuracy on several benchmarks, but the full multimodal model is not used on two of the six lon...
-
Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models
Fine-tuned time-series foundation models, especially Chronos with LoRA, outperform trained-from-scratch deep forecasters on multi-signal building energy forecasting with limited data.
-
Forecasting Multivariate Urban Data via Decomposition and Spatio-Temporal Graph Analysis
A graph neural network that learns one dependency graph per decomposed time-series component improves long-term urban forecasts by a few percent over baselines.
-
Enhancing LLMs for Time Series Forecasting via Structure-Guided Cross-Modal Alignment
SGCMA transfers an HMM state-transition prior learned from text into time series patches, then aligns patch embeddings to language tokens in each state, enabling a frozen GPT-2 to forecast as well as or better than tu...
-
IISE PG&E Energy Analytics Challenge 2025: Hourly-Binned Regression Models Beat Transformers in Load Forecasting
On the ESD 2025 PG&E dataset, hourly-binned XGBoost models with PCA weather covariates achieved lower MAPE than transformer, LSTM, TFT, and TimeGPT models in day-ahead annual load forecasting.
-
Foundation Time-Series AI Model for Realized Volatility Forecasting
Incremental fine-tuning of the TimesFM foundation model improves one-day-ahead realized volatility forecasts and beats HAR, ARFIMA, CHAR, and RGARCH benchmarks on average losses across 21 global equity indices.
-
How Effective are Large Time Series Models in Hydrology? A Study on Water Level Forecasting in Everglades
Chronos, a pretrained foundation model used zero-shot, beat 16 other models at forecasting Everglades water levels across five stations and horizons from 7 to 28 days.
-
Benchmarking Time Series Forecasting Models: From Statistical Techniques to Foundation Models in Real-World Applications
Gradient-boosted ML models achieved the best accuracy in a 4-restaurant hourly sales forecasting benchmark, with zero-shot Chronos-Bolt foundation models competitive in 3 of 4 cases.
-
Using Causality for Enhanced Prediction of Web Traffic Time Series
A causal cross-mapping module claims to improve web traffic prediction, but the theoretical justification is missing and the implementation deviates from the underlying CCM theory.
-
BEAT: Balanced Frequency Adaptive Tuning for Long-Term Time-Series Forecasting
BEAT adaptively scales gradients of frequency-specific networks during training to balance learning speeds, with reported gains on some long-term forecasting benchmarks.
-
PaMMA-Net: Plasmas magnetic measurement evolution based on data-driven incremental accumulative prediction
PaMMA-Net predicts tokamak magnetic measurements by learning to forecast their increments rather than absolute values, using a transformer decoder and spectrogram-based data augmentation, and outperforms generic time-...
-
Uncertainty-Aware Digital Twins: Robust Model Predictive Control using Time-Series Deep Quantile Learning
A robust MPC framework that uses one-shot multi-step TiDE predictions and learned quantile bounds as safety tubes, demonstrated on a DED additive-manufacturing simulator.
-
Interpretable deep convolutional model for nonlinear multivariate time series in complex systems
DCIts is a convolutional model whose per-sample transition tensor recovers signed, lag-resolved causal coefficients matching the ground-truth generators of eight synthetic multivariate time series.
-
Cherry-Picking in Time Series Forecasting: How to Select Datasets to Make Your Model Shine
Judiciously choosing just four datasets can make 46% of forecasting models appear best-in-class and 77% top-three, so dataset selection alone can distort reported performance.
-
Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions
A large empirical study finds that off-the-shelf RNNs are competitive with but not superior to ETS and ARIMA, and proposes a stacked LSTM plus COCOB configuration as a good default.
-
PhysAttNet: Enhancing Predictive Performance in Industrial and Astrophysical Time Series via Physics-Informed Attention
A CNN forecaster with attention regularized toward smooth peaks shows small gains on blazar flare forecasting, but its sparsity term is constant under softmax and its claimed broad accuracy gains are unsupported.
-
Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail
A smooth-baseline-plus-sparse-shock decomposition lowers forecast variance and safety stock but increases stockout costs, reducing total inventory cost only when holding costs exceed ~20% of stockout costs.
-
Towards Reliable Zero-Shot Crowd Forecasting: Evaluating Time Series Foundation Models for Special Event Pedestrian Forecasting
Zero-shot time-series foundation models, especially Chronos-2 with increasing context, can produce probabilistically reliable crowd-flow forecasts up to about 30–45 minutes ahead for a five-day special event, though t...
-
Cross-device Zero-shot Label Transfer via Alignment of Time Series Foundation Model Embeddings
A framework using adversarial alignment of frozen time-series foundation model embeddings transfers labels to a simulated target domain, but the target is synthetic and no real consumer device data is tested.
-
N-BEATS-MOE: N-BEATS with a Mixture-of-Experts Layer for Heterogeneous Time Series Forecasting
Adding a gating network on top of N-BEATS block outputs gives modest SMAPE improvements on some heterogeneous benchmark series, but the gains are small and not statistically validated.
-
Towards Measuring and Modeling Geometric Structures in Time Series Forecasting via Image Modality
A new image-based similarity metric (TGSI) and a three-part training loss (SATL) that together aim to improve the geometric fidelity of time series forecasts.
-
MamNet: A Novel Hybrid Model for Time-Series Forecasting and Frequency Pattern Analysis in Network Traffic
MamNet combines Mamba time-domain modeling with Fourier frequency features and claims 2-4% gains over five baselines on UNSW-NB15 and CAIDA.
-
FAF: A Feature-Adaptive Framework for Few-Shot Time Series Forecasting
A feature-adaptive meta-learning framework for few-shot time series forecasting reports large gains, but its evaluation uses one to nine test tasks per dataset, lacks error bars, and contains numerical and preprocessi...
-
PPTNet: A Hybrid Periodic Pattern-Transformer Architecture for Traffic Flow Prediction and Congestion Identification
PPTNet forecasts highway density and speed using FFT-selected periodic patterns, 2D Inception convolutions, and a Transformer decoder, then converts forecasts into congestion probabilities with a Mamdani fuzzy system.
-
Quantifying Cryptocurrency Unpredictability: A Comprehensive Study of Complexity and Forecasting
Major cryptocurrency price series are indistinguishable from Brownian noise in complexity analysis, and no forecasting model beats a naive random-walk baseline.
-
STAN: Smooth Transition Autoregressive Networks
A STAR-inspired gated neural network shows small short-horizon RMSE gains on PJM hourly load data, but the claimed advantage over STAR itself is never tested.
-
An Investigation into Seasonal Variations in Energy Forecasting for Student Residences
No single ML model wins year-round for student-residence energy forecasting; the proposed hypernetwork-LSTM and MiniAutoEncXGBoost do well on specific seasons and buildings.
-
Forecasting Anonymized Electricity Load Profiles
Microaggregating smart meter data before forecasting leaves aggregated load forecasts accurate or improves them, with volatility dropping as group size k increases.
-
CORAL: Concept Drift Representation Learning for Co-evolving Time-series
CORAL learns block-diagonal kernel self-representation matrices per time window to identify, track, and forecast concept drift in co-evolving time series, with modest reported RMSE gains over baselines.
-
CryptoMamba: Leveraging State Space Models for Accurate Bitcoin Price Prediction
CryptoMamba applies a Mamba state space model to Bitcoin price prediction, reporting better test-set RMSE and trading returns than baselines, but the generalization claim is undermined by validation results and missin...
-
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
A simple gradient-free 'directional gradient approximation' attack makes LLM time series forecasters degrade more than equivalent random noise, across GPT-3.5, GPT-4, LLaMa, Mistral, TimeGPT, and TimeLLM.
-
Wavelet-Enhanced Neural ODE and Graph Attention for Interpretable Energy Forecasting
The proposed hybrid forecaster claims consistent superiority on ETT and EIA energy datasets, but the benchmark setup and reported numbers do not support that claim.
-
A Review of the Long Horizon Forecasting Problem in Time Series Analysis
A survey of long-horizon forecasting with new ETTm2 ablations showing per-timestep error growth that is absent for xLSTM and Triformer.
-
Entanglement for Pattern Learning in Temporal Data with Logarithmic Complexity: Benchmarking on IBM Quantum Hardware
A fixed, untrained 10-qubit circuit with CNOT entanglement forecasts Z500 weather data with MSE near classical AR models, but the claimed logarithmic training complexity is an artifact of defining the window size as l...
-
Hierarchical Forecast Reconciliation on Networks: A Network Flow Optimization Formulation
FlowRec recasts reconciliation as flow optimization, but key theorems are wrong or restate known projection results, so the advertised O(n^2 log n) and dynamic-update guarantees are unsupported.
-
Zero Shot Time Series Forecasting Using Kolmogorov Arnold Networks
A KAN-based doubly residual N-BEATS model with adversarial domain adaptation is claimed to improve zero-shot day-ahead electricity price forecasts on Nord Pool by 13% over N-BEATS and 24% over KAN.
-
EDformer: Embedded Decomposition Transformer for Interpretable Multivariate Time Series Predictions
EDformer combines moving-average decomposition with an iTransformer-style variate-token encoder and claims state-of-the-art forecasting, but its reported benchmark results do not consistently support that claim.
Discussion (0). Continue with ORCID to comment.