REVIEW 3 major objections 5 minor 1 cited by
TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read TAT is a transformer forecaster built around Temporal Alignment Attention, which aligns demand series with known holiday and promotion calendars in both lookback and horizon; the paper reports up to 30% accuracy improvement on peak demand…
desk verdict A plausible new attention module with real gains on short-horizon peaks, but the 'up to 30%' headline is measured against a weak baseline and the strongest comparisons show at most ~19% — still worth a serious referee, not a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Temporal Alignment Attention (TAA) is a multi-head scaled dot-product attention mechanism with deliberately different query, key, and value choices than vanilla self-attention. In the encoder, Q is the embedded observed demand sequence, K is the concatenation of embedded historical context and broadcast static metadata, and V includes demand, context, and static features, producing an L-by-L alignment score matrix. In the decoder, Q is the decoder initialization sequence, K is the concatenation of embedded known future context and static metadata, and V includes observed and future context, producing an H-by-H alignment matrix. The mechanism lets the network learn to assign larger weights to positions where demand peaks and promotional signals coincide, which the paper argues is exactly what generic self-attention and the temporally compressed attention of MQTransformer fail to do.
What would settle it
A reader could settle the claim by retraining MQCNN and MQTransformer with the same 104-week lookback as TAT and recomputing Event A P50 and P90; if their peak-event errors drop to TAT's level, the improvement is not attributable to temporal alignment. Running the same peak-date quantile-loss comparison on a public dataset with a known promotion calendar would also test whether the result is specific to the two proprietary datasets.
Extended reading notes
Core claim
The paper's central claim is that peak demand forecasting improves when attention is context-dependent and temporally aligned rather than generic self-attention. In TAT's encoder, the embedded observed demand sequence serves as the query and the embedded historical holiday/promotion features, with broadcast static metadata, serve as the key, so attention scores emphasize time positions where demand spikes and event indicators coincide; in the decoder, the forecast representation is the query and the known future context is the key, aligning each horizon with the relevant upcoming event. This design is paired with a transposed self-attention step that turns encoder tokens into a decoder initialization and an MLP posterior calibration that rescales decoder outputs from future context. Empirically, TAT reports peak-event improvements such as Region 1 Event A P90 error falling from 0.454 for TFT to 0.309, with P50 also improved, while overall P50 and P90 remain competitive on both datasets. The ablation identifies Temporal Alignment Attention as the main driver: removing it raises Event A P50 error by about 15% and P90 error by about 40%.
Load-bearing premise
The load-bearing premise is that comparing TAT at a 104-week lookback against MQCNN and MQTransformer at 208 weeks is fair; if the longer history materially helps those baselines, the peak-forecast advantage attributed to TAA would be overstated.
Editorial extensions
If this is right
- Peak-event forecasts for short-horizon events improve most, with the paper reporting P50 error reductions of up to roughly 10% and P90 reductions of up to roughly 20% relative to the strongest baselines on Event A.
- The peak gains do not come at the cost of ordinary-period accuracy: TAT remains competitive or slightly leading on overall all-horizon P50 and P90 on both datasets.
- Removing Temporal Alignment Attention is the most damaging ablation, raising Event A P50 error by about 15% and P90 error by about 40%, so the alignment mechanism is the main driver of the reported improvement.
- Because the context features are known in advance, TAT can be used operationally wherever promotion and holiday calendars are set ahead of time, not only in the retrospective evaluation setting.
- The posterior calibration module gives a concrete mechanism for rescaling peak-period predictions from future context features, which could be applied to other forecasters as a standalone refinement step.
Reading between the lines
- Editorial extension: TAA is domain-agnostic, so a natural test is applying it to other settings with known event calendars, such as energy demand around temperature alerts, travel bookings around holidays, or store traffic around local promotions.
- Editorial extension: the paper justifies but does not ablate the lookback asymmetry, where TAT sees 104 weeks while MQCNN and MQTransformer see 208; a controlled same-lookback comparison would separate the effect of TAA from the effect of ignoring distant history.
- Editorial extension: inspecting the learned posterior calibration multipliers across event types could reveal whether TAT learns interpretable event-specific scaling factors, something the paper does not report.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Temporal-Aligned Transformer (TAT), a multi-horizon demand forecasting model whose key component is Temporal Alignment Attention (TAA), an attention mechanism that uses known context features (holidays, promotions) as keys while querying target demand series, in both an encoder and a decoder. A posterior calibration MLP rescales decoder outputs based on future context. The model is evaluated on two proprietary e-commerce demand datasets against eight baselines (RNN, N-BEATS, PatchTST, iTransformer, MQCNN, TFT, MQT, TSMixer), reporting both overall quantile losses (Table 2) and peak-event target-date losses for two annual events (Table 1). The central claim is that TAT brings 'up to 30% accuracy improvement on peak demand forecasting' while maintaining competitive overall performance.
Significance. If its claims held, TAT would be a meaningful step for industrial peak-demand forecasting, a practical problem where exogenous event information is available in advance. The TAA idea is simple and plausible: using known context as keys for attention over the target series is a natural way to align peaks with promotional drivers. The paper also has strengths: evaluation on two large-scale industrial datasets, comparison against a broad baseline set, multiple random seeds with reported standard deviations for the closest baselines, and an ablation over three components. However, the headline result as stated is not supported by the paper's own tables. The maximum improvement over the strongest baseline is about 19%, not 30%, and the model is worse than the strongest baselines on one of the eight peak-event cells. The scientific contribution is therefore a potentially useful architectural variant with a modest and uneven advantage, not the SOTA magnitude advertised in the abstract. The comparison is also confounded by different lookback windows across models.
major comments (3)
- [Abstract, §1, Table 1, §4.2.1] The abstract and Section 1 claim 'up to 30% accuracy improvement' over state-of-the-art methods, but Table 1 does not support this against the strongest baselines. The largest gain over the strongest baseline is 18.7% (Region 1, Event A, P90: TAT 0.309 vs. MQT 0.380, i.e., 1 − 0.309/0.380); the 30% figure is only realized against TFT (0.454 → 0.309 = 31.9%), which is itself outperformed by MQT (0.380) and TSMixer (0.445) in that cell. Furthermore, on Region 2, Event B, P90, TAT (0.659) is worse than both MQCNN (0.641) and MQT (0.639). The claim in §4.2.1 of 'up to 10% lower P50 error and 20% lower P90 error' for Event A also does not match the best-vs-strongest-baseline numbers (8.5% and 18.7%, respectively). The headline and contribution claims must be rewritten to state the maximum improvement against the strongest baseline and to acknowledge the negative cell.
- [Appendix B.2, Table 1] The comparison between TAT and MQCNN/MQT is confounded by lookback length: TAT uses L=104 while MQCNN and MQT use L=208, and the paper provides only a qualitative justification for this asymmetry. Since the paper also asserts that a shorter lookback is preferred for sequential models, the reported gains could in principle be driven by input length rather than by TAA. I request a controlled experiment (e.g., TAT with L=208, or MQCNN/MQT with L=104), or, failing that, a clear statement that the SOTA claim is conditional on the model-specific lookback settings.
- [§4.3, Figure 4] The ablation results for removing TAA are reported inconsistently: the text says TAT\TAA has 'approximately 15% higher P50 error and 40% higher P90 error' on Event A, while the Figure 4 caption says 'incremental improvements of 12% in P50 and 26% in P90.' Because this ablation is the primary evidence that TAA is the driving component, the numbers must be reconciled, and the figure should display the actual values so the claim can be verified.
minor comments (5)
- [§3.2.4, Eq. (4)] The decoder is called with 𝑒Dec[𝐿], but the decoder input was defined as 𝑒Dec[𝐻] in Eq. (3); this appears to be a typo.
- [§2] The expression for the observed window is malformed: '𝒙𝑏 𝑡−𝐿:𝑡 =[𝑥𝑏 𝑡−𝐿,...,𝑥 𝑏 𝐿]' should be '[𝑥𝑏_{𝑡−𝐿}, ..., 𝑥𝑏_𝑡]'.
- [Appendix B.1, Table 3] 'Lookback Window 208' is listed as a dataset statistic, but Appendix B.2 explains that lookback is a model-specific choice; Table 3 should be corrected or clarified.
- [Appendix C] Standard deviations are reported only for MQCNN, MQT, TSMixer, and TAT; the paper should state whether the differences in Table 1, especially the Region 2, Event B, P90 cell where TAT is worse, are statistically significant or within noise, rather than relying on visual inspection.
- [§4.1.4, Appendix B.2] The paper reports results averaged over three independent runs, but no code or data release is mentioned; providing a reproducibility statement or public baseline code would strengthen the evaluation, which is otherwise limited to proprietary data.
Circularity Check
No circularity: TAT is an empirical model evaluation with test-set benchmarks, and no fitted parameter or definition is recycled as a prediction.
full rationale
The paper contains no derivation chain in which an output is defined in terms of the input or in which a fitted parameter is renamed as a prediction. TAT is an encoder-decoder with Temporal Alignment Attention and a posterior-calibration MLP; all components are trained by the quantile loss of Eq. (5) on training data and evaluated on held-out test periods (Section 4.2). The calibration layer y_hat = y_Dec * (1 + Calib(x_c[H])) in Eq. (4) takes known future context as input and is learned jointly; it is not fit to event-date test labels, so the peak-forecast results are genuine test-set measurements. TAA attention in Eqs. (1)-(2) is an architectural choice, not an equation that assumes the conclusion. Self-citations to SPADE [35], MQRetNN [37], GEANN [38], and LLMForecaster [40] appear only as related work and do not provide the load-bearing justification for TAT. The asymmetric lookback (L=104 for TAT vs L=208 for MQCNN/MQT, Section B.2) is an experimental-design choice that could affect fairness but is not circularity. The abstract's 'up to 30%' improvement is not consistently supported by Table 1 against the strongest baselines (e.g., 18.7% vs MQT on Region 1 Event A P90; TAT is worse than MQCNN/MQT on Region 2 Event B P90), but this is a claim-support issue, not a circular reduction. No circular step can be exhibited.
Assumptions & free parameters
assumptions (5)
- standard math Multi-head scaled dot-product attention (Vaswani et al., 2017) is a suitable mechanism for temporal alignment.
- standard math Quantile loss minimization yields calibrated prediction intervals.
- domain assumption Known future context x_c (holidays, promotions, discounts) is deterministic and available for the entire forecasting horizon.
- domain assumption Demand peaks are primarily driven by the known context features, so aligning to x_c is sufficient for peak accuracy.
- domain assumption The test-year distribution resembles the training-year distribution for the same products and regions.
Cite this review
Pith. "Pith review of TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting." pith.science (2026). https://pith.science/paper/BFECK2XY
@misc{pith2026250710349,
author = {Pith},
title = {Pith review of: TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting},
year = {2026},
howpublished = {\url{https://pith.science/paper/BFECK2XY}},
note = {Machine review of arXiv:2507.10349}
}
read the original abstract
Multi-horizon time series forecasting has many practical applications such as demand forecasting. Accurate demand prediction is critical to help make buying and inventory decisions for supply chain management of e-commerce and physical retailers, and such predictions are typically required for future horizons extending tens of weeks. This is especially challenging during high-stake sales events when demand peaks are particularly difficult to predict accurately. However, these events are important not only for managing supply chain operations but also for ensuring a seamless shopping experience for customers. To address this challenge, we propose Temporal-Aligned Transformer (TAT), a multi-horizon forecaster leveraging apriori-known context variables such as holiday and promotion events information for improving predictive performance. Our model consists of an encoder and decoder, both embedded with a novel Temporal Alignment Attention (TAA), designed to learn context-dependent alignment for peak demand forecasting. We conduct extensive empirical analysis on two large-scale proprietary datasets from a large e-commerce retailer. We demonstrate that TAT brings up to 30% accuracy improvement on peak demand forecasting while maintaining competitive overall performance compared to other state-of-the-art methods.
Figures
Forward citations
Cited by 1 Pith paper
-
Burst Aware Forecasting of User Traffic Demand in LEO Satellite Networks
A transformer with burst-distance embeddings, auxiliary burst decoders, and asymmetric loss cuts burst MSE by up to 94% on synthetic LEO traffic.
Reference graph
Works this paper leans on
-
[1]
M Zied Babai, John E Boylan, and Bahman Rostami-Tabar. 2022. Demand fore- casting in supply chains: a review of aggregation and hierarchical approaches. International Journal of Production Research 60, 1 (2022), 324–348
work page 2022
-
[2]
Joos-Hendrik Böse, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Dustin Lange, David Salinas, Sebastian Schelter, Matthias Seeger, and Yuyang Wang. 2017. Probabilistic demand forecasting at scale. Proceedings of the VLDB Endowment 10, 12 (2017), 1694–1705
work page 2017
-
[3]
Frank Chen, Zvi Drezner, Jennifer K Ryan, and David Simchi-Levi. 2000. Quan- tifying the bullwhip effect in a simple supply chain: The impact of forecasting, lead times, and information. Management science 46, 3 (2000), 436–443
work page 2000
-
[4]
Chen, Lee Dicker, Carson Eisenach, and Dhruv Madeka
Kevin C. Chen, Lee Dicker, Carson Eisenach, and Dhruv Madeka. 2022. MQ- Transformer: Multi-horizon forecasts with context dependent attention and optimal bregman volatility. In KDD 2022, Vol. Workshop on Mining and Learning from Time Series – Deep Forecasting: Models, Interpretability, and Applica- tions. Association for Computing Machinery, New York, NY,...
work page 2022
-
[5]
Si-An Chen, Chun-Liang Li, Sercan O Arik, Nathanael Christian Yoder, and Tomas Pfister. 2023. TSMixer: An All-MLP Architecture for Time Series Forecast-ing. Transactions on Machine Learning Research - (2023), –. https://openreview.net/ forum?id=wbpxTuXgm0
work page 2023
-
[6]
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. In 28th International Conference on Neural Information Processing Systems (NIPS 2014). Curran Associates Inc., Montreal, Canada
work page 2014
-
[7]
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2023. A decoder-only foundation model for time-series forecasting. arXiv preprint arXiv:2310.10688 (2023)
arXiv 2023
-
[8]
Wei Fan, Pengyang Wang, Dongkun Wang, Dongjie Wang, Yuanchun Zhou, and Yanjie Fu. 2023. Dish-ts: a general paradigm for alleviating distribution shift in time series forecasting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37. 7522–7529
work page 2023
Show all 45 references
-
[9]
Clive William John Granger and Paul Newbold. 2014. Forecasting economic time series. Academic Press, United States
2014
-
[10]
Olivares
Rob J Hyndman, George Athanasopoulos, Azul Garza, Cristian Challu, Max Mergenthaler, and Kin G. Olivares. 2024. Forecasting: Principles and Prac- tice, the Pythonic Way . OTexts, Melbourne, Australia. available at https://otexts.com/fpppy/
2024
-
[11]
Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, et al . 2024. Time- LLM: Time series forecasting by reprogramming large language models. In 12th International Conference on Learning Representations...
2024
-
[12]
Harshavardhan Kamarthi and B Aditya Prakash. 2023. Large Pre-trained time series models for cross-domain Time series analysis tasks. arXiv preprint arXiv:2311.11413 (2023)
2023 arXiv
-
[13]
Taesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park, Jang-Ho Choi, and Jaegul Choo. 2021. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International conference on learning representations
2021
-
[14]
Guokun Lai, Wei-Cheng Chang, Yiming Yang, and Hanxiao Liu. 2018. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval. Association for Computing Machinery, An...
2018
-
[15]
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. 2019. Enhancing the locality and breaking the memory bottle- neck of transformer on time series forecasting. Advances in neural information processing systems 32 (2019)
2019
-
[16]
Bryan Lim, Sercan Ö Arık, Nicolas Loeff, and Tomas Pfister. 2021. Temporal fusion transformers for interpretable multi-horizon time series forecasting.International Journal of Forecasting 37, 4 (2021), 1748–1764
2021
-
[17]
Haoxin Liu, Harshavardhan Kamarthi, Lingkai Kong, Zhiyuan Zhao, Chao Zhang, and B Aditya Prakash. 2024. Time-series forecasting for out-of-distribution generalization using invariant learning. arXiv preprint arXiv:2406.09130 (2024)
2024 arXiv
-
[18]
Haoxin Liu, Chenghao Liu, and B Aditya Prakash. 2024. A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization. arXiv preprint arXiv:2411.06018 (2024)
2024 arXiv
-
[19]
Haoxin Liu, Shangqing Xu, Zhiyuan Zhao, Lingkai Kong, Harshavardhan Ka- marthi, Aditya B Sasanur, Megha Sharma, Jiaming Cui, Qingsong Wen, Chao Zhang, et al. 2024. Time-MMD: Multi-Domain Multimodal Dataset for Time Series Analysis. In Thirty-Eighth Annual Conference on Neural ...
2024
-
[20]
Aditya Prakash
Haoxin Liu, Zhiyuan Zhao, Jindong Wang, Harshavardhan Kamarthi, and B. Aditya Prakash. 2024. LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting. InFindings of the As- sociation for Computational Linguistics: ACL 2024 , Lun-Wei Ku...
2024 doi
-
[21]
Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long. 2024. iTransformer: Inverted transformers are effective for time series forecasting. In 12th International Conference on Learning Representations . ICLR 2024, Vienna, Austria, –. https://op...
2024
-
[22]
Zhiding Liu, Mingyue Cheng, Zhi Li, Zhenya Huang, Qi Liu, Yanhu Xie, and Enhong Chen. 2023. Adaptive normalization for non-stationary time series fore- casting: A temporal slice perspective. Advances in Neural Information Processing Systems 36 (2023), 14273–14292
2023
-
[23]
Sarabeth M Mathis, Alexander E Webber, Tomás M León, Erin L Murray, Mon- ica Sun, Lauren A White, Logan C Brooks, Alden Green, Addison J Hu, Roni Rosenfeld, et al. 2024. Title evaluation of FluSight influenza forecasting in the 2021–22 and 2022–23 seasons with a new target lab...
2024
-
[24]
Elizbar A Nadaraya. 1964. On estimating regression. Theory of Probability & Its Applications 9, 1 (1964), 141–142
1964
-
[25]
Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In International Conference on Learning Representations (ICLR 2023) . ICLR, Kigali Rwanda
2023
-
[26]
Olivares, Cristian Challu, Grzegorz Marcjasz, Rafał Weron, and Artur Dubrawski
Kin G. Olivares, Cristian Challu, Grzegorz Marcjasz, Rafał Weron, and Artur Dubrawski. 2023. Neural basis expansion analysis with exogenous variables: Forecasting electricity prices with NBEATSx. International Journal of Forecasting 39, 2 (2023), 884–900. https://doi.org/10.10...
2023 doi
-
[27]
Olivares, O
Kin G. Olivares, O. Nganba Meetei, Ruijun Ma, Rohan Reddy, Mengfei Cao, and Lee Dicker. 2024. Probabilistic hierarchical forecasting with deep Poisson mixtures. International Journal of Forecasting 40, 2 (2024), 470–489. https: //doi.org/10.1016/j.ijforecast.2023.04.007
2024 doi
-
[28]
Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio
Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio. 2020. N- BEATS: Neural basis expansion analysis for interpretable time series forecasting. In 8th International Conference on Learning Representations . ICLR 2020, Addis Ababa, Ethiopia, –. https://openrevi...
2020
-
[29]
David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. 2020. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. Inter- national journal of forecasting 36, 3 (2020), 1181–1191
2020
-
[30]
Alex J Smola and Bernhard Schölkopf. 2004. A tutorial on support vector regres- sion. Statistics and computing 14, 3 (2004), 199–222
2004
-
[31]
Senior, and Koray Kavukcuoglu
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016. WaveNet: A Generative Model for Raw Audio. CoRR abs/1609.03499 (2016). arXiv:1609.03499 http://arxiv.org/abs/1609.03499
2016 arXiv
-
[32]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...
2017
-
[33]
Ruofeng Wen, Kari Torkkola, Balakrishnan Narayanaswamy, and Dhruv Madeka
-
[34]
Christopher Williams and Carl Rasmussen. 1995. Gaussian processes for regres- sion. Neural information processing systems (NIPS 1995) 8 (1995), - pages
1995
-
[35]
Olivares, Boris Oreshkin, Sunny Ruan, Sitan Yang, Abhi- nav Katoch, Shankar Ramasubramanian, Youxin Zhang, Michael W
Malcolm Wolff, Kin G. Olivares, Boris Oreshkin, Sunny Ruan, Sitan Yang, Abhi- nav Katoch, Shankar Ramasubramanian, Youxin Zhang, Michael W. Mahoney, Dmitry Efimov, and Vincent Quenneville-Bélair. 2024. ♠ SPADE♠ Split Peak Attention DEcomposition. In Thirty-Eighth Annual Confer...
2024 arXiv
-
[36]
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: De- composition transformers with auto-correlation for long-term series forecasting. Advances in neural information processing systems 34 (2021), 22419–22430
2021
-
[37]
Sitan Yang, Carson Eisenach, and Dhruv Madeka. 2022. MQRetNN: Multi-Horizon Time Series Forecasting with Retrieval Augmentation. arXiv:2207.10517 [cs.LG] https://arxiv.org/abs/2207.10517
2022 arXiv
-
[38]
Sitan Yang, Malcolm Wolff, Shankar Ramasubramanian, Vincent Quenneville- Belair, Ronak Metha, and Michael W. Mahoney. 2023. GEANN: Scal- able Graph Augmentations for Multi-Horizon Time Series Forecasting. arXiv:2307.03595 [cs.LG] https://arxiv.org/abs/2307.03595
2023 arXiv
-
[39]
Hsiang-Fu Yu, Nikhil Rao, and Inderjit S Dhillon. 2016. Temporal regularized matrix factorization for high-dimensional time series prediction. Advances in neural information processing systems 29 (2016)
2016
-
[40]
Hanyu Zhang, Chuck Arvin, Dmitry Efimov, Michael W Mahoney, Dominique Perrault-Joncas, Shankar Ramasubramanian, Andrew Gordon Wilson, and Mal- colm Wolff. 2024. LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data. InThirty-Eighth Annual Conference ...
2024 arXiv
-
[41]
Aditya Prakash
Zhiyuan Zhao, Haoxin Liu, Alexander Rodriguez, and B. Aditya Prakash. 2025. Performative Time-Series Forecasting. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
2025
-
[42]
Zhiyuan Zhao, Juntong Ni, Shangqing Xu, Haoxin Liu, Wei Jin, and B Aditya Prakash. 2025. TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level Effectiveness. arXiv preprint arXiv:2506.06482 (2025)
2025
-
[43]
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond efficient transformer for long se- quence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35-12. Association fo...
2021
-
[44]
Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International conference on machine learning . PMLR, Baltimore, Maryland USA, 27268–27286. A Related Work Time...
2022
-
[2017]
In 31st Conference on Neural Information Processing Systems NIPS 2017 , Vol
A Multi-Horizon Quantile Recurrent Forecaster. In 31st Conference on Neural Information Processing Systems NIPS 2017 , Vol. Time Series Workshop. Curran Associates Inc., Long Beach California USA, –. arXiv:1711.11053 [stat.ML] https://arxiv.org/abs/1711.11053
2017 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.