Pith. sign in

REVIEW 2 major objections 5 minor 54 references

Forecasts get better when uncertainty is allowed to remember and evolve over time, not treated as independent noise at each step.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 18:57 UTC pith:VJUBJBWS

load-bearing objection Solid empirical methods paper: recurrent volatility-state transfer inside a non-AR forecasting VAE improves CRPS/NMAE and calibration on nine benchmarks, with clean ablations and transferable plug-ins. the 2 major comments →

arxiv 2603.24254 v3 pith:VJUBJBWS submitted 2026-03-25 cs.LG cs.AI

Beyond Static Uncertainty: Modeling Temporal Uncertainty Dynamics for Probabilistic Time Series Forecasting

classification cs.LG cs.AI
keywords probabilistic time series forecastingtemporal uncertainty dynamicsheteroscedasticityvariational autoencodervolatility modelinglocation-scale decoderinverse-variance weighting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Real-world time series do not have random scatter that restarts every moment. Volatility clusters, lingers, and jumps at structural breaks. Most probabilistic forecasters still treat the width of each confidence interval as a local, memoryless quantity. This paper argues that is a real modeling gap: both the training objective (often plain MSE) and the architecture leave the evolution of uncertainty under-specified. The authors formalize temporal uncertainty dynamics and build VolDy-VAE, a non-autoregressive VAE whose decoder has a mean path and a recurrent scale path that carries a volatility hidden state from the look-back window into the future. High-volatility points then down-weight the mean update while their uncertainty is kept in the scale, and a simple regime-switching argument shows that a well-estimated scale recovers efficient inverse-variance weighting. On nine standard benchmarks the model improves both distributional scores and point accuracy while staying fast to run; the same idea also lifts GAN, Koopman-VAE, and Transformer backbones when plugged in.

Core claim

Modeling how predictive uncertainty evolves and persists across time—via a recurrent volatility state transferred into a non-autoregressive location-scale decoder—improves both probabilistic calibration and point forecasts relative to methods that treat scale as independent per step or train under a fixed-variance MSE assumption.

What carries the argument

VolDy-VAE’s dual-head decoder: a location MLP for the mean, plus a GRU-based scale path that maintains and transfers a volatility hidden state from look-back to forecast horizon, trained under heteroscedastic Gaussian negative log-likelihood so that gradients on the mean are attenuated by 1/σ^{2}.

Load-bearing premise

The claimed efficiency gain holds only when the recurrent scale head correctly recovers the true or consistently estimated time-varying variances; if it mis-tracks regime boundaries or oversmooths, inverse-variance weighting and adaptive attenuation no longer deliver the promised improvement.

What would settle it

On a controlled regime-switching series with known ground-truth σ_t, check whether the predicted scale still tracks the true volatility trajectory with high lag-1 coherence and whether CRPS and NMAE still beat an otherwise identical feed-forward (memoryless) scale head; if the recurrent head loses coherence or the gains vanish, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Probabilistic forecasters should treat volatility as a state that is transferred across the look-back/prediction boundary, not only as a per-step output head.
  • Replacing MSE with a location-scale NLL objective can improve both uncertainty calibration and point accuracy when scales are well estimated.
  • The same dual-path idea can be dropped into existing GAN, VAE, and Transformer forecasting backbones without redesigning the whole model.
  • Direct (non-autoregressive) generation remains compatible with temporally coherent uncertainty, so long latency of diffusion or recursive decoding is not required for better calibration.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If volatility memory is the missing piece, multi-series settings with contagion or shared shocks may need a shared or cross-variable volatility state, not only per-channel recurrence.
  • The same inverse-variance attenuation logic could improve robust training in other sequence tasks (imputation, anomaly detection) where outliers currently dominate MSE gradients.
  • Abrupt single-step regime flips remain a stress case; hybrid designs that combine the GRU path with an explicit switch indicator would be a natural next experiment.
  • Because the method is architecture-agnostic at the principle level, library-level location-scale modules could become a default add-on for long-horizon forecasting stacks.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper argues that many probabilistic time series forecasters treat predictive scale as a memoryless per-step quantity and therefore under-model volatility clustering and regime persistence. It formalizes temporal uncertainty dynamics and proposes VolDy-VAE: a non-autoregressive VAE whose decoder couples a location head with a GRU scale path that transfers a volatility hidden state from the look-back window into the forecast horizon, trained with heteroscedastic Gaussian NLL. A simplified regime-switching analysis shows that, when scales are known or consistently estimated, this objective reduces to inverse-variance weighting and is more efficient than MSE. Empirically, VolDy-VAE improves CRPS and NMAE over twelve baselines on nine ProbTS-style benchmarks (six seeds, Wilcoxon tests), with architecture ablations (MLP/LSTM/GRU scale heads), synthetic volatility recovery, efficiency measurements, and plug-in variants on TimeGAN, K2VAE, and PatchTST.

Significance. If the results hold, the paper offers a practical and transferable design principle for probabilistic forecasting: dedicated volatility-state transfer plus heteroscedastic NLL can improve both calibration and point accuracy without iterative sampling. Strengths include a carefully scoped statistical rationale (Appendix A-D / Theorem A.1), synthetic recovery of planted volatility (ρ>0.98), multi-seed evaluation under a public protocol, uncertainty-dynamics diagnostics (coverage, sharpness, QICE, lag-1 scale autocorrelation), and publicly released code. The work is incremental relative to heteroscedastic VAEs and GARCH-style recurrence, but the combination of non-autoregressive horizon generation with look-back-to-horizon volatility-state transfer is a clear, usable contribution for long-horizon PTSF.

major comments (2)
  1. Contribution 3 and §V-C claim that the VolDy principle transfers to TimeGAN, K2VAE, and PatchTST, but Table X shows that the TimeGAN and PatchTST plug-ins primarily replace the objective/head with location-scale Gaussian NLL; they do not clearly retain the recurrent volatility-state transfer that defines VolDy in §III-B3 (Eqs. 6–7 and the look-back→horizon h_vol initialization). Please either (i) implement and report the full recurrent scale path on those backbones, or (ii) restate the plug-in claim as transfer of heteroscedastic location-scale training rather than of temporal volatility dynamics, and keep the stronger claim only for K2VAE / the main model.
  2. The central architectural claim is not only a recurrent scale head, but specifically transfer of the look-back volatility state into the forecast horizon (§III-B3; Fig. 2). Tables VI–VII ablate MLP vs LSTM vs GRU under a shared NLL objective, which isolates temporal memory, but do not isolate state transfer itself (e.g., GRU initialized from the final look-back h_vol versus zero/random re-init on the future path only). A short ablation on ETTh1/ETTh2 would make the transfer mechanism load-bearing rather than inferred from the GRU vs MLP gap alone.
minor comments (5)
  1. Fig. 1 caption refers to panels (a)–(c) for MSE, independent scale, and VolDy, but the body text later says “adaptive attenuation effect (visualized in Fig. 1(b))” when discussing VolDy; align caption and in-text panel references.
  2. Eq. (9) / main-text NLL indexing: clarify that the GRU runs at patch granularity while the NLL is written over time steps (t,c); a one-sentence reshape note would prevent confusion with per-step recurrence.
  3. Table III and the full CRPS/NMAE tables mark some baseline entries as missing (“-”); briefly state whether these are OOM/timeouts or protocol exclusions so readers can interpret the ranking fairly.
  4. Appendix C limitations (abrupt single-step breaks, high-C spillovers) are appropriate; consider cross-referencing them from the main conclusions so the scoped claims of Theorem A.1 are visible without the supplement.
  5. Notation consistency: Z_f uture / Zfuture and VolDy-V AE / VolDy-VAE appear with spacing artifacts in several places; clean for camera-ready.

Circularity Check

0 steps flagged

No significant circularity: inverse-variance reduction is standard WLS under known/consistent scales; empirical claims rest on external benchmarks and ablations.

full rationale

The paper's only formal derivation (Appendix A-D / Theorem A.1) shows that Gaussian NLL with known or consistently estimated time-varying variances reduces to inverse-variance weighted least squares, while MSE remains unbiased but inefficient under regime-switching heteroscedasticity. That reduction is textbook weighted least squares once the scales are treated as given; it does not define the forecasting target from the reported CRPS/NMAE, nor does it smuggle the empirical gains into the architecture by construction. The adaptive-attenuation gradient (Eq. 10 / Prop. 1) likewise follows directly from differentiating the NLL and is not a self-definitional loop. Architecture choices (GRU scale head, RevIN, patching) are justified by ablations against MLP/LSTM alternatives and by ordinary hyperparameter search, not by baking held-out metrics into the model definition. Empirical claims are measured against external baselines (K2VAE, PatchTST, diffusion/flow models, etc.) on nine public benchmarks with multi-seed evaluation; plug-in gains on TimeGAN/K2VAE/PatchTST are likewise comparative. Self-citation is ordinary related-work positioning and is not load-bearing for uniqueness or for the central efficiency claim. Score 1 reflects only the mild, non-load-bearing character of the scoped theoretical rationale (which the paper itself labels as simplified and conditional), not any reduction of the main result to its inputs.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central empirical claim rests on standard VAE and Gaussian-likelihood machinery plus a small set of modeling choices (direct latent projection, single-layer GRU volatility state, RevIN, patching). Free parameters are ordinary architectural and optimization knobs selected by validation. No new physical entities are postulated; 'temporal uncertainty dynamics' is a named modeling principle rather than an unobserved particle or force.

free parameters (5)
  • patch length P = 24
    Fixed at 24 to match K2VAE for fair comparison; sensitivity table shows modest variation around this value.
  • latent / GRU hidden dimension D = 256
    Default 256 chosen as performance-efficiency trade-off on ETTh1 and Weather.
  • MLP encoder depth N = 3
    Default 3 layers; deeper encoders give only marginal CRPS gains.
  • KL weight beta and learning rate
    Standard VAE hyper-parameters tuned in [1e-4, 1e-3] with early stopping; not claimed to be universal.
  • stability epsilon xi in Softplus scale = 1e-6
    Small constant (e.g. 1e-6) to keep scales positive; conventional numerical safeguard.
axioms (5)
  • domain assumption Predictive observations are well-modeled by a diagonal heteroscedastic Gaussian whose mean and scale are functions of a shared latent code.
    Stated in Sec. III-B3 and Eq. (9); likelihood-family ablation later shows richer families did not help under the same budget.
  • ad hoc to paper High-level semantic evolution of the series can be approximated by a single global linear projection in latent space (non-autoregressive latent dynamics).
    Explicit modeling assumption in Sec. III-B2, Eq. (4).
  • domain assumption Volatility regimes are sufficiently persistent that a single-layer GRU can pool neighboring evidence without catastrophic lag at regime boundaries.
    Motivated by classical volatility clustering and used to justify the recurrent scale head (Sec. III-B3 and synthetic experiments).
  • standard math When scales are known or consistently estimated, Gaussian NLL is equivalent to inverse-variance weighted least squares and is more efficient than MSE under heteroscedasticity.
    Theorem A.1 / Appendix A-D; classical weighted-least-squares result specialized to a two-regime constant-mean model.
  • domain assumption Reversible instance normalization removes global level/scale shifts so that the latent and scale modules can focus on local dynamics.
    Adopted from Kim et al.; ablation in Table V shows large degradation when removed.
invented entities (2)
  • temporal uncertainty dynamics (as a named forecasting principle) no independent evidence
    purpose: To name the missing modeling target: evolution and persistence of volatility regimes rather than independent per-step scales.
    Framing device introduced in the abstract and Sec. I; not an unobserved physical quantity.
  • volatility hidden state h_vol transferred across the look-back/forecast boundary independent evidence
    purpose: Provides the recurrent memory that makes predicted scales temporally coherent.
    Architectural construct defined in Eqs. (6)-(7); its utility is tested by MLP vs GRU ablations and synthetic recovery.

pith-pipeline@v1.1.0-grok45 · 40911 in / 3294 out tokens · 37984 ms · 2026-07-13T18:57:26.946197+00:00 · methodology

0 comments
read the original abstract

Real-world time series exhibit temporally structured uncertainty: volatility clusters in turbulent regimes, dissipates in stable periods, and shifts abruptly around structural breaks. Yet many probabilistic forecasting methods estimate predictive uncertainty as an independent per-step quantity, leaving the evolution and persistence of volatility regimes under-modeled. We formalize this missing dimension as temporal uncertainty dynamics and instantiate it in the Volatility Dynamics Variational Autoencoder (VolDy-VAE), a non-autoregressive generative forecaster with a location-scale decoder. VolDy-VAE combines a location path for mean prediction with a recurrent scale path that transfers and evolves a volatility hidden state from the look-back window to the forecasting horizon, enabling temporally coherent predictive variances. This design yields an adaptive attenuation mechanism: high-variance observations receive lower influence on the location estimate while their uncertainty is preserved through explicit scale predictions. We further provide a simplified regime-switching analysis showing that, when variances are known or consistently estimated, the volatility-aware objective reduces to inverse-variance weighting, whereas MSE-based estimators remain unbiased but statistically inefficient. Experiments on nine benchmarks show that VolDy-VAE improves forecasting accuracy and uncertainty calibration over competitive probabilistic and point-forecasting baselines while maintaining low inference latency; plug-in studies further indicate that the VolDy principle can benefit GAN, Koopman VAE, and Transformer backbones. The source code is publicly available at https://github.com/wangyijunlyy/VolDy-VAE.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 3 linked inside Pith

  1. [1]

    Time-series forecasting with deep learning: a survey,

    B. Lim and S. Zohren, “Time-series forecasting with deep learning: a survey,”Philosophical Transactions of the Royal Society A: Mathemati- cal, Physical and Engineering Sciences, vol. 379, no. 2194, p. 20200209, 2021

  2. [2]

    Deep learning for time series forecasting: a survey,

    X. Kong, Z. Chen, W. Liu, K. Ning, L. Zhang, S. Muhammad Marier, Y . Liu, Y . Chen, and F. Xia, “Deep learning for time series forecasting: a survey,”International Journal of Machine Learning and Cybernetics, pp. 1–34, 2025

  3. [3]

    BRITS: Bidirectional Recurrent Imputation for Time Series,

    W. Cao, D. Wang, J. Li, H. Zhou, L. Li, and Y . Li, “BRITS: Bidirectional Recurrent Imputation for Time Series,” inAdvances in Neural Inf. Process. Syst., 2018, pp. 6776–6786

  4. [4]

    CSDI: Conditional Score- based Diffusion Models for Probabilistic Time Series Imputation,

    Y . Tashiro, J. Song, Y . Song, and S. Ermon, “CSDI: Conditional Score- based Diffusion Models for Probabilistic Time Series Imputation,” in Advances in Neural Inf. Process. Syst., 2021, pp. 24 804–24 816

  5. [5]

    SSD-TS: Exploring the Potential of Linear State Space Models for Diffusion Models in Time Series Imputation,

    H. Gao, W. Shen, X. Qiu, R. Xu, B. Yang, and J. Hu, “SSD-TS: Exploring the Potential of Linear State Space Models for Diffusion Models in Time Series Imputation,” inProc. ACM SIGKDD Int. Conf. Knowledge discovery & data mining, 2025, p. 649–660

  6. [6]

    Optimal Transport for Time Series Imputation,

    H. Wang, z. li, H. Li, X. Chen, M. Gong, BinChen, and Z. Chen, “Optimal Transport for Time Series Imputation,” inProc. Int. Conf. Learn. Representations, 2025, pp. 1–25

  7. [7]

    Deep learning for multivariate time series imputation: A survey,

    J. Wang, W. Du, Y . Yang, L. Qian, W. Cao, K. Zhang, W. Wang, Y . Liang, and Q. Wen, “Deep learning for multivariate time series imputation: A survey,”arXiv preprint arXiv:2402.04059, 2024

  8. [8]

    Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy,

    J. Xu, H. Wu, J. Wang, and M. Long, “Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy,” inProc. Int. Conf. Learn. Representations, 2022, pp. 1–20

  9. [9]

    Drift doesn’t matter: Dynamic decomposition with diffusion recon- struction for unstable multivariate time series anomaly detection,

    C. Wang, Z. Zhuang, Q. Qi, J. Wang, X. Wang, H. Sun, and J. Liao, “Drift doesn’t matter: Dynamic decomposition with diffusion recon- struction for unstable multivariate time series anomaly detection,” in Advances in Neural Inf. Process. Syst., 2023, pp. 10 758–10 774

  10. [10]

    The Elephant in the Room: Towards A Reliable Time-Series Anomaly Detection Benchmark,

    Q. Liu and J. Paparrizos, “The Elephant in the Room: Towards A Reliable Time-Series Anomaly Detection Benchmark,” inAdvances in Neural Inf. Process. Syst., 2024, pp. 108 231–108 261

  11. [11]

    A parameter-efficient federated framework for streaming time series anomaly detection via lightweight adaptation,

    H. Miao, R. Xu, Y . Zhao, S. Wang, J. Wang, P. S. Yu, and C. S. Jensen, “A parameter-efficient federated framework for streaming time series anomaly detection via lightweight adaptation,”IEEE Transactions on Mobile Computing, 2025

  12. [12]

    CATCH: Channel-Aware multivariate Time Series Anomaly Detection via Frequency Patching,

    X. Wu, X. Qiu, Z. Li, Y . Wang, J. Hu, C. Guo, H. Xiong, and B. Yang, “CATCH: Channel-Aware multivariate Time Series Anomaly Detection via Frequency Patching,” inProc. Int. Conf. Learn. Representations, 2025, pp. 56 894–56 922

  13. [13]

    Strictly Proper Scoring Rules, Prediction, and Estimation,

    T. Gneiting and A. E. Raftery, “Strictly Proper Scoring Rules, Prediction, and Estimation,”J. Amer. Statist. Assoc., vol. 102, no. 477, pp. 359–378, 2007

  14. [14]

    Probabilistic forecasts, calibration and sharpness,

    T. Gneiting, F. Balabdaoui, and A. E. Raftery, “Probabilistic forecasts, calibration and sharpness,”Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 69, no. 2, pp. 243–268, 2007

  15. [15]

    Probabilistic electric load forecasting: A tutorial review,

    T. Hong and S. Fan, “Probabilistic electric load forecasting: A tutorial review,”Int. J. Forecasting, vol. 32, no. 3, pp. 914–938, 2016

  16. [16]

    An Accurate and Interpretable Framework for Trustworthy Process Monitoring,

    H. Wang, Z. Wang, Y . Niu, Z. Liu, H. Li, Y . Liao, Y . Huang, and X. Liu, “An Accurate and Interpretable Framework for Trustworthy Process Monitoring,”IEEE Trans. Artif. Intell., vol. 5, no. 5, pp. 2241–2252, 2024

  17. [17]

    Solar wind speed prediction via graph attention network,

    Y . Sun, Z. Xie, H. Wang, X. Huang, and Q. Hu, “Solar wind speed prediction via graph attention network,”Space Weather, vol. 20, no. 7, 2022

  18. [18]

    Financial time series forecasting with deep learning: A systematic literature review: 2005–2019,

    O. B. Sezer, M. U. Gudelek, and A. M. Ozbayoglu, “Financial time series forecasting with deep learning: A systematic literature review: 2005–2019,”Appl. Soft Comput., vol. 90, p. 106181, 2020

  19. [19]

    Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting,

    Y . Li, R. Yu, C. Shahabi, and Y . Liu, “Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting,” inProc. Int. Conf. Learn. Representations, 2018, pp. 1–16

  20. [20]

    DGraph: A Large-Scale Financial Dataset for Graph Anomaly Detection,

    X. Huang, Y . Yang, Y . Wang, C. Wang, Z. Zhang, J. Xu, L. Chen, and M. Vazirgiannis, “DGraph: A Large-Scale Financial Dataset for Graph Anomaly Detection,” inAdvances in Neural Inf. Process. Syst., 2022, pp. 22 765–22 777

  21. [21]

    ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Pre- diction Horizons,

    J. Zhang, X. Wen, Z. Zhang, S. Zheng, J. Li, and J. Bian, “ProbTS: Benchmarking Point and Distributional Forecasting across Diverse Pre- diction Horizons,” inAdvances in Neural Inf. Process. Syst., 2024, pp. 48 045–48 082

  22. [22]

    Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series Forecasting,

    K. Rasul, C. Seward, I. Schuster, and R. V ollgraf, “Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series Forecasting,” inProc. Int. Conf. Mach. Learn., 2021, pp. 8857–8868

  23. [23]

    Non-stationary Diffusion for Probabilistic Time Series Forecasting,

    W. Ye, Z. Xu, and N. Gui, “Non-stationary Diffusion for Probabilistic Time Series Forecasting,” inProc. Int. Conf. Mach. Learn., 2025, pp. 1–19

  24. [24]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,”arXiv preprint arXiv:2312.00752, 2023

  25. [25]

    Mamba time series forecasting with uncertainty quantification,

    P. Pessoa, P. Campitelli, D. P. Shepherd, S. B. Ozkan, and S. Press ´e, “Mamba time series forecasting with uncertainty quantification,”Ma- chine Learning: Science and Technology, vol. 6, p. 035012, 2025

  26. [26]

    Auto-Encoding Variational Bayes,

    D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in Proc. Int. Conf. Learn. Representations, 2014, pp. 1–14

  27. [27]

    TimeV AE: A variational auto-encoder for multivariate time series generation,

    A. Desai, C. Freeman, Z. Wang, and I. Beaver, “TimeV AE: A variational auto-encoder for multivariate time series generation,”arXiv preprint arXiv:2111.08095, 2021

  28. [28]

    K 2V AE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting,

    X. Wu, X. Qiu, H. Gao, J. Hu, B. Yang, and C. Guo, “K 2V AE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting,” inProc. Int. Conf. Mach. Learn., 2025, pp. 1–22

  29. [29]

    What uncertainties do we need in bayesian deep learning for computer vision?

    A. Kendall and Y . Gal, “What uncertainties do we need in bayesian deep learning for computer vision?” inAdvances in Neural Inf. Process. Syst., 2017, pp. 5580–5590. ARXIV PREPRINT 12

  30. [30]

    Multi- Scale Attention Flow for Probabilistic Time Series Forecasting,

    S. Feng, C. Miao, K. Xu, J. Wu, P. Wu, Y . Zhang, and P. Zhao, “Multi- Scale Attention Flow for Probabilistic Time Series Forecasting,”IEEE Trans. Knowl. Data Eng., vol. 36, no. 5, pp. 2056–2068, 2024

  31. [31]

    DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks,

    D. Salinas, V . Flunkert, J. Gasthaus, and T. Januschowski, “DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks,”Int. J. Forecasting, vol. 36, no. 3, pp. 1181–1191, 2020

  32. [32]

    GluonTS: Proba- bilistic and neural time series modeling in python,

    A. Alexandrov, K. Benidis, M. Bohlke-Schneider, V . Flunkert, J. Gasthaus, T. Januschowski, D. C. Maddix, S. Rangapuram, D. Salinas, J. Schulz, L. Stella, A. C. T ¨urkmen, and Y . Wang, “GluonTS: Proba- bilistic and neural time series modeling in python,”J. Mach. Learn. Res., vol. 21, no. 116, pp. 1–6, 2020

  33. [33]

    Deep state space models for time series forecasting,

    S. S. Rangapuram, M. W. Seeger, J. Gasthaus, L. Stella, Y . Wang, and T. Januschowski, “Deep state space models for time series forecasting,” inAdvances in Neural Inf. Process. Syst., 2018, pp. 7796–7805

  34. [34]

    Multivariate Time Series Forecasting With Dynamic Graph Neural ODEs,

    M. Jin, Y . Zheng, Y .-F. Li, S. Chen, B. Yang, and S. Pan, “Multivariate Time Series Forecasting With Dynamic Graph Neural ODEs,”IEEE Trans. Knowl. Data Eng., vol. 35, no. 9, pp. 9168–9180, 2023

  35. [35]

    LightCTS*: Lightweight correlated time series forecasting enhanced with model distillation,

    Z. Lai, D. Zhang, H. Li, C. S. Jensen, H. Lu, and Y . Zhao, “LightCTS*: Lightweight correlated time series forecasting enhanced with model distillation,”IEEE Trans. Knowl. Data Eng., vol. 36, no. 12, pp. 8695– 8710, 2024

  36. [36]

    Generalized additive models for location, scale and shape,

    R. A. Rigby and D. M. Stasinopoulos, “Generalized additive models for location, scale and shape,”Journal of the Royal Statistical Society Series C: Applied Statistics, vol. 54, no. 3, pp. 507–554, 2005

  37. [37]

    Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction,

    O. Makansi, E. Ilg, O. Cicek, and T. Brox, “Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction,” inProc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 7144–7153

  38. [38]

    Variational Heteroscedastic Gaussian Process Regression,

    M. L ´azaro-Gredilla and M. K. Titsias, “Variational Heteroscedastic Gaussian Process Regression,” inProc. Int. Conf. Mach. Learn., 2011, pp. 841–848

  39. [39]

    Generalized autoregressive conditional heteroskedastic- ity,

    T. Bollerslev, “Generalized autoregressive conditional heteroskedastic- ity,”Journal of Econometrics, vol. 31, no. 3, pp. 307–327, 1986

  40. [40]

    Multi-Resolution Expansion of Analysis in Time-Frequency Domain for Time Series Forecasting,

    K. Yan, C. Long, H. Wu, and Z. Wen, “Multi-Resolution Expansion of Analysis in Time-Frequency Domain for Time Series Forecasting,” IEEE Trans. Knowl. Data Eng., vol. 36, no. 11, pp. 6667–6680, 2024

  41. [41]

    PromptCast: A new prompt-based learning paradigm for time series forecasting,

    H. Xue and F. D. Salim, “PromptCast: A new prompt-based learning paradigm for time series forecasting,”IEEE Trans. Knowl. Data Eng., vol. 36, no. 11, pp. 6851–6864, 2024

  42. [42]

    Diffusion-TS: Interpretable Diffusion for General Time Series Generation,

    X. Yuan and Y . Qiao, “Diffusion-TS: Interpretable Diffusion for General Time Series Generation,” inProc. Int. Conf. Learn. Representations, 2024, pp. 1–29

  43. [43]

    The Capacity and Robustness Trade- Off: Revisiting the channel independent strategy for multivariate time series forecasting,

    L. Han, H.-J. Ye, and D.-C. Zhan, “The Capacity and Robustness Trade- Off: Revisiting the channel independent strategy for multivariate time series forecasting,”IEEE Trans. Knowl. Data Eng., vol. 36, no. 11, pp. 7129–7142, 2024

  44. [44]

    Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series ,

    S. Narayan Shukla and B. Marlin, “Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series ,” inProc. Int. Conf. Learn. Representations, 2022, pp. 1–20

  45. [45]

    A Time Series is Worth 64 Words: Long-term Forecasting with Transformers,

    Y . Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A Time Series is Worth 64 Words: Long-term Forecasting with Transformers,” inProc. Int. Conf. Learn. Representations, 2023, pp. 1–24

  46. [46]

    Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift,

    T. Kim, J. Kim, Y . Tae, C. Park, J. Choi, and C. Jaegul, “Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift,” inProc. Int. Conf. Learn. Representations, 2022, pp. 1–25

  47. [47]

    Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks,

    S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks,” inAdvances in Neural Inf. Process. Syst., 2015, pp. 1171–1179

  48. [48]

    Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning,

    Y . Gal and Z. Ghahramani, “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning,” inProc. Int. Conf. Mach. Learn., 2016, pp. 1050–1059

  49. [49]

    FITS: Modeling Time Series with 10k Parameters,

    Y . Xu, A. Zeng, and Q. Xu, “FITS: Modeling Time Series with 10k Parameters,” inProc. Int. Conf. Learn. Representations, 2024, pp. 1–24

  50. [50]

    iTransformer: Inverted Transformers Are Effective for Time Series Forecasting,

    Y . Liu, T. Hu, H. Zhang, H. Wu, S. Wang, L. Ma, and M. Long, “iTransformer: Inverted Transformers Are Effective for Time Series Forecasting,” inProc. Int. Conf. Learn. Representations, 2024, pp. 1–25

  51. [51]

    Koopa: Learning non-stationary time series dynamics with koopman predictors,

    Y . Liu, C. Li, J. Wang, and M. Long, “Koopa: Learning non-stationary time series dynamics with koopman predictors,” inAdvances in Neural Inf. Process. Syst., 2023, pp. 12 271–12 290

  52. [52]

    Predict, Refine, Synthesize: Self-guiding diffusion models for probabilistic time series forecasting,

    M. Kollovieh, A. F. Ansari, M. Bohlke-Schneider, J. Zschiegner, H. Wang, and Y . B. Wang, “Predict, Refine, Synthesize: Self-guiding diffusion models for probabilistic time series forecasting,” inAdvances in Neural Inf. Process. Syst., 2023, pp. 28 341–28 364

  53. [53]

    Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows,

    K. Rasul, A. Sheikh, I. Schuster, U. M. Bergmann, and R. V ollgraf, “Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows,” inProc. Int. Conf. Learn. Representations, 2021, pp. 1–19

  54. [54]

    PyTorch: An Imperative Style, High-Performance Deep Learning Library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” inAdvances in Neural Inf. Process. Syst., 2019...