Pith. sign in

REVIEW 3 major objections 5 minor 86 references

Frequency-Domain Multi-Modality Transportation Modeling

T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read Multi-modality traffic forecasts improve when each mode is filtered in the frequency domain and only high-consensus frequencies are shared across modes.

desk verdict Solid plug-and-play frequency module that moves multi-modality traffic numbers; the selective-synergy story is only partly isolated from plain residual capacity. read the letter →

arxiv 2607.08475 v1 pith:TVQH4W5X submitted 2026-07-09 cs.LG

classification cs.LG
keywords frequency-domainmulti-modalitytransportationmodelingspectralfilteringcross-modalitysynergytrafficforecastingplug-and-playmoduletimeseries
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Urban systems mix several transportation modes—bike inflows, taxi outflows, and the like—that share broad daily rhythms but diverge in fine-grained fluctuations. Most forecasting models either treat each mode alone or fuse them with a single time-domain rule, which can mix noise and cause negative transfer. This paper claims that the right place to coordinate those modes is the frequency domain: first refine each modality’s spectrum on its own, then let modalities contribute to a shared consensus only where their spectral strength is relatively high. The resulting module, FreMo, is lightweight, plug-and-play, and consistently lowers forecast error on three city-scale bike/taxi datasets while also lifting several general time-series backbones. A sympathetic reader cares because the same spectral idea can turn coarse multi-source fusion into selective, frequency-aware synergy without redesigning the underlying predictor.

What carries the argument

FreMo: a plug-and-play frequency-domain block whose Modality-Wise Frequency Filter (MFF) produces phase-preserving soft gates from amplitude spectra plus node embeddings, and whose Frequency-Guided Synergy Integrator (FSI) forms a Softmax-weighted consensus across modalities at each frequency bin and residual-injects it under a single learnable scalar γ.

What would settle it

On a held-out city or period, replace the amplitude-based Softmax weights in FSI with random or uniform weights (or with phase-aware reliability scores) and check whether the reported MAE/RMSE gains over the no-synergy and shared-weight ablations disappear or reverse.

Watch

Extended reading notes

Core claim

The paper establishes that multi-modality transportation forecasting benefits from an explicit two-stage frequency-domain procedure: a Modality-Wise Frequency Filter that learns per-modality, per-node soft gates on spectral amplitudes, followed by a Frequency-Guided Synergy Integrator that builds a frequency-wise consensus via Softmax over those amplitudes and injects it back through a shared residual scalar. This selective synergy outperforms both strong uni-modality models and prior multi-modality fusion methods on NYC, DC, and Chicago bike/taxi flows, and the same module improves AGCRN, TimesNet, iTransformer, and STAEformer when inserted without architectural change.

Load-bearing premise

That the amplitude of each frequency bin, after simple pooling, is a good enough signal both for deciding which frequencies to keep inside a modality and for deciding which modality should dominate the shared consensus at that frequency.

Editorial extensions

If this is right

  • Frequency-wise selective fusion can reduce negative transfer among transportation modes that share low-frequency trends but differ at high frequencies.
  • Existing time-series backbones can be upgraded for multi-modality traffic by inserting FreMo without redesigning their spatial or temporal layers.
  • Modality-specific frequency gates learned from amplitude plus node embeddings capture structured node-level spectral differences that a single shared filter would miss.
  • A single global residual scalar is sufficient to control synergy injection; more flexible per-modality or dynamic gates do not help and can hurt.
  • The same spectral refinement-plus-consensus pattern is claimed to generalize across cities of different scale and different input/output horizons.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same amplitude-driven consensus idea could be tested on other multi-source urban series (crime, air quality, energy) where low-frequency coherence is high but high-frequency noise is modality-specific.
  • If spectral amplitude is only a weak reliability proxy, replacing it with a learned or phase-sensitive score might further cut large-error events that currently show up mainly as RMSE reductions.
  • Because FreMo is residual and phase-preserving, it could be stacked or applied at multiple encoder depths to see whether early versus late frequency synergy matters more.
  • Node clustering by learned frequency gates suggests a diagnostic: maps of gate profiles might reveal which districts are dominated by periodic commuting versus bursty last-mile traffic.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes FreMo, a plug-and-play frequency-domain module for multi-modality transportation forecasting (Bike/Taxi In/Outflow). After a TCN + latent-attention spatio-temporal encoder, MFF applies modality- and node-specific soft gates (learned from pooled spectral amplitudes plus node embeddings) to rFFT spectra in a phase-preserving way; FSI then treats filtered amplitudes as reliability scores, Softmaxes them across modalities per frequency bin to form a consensus spectrum, and residual-injects it with a shared scalar γ before irFFT and a gated predictor. On NYC, DC, and Chicago, FreMo reports consistent MAE/RMSE gains over ten uni- and multi-modality baselines, with ablations, γ variants, plug-and-play lifts on four backbones, hyperparameter sweeps, efficiency numbers, and qualitative visualizations of learned weights.

Significance. If the gains are real and the frequency-guided synergy story holds, FreMo is a useful, lightweight, architecture-agnostic enhancer for multi-modality urban forecasting: it is simple, code is released, overhead is small (+0.028–0.054M params), and it improves both specialized multi-modality models and general time-series backbones (AGCRN, TimesNet, iTransformer, STAEformer). The empirical package (three cities, four modalities, ablations, plug-and-play, case studies) is stronger than many KDD-style forecasting papers. The main conceptual contribution is the explicit split between modality-wise spectral refinement and frequency-wise selective consensus, which is a clean framing even if the reliability proxy is heuristic.

major comments (3)
  1. The central claim that FSI provides selective, non-negative-transfer synergy rests on treating pooled filtered amplitude as a reliability score (Eqs. 9–11) and Softmaxing it across modalities. Fig. 1(c) is only motivational; α is never validated against spectral coherence, phase consistency, or an oracle of which modality is more predictive at each bin. Table 3 removes FSI or shares generators, but never replaces the amplitude proxy (e.g., with phase-aware coherence, a learned reliability head, random/uniform α, or fixed low-frequency-only consensus). Without that control, the edge over MoSSL and the plug-and-play lifts could be explained by MFF gating plus residual capacity rather than by the claimed frequency-guided synergy. A targeted ablation or correlation of α with true cross-modality coherence is needed to support the load-bearing narrative.
  2. Forecasting setup is short-horizon only (input 16 → output 1/2/3; Table 1). Multi-modality transportation claims often care about longer horizons where low-frequency consensus should matter more. The paper does not show whether FreMo’s advantage grows, shrinks, or reverses as O increases, nor whether FSI’s low-frequency concentration (Fig. 6a) remains beneficial. At least one longer-horizon experiment (or explicit limitation) is needed for the generalization claim in the abstract and conclusion.
  3. Baselines and multi-modality protocol: uni-modality models (GWN, AGCRN, MTGNN, TimesNet, iTransformer, STAEformer) appear to be run per modality or as independent channels; multi-modality baselines are fewer and partly from overlapping prior work (MoSSL, etc.). It is unclear whether all methods receive identical multi-modality tensors, the same train/val/test splits, and the same early-stopping. Table 2 averages three runs but reports no std/CI. For a SOTA claim, clarify the multi-modality adaptation of uni-modality baselines and add variance or a significance test on the main table.
minor comments (5)
  1. Notation: F is used both for rFFT and for the number of frequency bins; H^F vs H_F and m,n,: indexing are dense. A short notation table would help.
  2. Fig. 6 caption and body mix Chinese notes with English; clean for camera-ready.
  3. Algorithm 1 uses g_m while the text uses G_m; align naming.
  4. Appendix C chooses different (d,L) per city for the main table; state this clearly in §5.1 so Table 2 is not read as a single fixed hyperparameter setting.
  5. Related work is thorough but could more sharply contrast FreMo with frequency-domain MLPs / FilterNet / FEDformer on the multi-modality (not uni-modality) axis.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: end-to-end empirical architecture whose gates and synergy weights are free parameters optimized on held-out MAE, not algebraic identities or self-defined predictions.

full rationale

FreMo is a standard plug-and-play neural module (MFF frequency gates via learned generators on amplitude+node emb, FSI Softmax over pooled spectral amplitudes, residual scalar γ) trained end-to-end by MAE on multi-modality tensors and evaluated on held-out horizons of NYC/DC/Chicago Bike/Taxi flows. Performance claims (Table 2, ablations Table 3, plug-and-play Table 4) are empirical comparisons against external and self-baselines; nothing reduces by construction to a fitted constant or definitional identity. Self-citations (MoSSL, TTS-Norm, Gaussian-mixture papers by overlapping authors) appear only as related-work baselines and datasets, not as load-bearing uniqueness theorems or ansatzes that force the reported gains. Spectral-coherence motivation (Fig. 1) and amplitude-as-proxy (Eqs. 5–11) are modeling assumptions whose validity is open to correctness critique, but they do not create circular derivation: the weights remain free parameters, not tautological renamings of the evaluation metric. Score 1 reflects only the ordinary presence of non-load-bearing self-baselines.

Assumptions & free parameters 5 free parameters · 4 assumptions · 2 invented entities

The central empirical claim rests on standard FFT identities, the modeling choice that amplitude is a reliability proxy, a small set of architectural free parameters (hidden/latent dims, learnable gamma, per-modality gate generators), and the usual supervised MAE objective. No new physical entities are postulated; the invented pieces are architectural modules whose only evidence is the reported forecast gains.

free parameters (5)
  • hidden dimension d = 64 (or 32 for Chicago)
    Chosen by sensitivity sweep; main results use d=64 (NYC/DC) or 32 (Chicago).
  • latent dimension L (spatial bottleneck and node emb) = 64 or 32
    Hand-tuned; main results use L=64 (NYC) or 32 (DC/Chicago).
  • synergy residual scalar gamma = learned (init 0)
    Learnable scalar initialized at 0, shared across modalities; controls how much consensus is injected.
  • modality-wise frequency gate generators G_m
    Per-modality networks mapping pooled amplitude + node embedding to [0,1]^F gates; parameters fit end-to-end.
  • encoder depth / kernel / batch / optimizer settings = 4 layers, kernel 3, batch 16
    Four encoder layers, TCN kernel 3, batch 16, Adam; fixed by authors for all main tables.
assumptions (4)
  • standard math Real FFT / inverse rFFT preserve the information needed for subsequent linear filtering and residual injection without destroying temporal alignment when gates are real-valued.
    Invoked in Eqs. 4, 8, 13 and Algorithm 1; standard Fourier property.
  • domain assumption Spectral amplitude after channel pooling is a sufficient proxy for both intra-modality informativeness and inter-modality reliability at each frequency bin.
    Core of MFF weight generation (Eqs. 5–7) and FSI Softmax (Eqs. 9–10); not independently validated beyond ablations.
  • ad hoc to paper A single shared residual scale gamma is preferable to per-modality or dynamic gates for injecting consensus.
    Justified only by the gamma ablation in Figure 3; design choice specific to FreMo.
  • domain assumption Multi-modality traffic tensors on the three chosen cities are representative enough that gains generalize to ‘diverse forecasting scenarios’.
    Stated in abstract and conclusion; only NYC/DC/Chicago bike+taxi data are shown.
invented entities (2)
  • Modality-Wise Frequency Filter (MFF)
    purpose: Learn per-modality, per-node soft frequency gates from amplitude + node embedding and apply phase-preserving filtering.
    Architectural module introduced in §4.2; evidence is ablation and weight visualizations, not external measurement.
  • Frequency-Guided Synergy Integrator (FSI)
    purpose: Form frequency-wise Softmax weights over modalities, build a consensus spectrum, and residual-inject it via gamma.
    Architectural module introduced in §4.3; evidence is ablation deltas and case-study weight maps.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Frequency-Domain Multi-Modality Transportation Modeling." pith.science (2026). https://pith.science/paper/TVQH4W5X

@misc{pith2026260708475,
  author       = {Pith},
  title        = {Pith review of: Frequency-Domain Multi-Modality Transportation Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TVQH4W5X}},
  note         = {Machine review of arXiv:2607.08475}
}
read the original abstract

Multi-modality transportation refers to urban systems composed of multiple transportation modes, such as traffic flow and public transit, whose dynamics are coupled by shared temporal patterns. Accurate multi-modality transportation forecasting remains challenging because (1) different modalities exhibit distinct spectral characteristics and (2) interact unevenly across frequencies, whereas most existing methods operate primarily in the time domain or rely on coarse feature fusion. To address these limitations, we propose a lightweight yet effective Frequency-Domain Multi-Modality modeling (FreMo) that explicitly exploits the frequency domain to enable adaptive and selective cross-modality synergy. FreMo disentangles modality-wise spectral refinement from cross-modality synergy and supports plug-and-play integration with general time series backbones. Specifically, FreMo introduces a Modality-Wise Frequency Filter (MFF) to adaptively refine spectral components within each modality, emphasizing informative frequencies while suppressing noise. FreMo further incorporates a Frequency-Guided Synergy Integrator (FSI) that selectively aggregates information across modalities based on their relative contribution at each frequency, facilitating effective cross-modality knowledge sharing while mitigating negative transfer. Extensive experiments on real-world datasets show that FreMo consistently outperforms state-of-the-art baselines, with superior performance and generalization across diverse forecasting scenarios. The code is available at https://github.com/beginner-sketch/FreMo.

Figures

Figures reproduced from arXiv: 2607.08475 by the authors.

Figure 1
Figure 1. Multi-modality transportation data with represen [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An overview of the proposed FreMo. Left: As a plug-and-play component, FreMo can be integrated into time series [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Performance of different synergy injection strate [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Hyperparameter sensitivity on the NYC. Frequency (Low → High) 0 2 4 6 Spectral Energy Bike Inflow Frequency (Low → High) 0 2 4 6 8 Spectral Energy Bike Outflow 0 1 2 3 4 5 6 7 8 Frequency (Low → High) 0 2 4 Spectral Energy Taxi Inflow 0 1 2 3 4 5 6 7 8 Frequency (Low →…
Figure 5
Figure 5. Figure 5: The learned modality-wise frequency weights on [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Node-level frequency gating and synergy alloca [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Hyperparameter sensitivity studies of hidden dimension [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

86 extracted references · 86 canonical work pages

  1. [1]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normaliza- tion.arXiv preprint arXiv:1607.06450(2016)

  2. [2]

    Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting.Advances in Neural Information Processing Systems33 (2020), 17804–17815

  3. [3]

    Wanlin Cai, Yuxuan Liang, Xianggen Liu, Jianshuai Feng, and Yuankai Wu. 2024. Msgnet: Learning multi-scale inter-series correlations for multivariate time series forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 11141–11149

  4. [4]

    Peng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu, Yihang Wang, Qingsong Wen, Bin Yang, and Chenjuan Guo. 2024. Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series Forecasting. InInternational Conference on Learning Representations

  5. [5]

    Mingyue Cheng, Qi Liu, Zhiding Liu, Zhi Li, Yucong Luo, and Enhong Chen

  6. [6]

    InProceedings of the ACM web conference

    Formertime: Hierarchical multi-scale representations for multivariate time series classification. InProceedings of the ACM web conference. 1437–1445

  7. [7]

    Tao Dai, Beiliang Wu, Peiyuan Liu, Naiqi Li, Jigang Bao, Yong Jiang, and Shu-Tao Xia. 2024. Periodicity decoupling framework for long-term series forecasting. In International Conference on Learning Representations

  8. [8]

    Shohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V Smith, and Flora D Salim. 2022. Cocoa: Cross modality contrastive learning for sensor data.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies6, 3 (2022), 1–28

Show all 86 references
  1. [9]

    Jiewen Deng, Jinliang Deng, Renhe Jiang, and Xuan Song. 2023. Learning Gauss- ian Mixture Representations for Tensor Time Series Forecasting. InProceedings of the International Joint Conference on Artificial Intelligence. 2077–2085

  2. [10]

    Jiewen Deng, Jinliang Deng, Du Yin, Renhe Jiang, and Xuan Song. 2023. Tts-norm: Forecasting tensor time series via multi-way normalization.ACM Transactions on Knowledge Discovery from Data18, 1 (2023), 1–25

  3. [11]

    Jiewen Deng, Renhe Jiang, Jiaqi Zhang, and Xuan Song. 2024. Multi-Modality Spatio-Temporal Forecasting via Self-Supervised Learning. InProceedings of the International Joint Conference on Artificial Intelligence. 2018–2026

  4. [12]

    Leyan Deng, Defu Lian, Zhenya Huang, and Enhong Chen. 2022. Graph con- volutional adversarial networks for spatiotemporal anomaly detection.IEEE Transactions on Neural Networks and Learning Systems33, 6 (2022), 2416–2428

  5. [13]

    Wei Fan, Pengyang Wang, Dongkun Wang, Dongjie Wang, Yuanchun Zhou, and Yanjie Fu. 2023. Dish-ts: a general paradigm for alleviating distribution shift in time series forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 7522–7529

  6. [14]

    Shen Fang, Qi Zhang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan. 2019. GSTNet: Global Spatial-Temporal Network for Traffic Flow Prediction.. InIJCAI. 2286–2293

  7. [15]

    Yuchen Fang, Yuxuan Liang, Bo Hui, Zezhi Shao, Liwei Deng, Xu Liu, Xinke Jiang, and Kai Zheng. 2025. Efficient large-scale traffic forecasting with transformers: A spatial data management perspective. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data M...

  8. [16]

    Ziquan Fang, Lu Pan, Lu Chen, Yuntao Du, and Yunjun Gao. 2021. MDTP: a multi-source deep traffic prediction framework over spatio-temporal trajectory data.Proceedings of the VLDB Endowment14, 8 (2021), 1289–1297

  9. [17]

    En Fu and Yanyan Hu. 2025. Frequency-Masked Embedding Inference: A Non- Contrastive Approach for Time Series Representation Learning. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 16639–16647

  10. [18]

    Kan Guo, Yongli Hu, Yanfeng Sun, Sean Qian, Junbin Gao, and Baocai Yin. 2021. Hierarchical Graph Convolution Network for Traffic Forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 151–159

  11. [19]

    Shengnan Guo, Youfang Lin, Ning Feng, Chao Song, and Huaiyu Wan. 2019. Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 922–929

  12. [20]

    Shengnan Guo, Youfang Lin, Shijie Li, Zhaoming Chen, and Huaiyu Wan. 2019. Deep spatial–temporal 3D convolutional neural networks for traffic data fore- casting.IEEE Transactions on Intelligent Transportation Systems20, 10 (2019), 3913–3926

  13. [21]

    Jindong Han, Hao Liu, Hengshu Zhu, Hui Xiong, and Dejing Dou. 2021. Joint air quality and weather prediction based on multi-adversarial spatiotemporal networks. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 4081–4089

  14. [22]

    Liangzhe Han, Bowen Du, Leilei Sun, Yanjie Fu, Yisheng Lv, and Hui Xiong

  15. [23]

    InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    Dynamic and multi-faceted spatio-temporal deep learning for traffic speed forecasting. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 547–555

  16. [24]

    Chao Huang, Chuxu Zhang, Jiashu Zhao, Xian Wu, Dawei Yin, and Nitesh Chawla

  17. [25]

    InThe World Wide Web conference

    Mist: A multiview and multimodal spatial-temporal learning framework for citywide abnormal event forecasting. InThe World Wide Web conference. 717–728. KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Jiewen Deng, Hangchen Liu, Junchen Li, Boyuan Zhang, and Renhe Jiang

  18. [26]

    Chao Huang, Junbo Zhang, Yu Zheng, and Nitesh V. Chawla. 2018. DeepCrime: Attentive Hierarchical Recurrent Networks for Crime Prediction. InProceedings of the ACM International Conference on Information and Knowledge Management. 1423–1432

  19. [27]

    Qihe Huang, Lei Shen, Ruixin Zhang, Jiahuan Cheng, Shouhong Ding, Zhengyang Zhou, and Yang Wang. 2024. Hdmixer: Hierarchical dependency with extend- able patch for multivariate time series forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 1...

  20. [28]

    Qihe Huang, Lei Shen, Ruixin Zhang, Shouhong Ding, Binwu Wang, Zhengyang Zhou, and Yang Wang. 2023. Crossgnn: Confronting noisy multivariate time series via cross interaction refinement.Advances in Neural Information Processing Systems36 (2023), 46885–46902

  21. [29]

    Jiahao Ji, Jingyuan Wang, Chao Huang, Junjie Wu, Boren Xu, Zhenhe Wu, Junbo Zhang, and Yu Zheng. 2023. Spatio-temporal self-supervised learning for traffic flow prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4356–4364

  22. [30]

    Jiawei Jiang, Chengkai Han, Wayne Xin Zhao, and Jingyuan Wang. 2023. Pdformer: Propagation delay-aware dynamic long-range transformer for traffic flow prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4365–4373

  23. [31]

    Renhe Jiang, Zekun Cai, Zhaonan Wang, Chuang Yang, Zipei Fan, Quanjun Chen, Kota Tsubouchi, Xuan Song, and Ryosuke Shibasaki. 2021. DeepCrowd: A deep model for large-scale citywide crowd density and flow prediction.IEEE Transactions on Knowledge and Data Engineering35, 1 (2021...

  24. [32]

    Renhe Jiang, Du Yin, Zhaonan Wang, Yizhuo Wang, Jiewen Deng, Hangchen Liu, Zekun Cai, Jinliang Deng, Xuan Song, and Ryosuke Shibasaki. 2021. Dl-traff: Survey and benchmark of deep learning models for urban traffic prediction. In Proceedings of the ACM International Conference ...

  25. [33]

    Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time series forecasting by reprogramming large language models. In International Conference on Learning Representations

  26. [34]

    Hyunwook Lee and Sungahn Ko. 2024. TESTAM: A Time-Enhanced Spatio- Temporal Attention Model with Mixture of Experts. InInternational Conference on Learning Representations

  27. [35]

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks. InInternational Conference on Machine Learning. 3744–3753

  28. [36]

    Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. InInternational Conference on Learning Representations

  29. [37]

    Shengsheng Lin, Weiwei Lin, Xinyi Hu, Wentai Wu, Ruichao Mo, and Haocheng Zhong. 2024. Cyclenet: Enhancing time series forecasting through modeling periodic patterns.Advances in Neural Information Processing Systems37 (2024), 106315–106345

  30. [38]

    Hangchen Liu, Zheng Dong, Renhe Jiang, Jiewen Deng, Jinliang Deng, Quanjun Chen, and Xuan Song. 2023. Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. InProceedings of the ACM International Conference on Information and Knowledge manag...

  31. [39]

    Hao Liu, Qiyu Wu, Fuzhen Zhuang, Xinjiang Lu, Dejing Dou, and Hui Xiong. 2021. Community-aware multi-task transportation demand prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 320–327

  32. [40]

    Peiyuan Liu, Hang Guo, Tao Dai, Naiqi Li, Jigang Bao, Xudong Ren, Yong Jiang, and Shu-Tao Xia. 2025. Calf: Aligning llms for time series forecasting via cross- modal fine-tuning. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 18915–18923

  33. [41]

    Qinghua Liu and John Paparrizos. 2024. The elephant in the room: Towards a re- liable time-series anomaly detection benchmark.Advances in Neural Information Processing Systems37 (2024), 108231–108261

  34. [42]

    Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long. 2024. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. InInternational Conference on Learning Representations

  35. [43]

    Wang Lu, Jindong Wang, Xinwei Sun, Yiqiang Chen, and Xing Xie. 2023. Out-of- distribution Representation Learning for Time Series Classification. InInterna- tional Conference on Learning Representations

  36. [44]

    Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam

    Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In International Conference on Learning Representations

  37. [45]

    Boris N Oreshkin, Arezou Amini, Lucy Coyle, and Mark Coates. 2021. FC-GAGA: Fully connected gated graph architecture for spatio-temporal traffic forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 9233–9241

  38. [46]

    Zheyi Pan, Yuxuan Liang, Weifeng Wang, Yong Yu, Yu Zheng, and Junbo Zhang

  39. [47]

    InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    Urban traffic prediction from spatio-temporal data using deep meta learning. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1720–1730

  40. [48]

    Xihao Piao, Zheng Chen, Taichi Murayama, Yasuko Matsubara, and Yasushi Saku- rai. 2024. Fredformer: Frequency debiased transformer for time series forecasting. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2400–2410

  41. [49]

    Xiangfei Qiu, Xingjian Wu, Yan Lin, Chenjuan Guo, Jilin Hu, and Bin Yang

  42. [50]

    In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining

    Duet: Dual clustering enhanced multivariate time series forecasting. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1185–1196

  43. [51]

    Zezhi Shao, Zhao Zhang, Fei Wang, and Yongjun Xu. 2022. Pre-training enhanced spatial-temporal graph neural network for multivariate time series forecasting. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1567–1577

  44. [52]

    Senior, and Koray Kavukcuoglu

    Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016. WaveNet: A Generative Model for Raw Audio. InProceedings of the ISCA Workshop on Speech Synthesis Workshop

  45. [53]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in Neural Information Processing Systems30 (2017)

  46. [54]

    Chengsen Wang, Zirui Zhuang, Qi Qi, Jingyu Wang, Xingyu Wang, Haifeng Sun, and Jianxin Liao. 2023. Drift doesn’t matter: Dynamic decomposition with diffusion reconstruction for unstable multivariate time series anomaly detection. Advances in Neural Information Processing Syste...

  47. [55]

    Shiyu Wang, Jiawei Li, Xiaoming Shi, Zhou Ye, Baichuan Mo, Wenze Lin, Sheng- tong Ju, Zhixuan Chu, and Ming Jin. 2025. TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis. InInternational Conference on Learning Representations

  48. [56]

    Shiyu Wang, Haixu Wu, Xiaoming Shi, Tengge Hu, Huakun Luo, Lintao Ma, James Y Zhang, and JUN ZHOU. 2024. TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting. InInternational Conference on Learning Representations

  49. [57]

    Zhaonan Wang, Renhe Jiang, Hao Xue, Flora D Salim, Xuan Song, and Ryosuke Shibasaki. 2022. Event-aware multimodal mobility nowcasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 4228–4236

  50. [58]

    Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. InInternational Conference on Learning Representations

  51. [59]

    Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: De- composition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems34 (2021), 22419–22430

  52. [60]

    Xian Wu, Chao Huang, Chuxu Zhang, and Nitesh V Chawla. 2020. Hierarchically structured transformer networks for fine-grained spatial event forecasting. In Proceedings of the Web Conference. 2320–2330

  53. [61]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 753–763

  54. [62]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph wavenet for deep spatial-temporal graph modeling. InProceedings of the International Joint Conference on Artificial Intelligence. 1907–1913

  55. [63]

    Lianghao Xia, Chao Huang, Yong Xu, Peng Dai, Liefeng Bo, Xiyue Zhang, and Tianyi Chen. 2021. Spatial-Temporal Sequential Hypergraph Network for Crime Prediction with Dynamic Multiplex Relation Learning. InProceedings of the International Joint Conference on Artificial Intellig...

  56. [64]

    Song Yang, Jiamou Liu, and Kaiqi Zhao. 2021. Space Meets Time: Local Spacetime Neural Network For Traffic Flow Forecasting. InIEEE International Conference on Data Mining. 817–826

  57. [65]

    Huaxiu Yao, Yiding Liu, Ying Wei, Xianfeng Tang, and Zhenhui Li. 2019. Learning from multiple cities: A meta-learning approach for spatial-temporal prediction. InThe World Wide Web conference. 2181–2191

  58. [66]

    Junchen Ye, Leilei Sun, Bowen Du, Yanjie Fu, Xinran Tong, and Hui Xiong. 2019. Co-prediction of multiple transportation demands based on deep spatio-temporal neural network. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 305–313

  59. [67]

    Weiwei Ye, Songgaojun Deng, Qiaosha Zou, and Ning Gui. 2024. Frequency adaptive normalization for non-stationary time series forecasting.Advances in Neural Information Processing Systems37 (2024), 31350–31379

  60. [68]

    Kun Yi, Jingru Fei, Qi Zhang, Hui He, Shufeng Hao, Defu Lian, and Wei Fan. 2024. Filternet: Harnessing frequency filters for time series forecasting.Advances in Neural Information Processing Systems37 (2024), 55115–55140

  61. [69]

    Kun Yi, Qi Zhang, Wei Fan, Hui He, Liang Hu, Pengyang Wang, Ning An, Long- bing Cao, and Zhendong Niu. 2023. FourierGNN: Rethinking multivariate time series forecasting from a pure graph perspective.Advances in Neural Information Processing Systems36 (2023), 69638–69660

  62. [70]

    Kun Yi, Qi Zhang, Wei Fan, Shoujin Wang, Pengyang Wang, Hui He, Ning An, Defu Lian, Longbing Cao, and Zhendong Niu. 2023. Frequency-domain MLPs are more effective learners in time series forecasting.Advances in Neural Information Processing Systems36 (2023), 76656–76679. Frequ...

  63. [71]

    Chengqing Yu, Fei Wang, Zezhi Shao, Tangwen Qian, Zhao Zhang, Wei Wei, and Yongjun Xu. 2024. Ginar: An end-to-end multivariate time series forecasting model suitable for variable missing. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3989–4000

  64. [72]

    Haitao Yuan, Guoliang Li, Zhifeng Bao, and Ling Feng. 2021. An effective joint prediction model for travel demands and traffic flows. InIEEE International Conference on Data Engineering. 348–359

  65. [73]

    Wenzhen Yue, Yong Liu, Xianghua Ying, Bowei Xing, Ruohao Guo, and Ji Shi

  66. [74]

    InProceedings of the International Joint Conference on Artificial Intelligence

    FreEformer: frequency enhanced transformer for multivariate time se- ries forecasting. InProceedings of the International Joint Conference on Artificial Intelligence. 3606–3614

  67. [75]

    Dongran Zhang, Jiangnan Yan, Kemal Polat, Adi Alhudhaif, and Jun Li. 2024. Multimodal joint prediction of traffic spatial-temporal data with graph sparse at- tention mechanism and bidirectional temporal convolutional network.Advanced Engineering Informatics62 (2024), 102533

  68. [76]

    Junbo Zhang, Yu Zheng, and Dekang Qi. 2017. Deep spatio-temporal residual net- works for citywide crowd flows prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 31. 1655–1661

  69. [77]

    Qianru Zhang, Chao Huang, Lianghao Xia, Zheng Wang, Zhonghang Li, and Siuming Yiu. 2023. Automated Spatio-Temporal Graph Contrastive Learning. In Proceedings of the ACM Web Conference. 295–305

  70. [78]

    Xiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, and Marinka Zitnik. 2022. Self-supervised contrastive pre-training for time series via time-frequency consis- tency.Advances in Neural Information Processing Systems35 (2022), 3988–4003

  71. [79]

    Yunhao Zhang and Junchi Yan. 2023. Crossformer: Transformer Utilizing Cross- Dimension Dependency for Multivariate Time Series Forecasting. InInternational Conference on Learning Representations

  72. [80]

    Lifan Zhao and Yanyan Shen. 2024. Rethinking Channel Dependence for Multi- variate Time Series Forecasting: Learning from Leading Indicators. InInterna- tional Conference on Learning Representations

  73. [81]

    Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020. Gman: A graph multi-attention network for traffic prediction. InProceedings of the AAAI conference on artificial intelligence, Vol. 34. 1234–1241

  74. [82]

    Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond efficient transformer for long se- quence time-series forecasting. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 11106–11115

  75. [83]

    Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. InInternational Conference on Machine Learning. 27268–27286

  76. [84]

    Tian Zhou, Peisong Niu, Liang Sun, Rong Jin, et al . 2023. One fits all: Power general time series analysis by pretrained lm.Advances in neural information processing systems36 (2023), 43322–43355

  77. [85]

    Xian Zhou, Yanyan Shen, Yanmin Zhu, and Linpeng Huang. 2018. Predicting Multi-step Citywide Passenger Demands Using Attention-based Neural Networks. InProceedings of the ACM International Conference on Web Search and Data Mining. 736–744

  78. [86]

    Zhibo Zhu, Ziqi Liu, Ge Jin, Zhiqiang Zhang, Lei Chen, Jun Zhou, and Jianyong Zhou. 2021. MixSeq: Connecting Macroscopic Time Series Forecasting with Microscopic Time Series Data.Advances in Neural Information Processing Systems 34 (2021), 12904–12916. A Complexity and Efficie...

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.