REVIEW 1 major objections 2 minor 29 references
MoE Enhanced Federated Learning for Spatiotemporal Prediction
T0 review · 1 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read A mixture-of-experts federated model improves traffic predictions for cities with scarce data by dynamically fusing source-city experts.
desk verdict MoE-FedTP layers a standard MoE personalization trick onto federated spatiotemporal models and reports gains on four traffic datasets, but the support for the gains is thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Mixture-of-experts architecture in which each expert is a partially shared network from one source city and the gating network learns to route or combine their outputs for a given target city.
What would settle it
MoE-FedTP would be falsified if it stopped outperforming the baselines on a fresh collection of cities whose traffic patterns differ in structure from the four original datasets.
Extended reading notes
Core claim
MoE-FedTP first runs spatiotemporal neural networks on source and target city data to extract features, then creates a set of expert networks from different source cities via partial parameter sharing and trains a gating mechanism that dynamically weights and fuses the experts to model city-specific traffic dynamics.
Load-bearing premise
Partial parameter sharing from a fixed collection of source-city experts plus a gating network will be enough to represent the range of traffic heterogeneities in new target cities.
Editorial extensions
If this is right
- Target cities obtain higher prediction accuracy than with standard federated averaging or single-source transfer.
- Only model updates are exchanged, so raw traffic observations remain local.
- The gating network produces city-specific weightings that reflect distinct spatiotemporal regimes.
- Communication cost stays lower than full-model sharing because only partial parameters move between cities.
Reading between the lines
- The same partial-sharing-plus-gating pattern could be tested on other privacy-sensitive spatiotemporal tasks such as energy load or air-quality forecasting.
- Increasing the number of source experts might extend coverage to cities whose patterns lie outside the current set.
- If the gating weights turn out to cluster cities by similarity, the framework could also support rapid expert selection for entirely new cities.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes MoE-FedTP, a personalized federated cross-city spatiotemporal prediction framework that extracts features via spatiotemporal neural networks, derives expert networks from source cities via partial parameter sharing, and uses a learned gating mechanism to dynamically fuse experts for modeling urban heterogeneity while preserving privacy. The central empirical claim is that MoE-FedTP consistently outperforms state-of-the-art cross-city and federated learning baselines on four real-world traffic datasets.
Significance. If the outperformance claim holds with rigorous validation, the work would offer a practical advance in privacy-preserving knowledge transfer for data-scarce cities in intelligent transportation systems, extending standard MoE personalization patterns to spatiotemporal federated settings.
major comments (1)
- [Abstract] Abstract: the claim of consistent outperformance on four datasets is presented without any quantitative metrics, error values, statistical tests, ablation results, or error analysis; this absence directly undermines assessment of the central empirical claim and its support for the design premise of partial sharing plus gating.
minor comments (2)
- Notation for the gating network and expert fusion is introduced without an explicit equation or diagram reference, making the adaptation mechanism harder to follow.
- The description of 'lightweight' MoE networks lacks a parameter count comparison to baselines, which would clarify the efficiency claim.
Simulated Author's Rebuttal
We thank the referee for the detailed feedback. We address the comment on the abstract below and will incorporate the suggested changes in the revised manuscript.
read point-by-point responses
-
Referee: [Abstract] Abstract: the claim of consistent outperformance on four datasets is presented without any quantitative metrics, error values, statistical tests, ablation results, or error analysis; this absence directly undermines assessment of the central empirical claim and its support for the design premise of partial sharing plus gating.
Authors: We agree that the abstract would be strengthened by the inclusion of quantitative metrics to support the outperformance claim. In the revised manuscript we will add concise performance highlights (e.g., average relative improvements in MAE and RMSE across the four datasets) while keeping the abstract within length limits. The full quantitative results, statistical significance tests, ablation studies, and error analyses are already reported in Sections 4 and 5; the abstract revision will simply provide readers with an immediate quantitative anchor for the central claim. revision: yes
Circularity Check
No circularity; empirical framework evaluated on external real-world datasets
full rationale
The paper proposes MoE-FedTP as a personalized federated learning architecture using spatiotemporal networks, source-city experts via partial parameter sharing, and a learned gating mechanism. Its central claim is empirical outperformance on four real-world traffic datasets against baselines. No derivation chain, equations, or predictions are presented that reduce by construction to fitted parameters or self-citations within the paper; the architecture is a standard MoE personalization pattern whose effectiveness is precisely what the external-dataset experiments are designed to test. The work is self-contained against external benchmarks with no load-bearing internal reductions.
Assumptions & free parameters
Cite this review
Pith. "Pith review of MoE Enhanced Federated Learning for Spatiotemporal Prediction." pith.science (2026). https://pith.science/paper/7FSPFADV
@misc{pith2026260610499,
author = {Pith},
title = {Pith review of: MoE Enhanced Federated Learning for Spatiotemporal Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/7FSPFADV}},
note = {Machine review of arXiv:2606.10499}
}
read the original abstract
Traffic prediction is fundamental to intelligent transportation systems and urban computing, yet many cities continue to suffer from traffic data scarcity due to limited sensor deployment and uneven urban development. Cross-city knowledge transfer has thus attracted increasing attention, enabling data-rich cities to assist data-scarce ones. However, centralized approaches raise privacy concerns, while existing federated methods struggle with pronounced spatiotemporal heterogeneity across cities. To address these challenges, we propose MoE-FedTP, a personalized federated cross-city spatiotemporal prediction framework based on lightweight Mixture-of-Experts (MoE) networks. MoE-FedTP first employs spatiotemporal neural networks to extract features from both source and target cities, then introduces a set of expert networks derived from different source cities through partial parameter sharing. A gating mechanism dynamically fuses the experts to capture diverse traffic dynamics, achieving fine-grained modeling of urban heterogeneity while preserving privacy. Experiments on four real-world traffic datasets show that MoE-FedTP consistently outperforms state-of-the-art cross-city and federated learning baselines, demonstrating its effectiveness in enhancing prediction accuracy for data-scarce cities.
Figures
Reference graph
Works this paper leans on
-
[1]
Lingxiao Cao, Bin Wang, Guiyuan Jiang, Yanwei Yu, and Junyu Dong. 2025. Spatiotemporal-aware Trend-Seasonality Decomposition Network for Traffic Flow Forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 11463–11471
2025
-
[2]
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS 2014 Workshop on Deep Learning, December 2014
2014
-
[3]
Xiao Han, Guojiang Shen, Xi Yang, and Xiangjie Kong. 2020. Congestion recog- nition for hybrid urban road systems via digraph convolutional network. Trans- portation Research Part C: Emerging Technologies 121 (2020), 102877
2020
-
[4]
Xiao Han, Zijian Zhang, Xiangyu Zhao, Yuanshao Zhu, Guojiang Shen, Xiangjie Kong, Xuetao Wei, Liqiang Nie, and Jieping Ye. 2025. Garlic: Gpt-augmented rein- forcement learning with intelligent control for vehicle dispatching. InProceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 255–263
2025
-
[5]
Xiao Han, Xiangyu Zhao, Liang Zhang, and Wanyu Wang. 2023. Mitigating action hysteresis in traffic signal control with traffic predictive reinforcement learning. In Proceedings of the 29th ACM SIGKDD conference on knowledge discovery and data mining. 673–684
2023
-
[6]
Xiao Han, Ding-Xuan Zhou, Guojiang Shen, Xiangjie Kong, and Yulong Zhao
-
[7]
IEEE Transactions on Vehicular Technology 73, 11 (2024), 16051–16062
Deep trajectory recovery approach of offline vehicles in the internet of vehicles. IEEE Transactions on Vehicular Technology 73, 11 (2024), 16051–16062
2024
-
[8]
Junfeng Hu, Xu Liu, Zhencheng Fan, Yifang Yin, Shili Xiang, Savitha Ramasamy, and Roger Zimmermann. 2024. Prompt-based spatio-temporal graph transfer learning. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 890–899
2024
Show all 29 references
-
[9]
Kipf and Max Welling
Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In Proc. of ICLR
2017
-
[10]
Zhenyu Lei, Yushun Dong, Jundong Li, and Chen Chen. 2025. ST-FiT: Inductive Spatial-Temporal Forecasting with Limited Training Data. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12031–12039
2025
-
[11]
Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In International Conference on Learning Representations
2018
-
[12]
Duanyang Liu, Longfeng Tang, Guojiang Shen, and Xiao Han. 2019. Traffic speed prediction: An attention-based method. Sensors 19, 18 (2019), 3836
2019
-
[13]
Xu Liu, Yutong Xia, Yuxuan Liang, Junfeng Hu, Yiwei Wang, Lei Bai, Chao Huang, Zhenguang Liu, Bryan Hooi, and Roger Zimmermann. 2023. Largest: A benchmark dataset for large-scale traffic forecasting. Advances in Neural Information Processing Systems 36 (2023), 75354–75371
2023
-
[14]
Bin Lu, Xiaoying Gan, Weinan Zhang, Huaxiu Yao, Luoyi Fu, and Xinbing Wang
-
[15]
In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
Spatio-temporal graph few-shot learning with cross-city knowledge trans- fer. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1162–1172
-
[16]
Chuizheng Meng, Sirisha Rambhatla, and Yan Liu. 2021. Cross-node federated graph neural network for spatio-temporal data modeling. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining . 1202–1211
2021
-
[17]
Veličković Petar, Cucurull Guillem, Casanova Arantxa, Romero Adriana, Lio Pietro, and B Yoshua. 2018. Graph attention networks. In Proc. of ICLR
2018
-
[18]
Weilin Ruan, Wenzhuo Wang, Siru Zhong, Wei Chen, Li Liu, and Yuxuan Liang
-
[19]
arXiv preprint arXiv:2411.09251 (2024)
Cross space and time: A spatio-temporal unitized model for traffic flow forecasting. arXiv preprint arXiv:2411.09251 (2024)
2024
-
[20]
Guojiang Shen, Xiao Han, KwaiSang Chin, and Xiangjie Kong. 2021. An attention- based digraph convolution network enabled framework for congestion recog- nition in three-dimensional road networks. IEEE Transactions on Intelligent Transportation Systems 23, 9 (2021), 14413–14426
2021
-
[21]
Yihong Tang, Ao Qu, Andy HF Chow, William HK Lam, Sze Chun Wong, and Wei Ma. 2022. Domain adversarial spatial-temporal network: A transferable framework for short-term traffic forecasting across cities. In Proceedings of the 31st ACM international conference on information & k...
2022
-
[22]
Guangyu Wang, Yujie Chen, Ming Gao, Zhiqiao Wu, Jiafu Tang, and Jiabi Zhao
-
[23]
arXiv preprint arXiv:2409.17440 (2024)
A time series is worth five experts: Heterogeneous mixture of experts for traffic flow prediction. arXiv preprint arXiv:2409.17440 (2024)
2024
-
[24]
Leye Wang, Xu Geng, Xiaojuan Ma, Feng Liu, and Qiang Yang. 2019. Cross-City Transfer Learning for Deep Spatio-Temporal Prediction. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence . Interna- tional Joint Conferences on Artificial In...
2019
-
[25]
Bing Yu, Haoteng Yin, and Zhanxing Zhu. 2018. Spatio-temporal graph convolu- tional networks: a deep learning framework for traffic forecasting. InProceedings of the 27th International Joint Conference on Artificial Intelligence . 3634–3640
2018
-
[26]
Yu Zhang, Hua Lu, Ning Liu, Yonghui Xu, Qingzhong Li, and Lizhen Cui. 2024. Personalized federated learning for cross-city traffic prediction. In 33rd Interna- tional Joint Conference on Artificial Intelligence, IJCAI . 5526–5534
2024
-
[27]
Yan Zhang, Xiaoye Miao, Bin Li, Yangyang Wu, and Yongheng Shang. 2025. Proxy- Validated Importance-Aware Federated Sample Selection with Meta Learning. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 3855–3866
2025
-
[28]
Yan Zhang, Guojiang Shen, Xiao Han, Wei Wang, and Xiangjie Kong. 2022. Spatio-temporal digraph convolutional network-based taxi pickup location rec- ommendation. IEEE Transactions on Industrial Informatics 19, 1 (2022), 394–403
2022
-
[29]
Yudong Zhang, Xu Wang, Pengkun Wang, Binwu Wang, Zhengyang Zhou, and Yang Wang. 2024. Modeling spatio-temporal mobility across data silos via per- sonalized federated learning. IEEE Transactions on Mobile Computing (2024)
2024
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.