Pith. sign in

REVIEW 3 major objections 5 minor 110 references

Client-Centric Federated Adaptive Optimization

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A client-centric federated learning framework achieves the best-known convergence rate for nonconvex objectives while tolerating asynchronous participation and heterogeneous local compute.

desk verdict Genuinely novel combination of buffered async aggregation with server-side adaptivity, but the 'arbitrary participation' claim in the abstract is not backed by the main theorem, which assumes uniform random participation. read the letter →

arxiv 2501.09946 v1 pith:AMZQF4VB submitted 2025-01-17 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords federatedlearningclient-centricadaptiveoptimizationasynchronousaggregationsystemheterogeneityclientdriftnonconvexconvergenceAMSGrad
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes Client-Centric Federated Adaptive Optimization (CC-FedAdagrad, CC-FedAdam, CC-FedAMS), a class of federated learning algorithms in which clients decide whether to participate, choose a device-dependent number of local steps, and send updates asynchronously to a server that updates the global model as soon as a fixed number of responses arrive. The paper's central claim is that despite this asynchronous, heterogeneous schedule, the algorithms converge for general nonconvex objectives at the rate O($\sqrt$(1/(AET))) + O($E^{2}$/T) + O($tau^{2}$/T), where A is the number of participating clients per round, E the local epochs, and tau the maximum delay; for long runs with moderate delay this reduces to O($\sqrt$(1/(AET))), matching the best-known rate in asynchronous federated learning. If correct, this means real-world system heterogeneity—stragglers, variable local compute, outdated global-model views—need not degrade the worst-case convergence rate, and server-side adaptive optimization can regularize client drift while preserving linear speedup in A and E. The paper reports experiments on CIFAR-10, CIFAR-100, Fashion-MNIST, and StackOverflow where the proposed methods outperform FedAvg-based baselines by large margins. The guarantee rests on the assumption that each client is included in each round uniformly at random with probability A/B, which is not the fully arbitrary participation the introduction emphasizes.

What carries the argument

The central object is the server-side adaptive update rule: the server maintains a momentum buffer m_t = (1-$\beta$)Delta_t + $\beta$ m_{t-1} and a non-decreasing second-moment estimate v_hat_t (AMSGrad-style), and updates x_{t+1} = x_t - eta m_t/$\sqrt$(v_hat_t + epsilon), where Delta_t is the average of client model differences normalized by their local step counts. Three supporting mechanisms carry the argument: (1) the normalized model update $\Delta$ = (x_tau - x_local)/E, which stops clients running more local epochs from biasing the aggregate; (2) a size-A client buffer, so the server updates once A responses arrive and no client waits for stragglers, with the maximum delay tau bounded; and (3) a Lyapunov sequence z_t = x_t + ($\beta$/(1-$\beta$))(x_t - x_{t-1}) that re-centers the momentum in the proof, letting the adaptive step be analyzed as a stochastic descent of the global loss.

What would settle it

Construct the two-client example in Remark 4.1.2 (f(x) = (1/2)[(x+K)^2 + (x-K)^2]) and run the algorithm with client 1 never participating; the paper's lower-bound argument predicts any algorithm is stuck with E[||grad f||^2] = $\Omega$($sigma_g^{2}$), so if the method converges past that floor the participation model is not the deciding factor. Alternatively, measure the average squared gradient norm on a small nonconvex problem under uniform independent participation with probability A/B and bounded delay; the rate should follow O(1/$\sqrt$(T)) for large T.

Watch

Extended reading notes

Core claim

The core claim is that server-side adaptive optimization—treating the averaged, normalized client updates as a pseudo-gradient and updating the global model with an Adagrad/Adam/AMSGrad-style rule—is compatible with client-centric system heterogeneity. Theorem 4.1 and Corollary 4.1.1 establish that under smoothness, bounded gradients, bounded local and global variance, bounded delay tau, and uniform independent participation with probability A/B, the average squared gradient norm is at most O($\sqrt$(1/(AET))) + O($E^{2}$/T) + O($tau^{2}$/T), with the first term dominating for large T and moderate tau. This matches the best-known convergence rate in asynchronous federated learning and gives linear speedup in both the number of participating clients A and the number of local epochs E. The proof uses a normalized model update to debias clients with more local steps, a size-A buffer to bound the delay, and a Lyapunov sequence that absorbs the server-side momentum so the adaptive step becomes a controlled descent step. The authors note that for fully arbitrary participation, an unavoidable $\Omega$($sigma_g^{2}$) term appears, so the advertised 'clients participate whenever they want' is analyzed only under a specific uniform random participation pattern.

Load-bearing premise

The theorem assumes each client is included in the round's update set uniformly at random with probability A/B, independently; the paper's motivating promise that clients self-determine participation is not what the guarantee actually covers, and the authors concede an unavoidable error floor under fully arbitrary participation in Remark 4.1.2.

Editorial extensions

If this is right

  • If the convergence guarantee is correct, federated systems can let clients choose their own participation and local compute without sacrificing the O(sqrt(1/(AET))) rate, as long as participation is uniform random with probability A/B and delays are bounded.
  • The linear speedup in A and E means that recruiting more clients per round or allowing more local epochs directly shortens the number of rounds needed to reach a target accuracy, preserving the communication-efficiency advantage of FedAvg.
  • The O(tau^2/T) term vanishes asymptotically when the maximum delay is moderate (tau <= (T/(AE))^{1/4}), so stale global-model views do not permanently hurt convergence.
  • Empirical results across four benchmarks show accuracy gains of roughly 10-30% over FedAvg-based baselines under the same asynchronous and heterogeneous schedule, with gains largest on tasks with sparse features such as StackOverflow.
  • Because the framework treats any adaptive optimizer as a plug-in server module, the same convergence analysis is claimed to extend to other adaptive optimizers beyond the three instantiations shown.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The uniform random participation assumption is doing real work: outside the theorem, the authors' own Remark 4.1.2 shows an unavoidable Omega(sigma_g^2) floor, so practitioners using the algorithm under participation patterns determined by battery or network conditions should expect that floor to appear unless they correct for the bias.
  • The normalized model update is the ingredient that makes heterogeneous local epochs benign; a natural ablation comparing CC-FedAdam with and without normalization under identical conditions would isolate how much of the reported accuracy gain comes from this debiasing rather than from adaptivity.
  • As the delay tau goes to zero and the buffer size A grows to the full client set, CC-FedAdam should reduce to the synchronous FedAdam of prior work; verifying that final accuracy tracks this limit would confirm that the gains come from adaptivity rather than from the asynchronous schedule alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Client-Centric Federated Adaptive Optimization (CC-FedAdam, CC-FedAdagrad, CC-FedAMS), a class of federated adaptive optimizers that combine asynchronous server aggregation, time-varying and device-dependent local computation, and client-determined participation. The main theoretical contribution is Theorem 4.1, which under Assumptions 1–4, bounded delay, and a uniform random participation model gives a convergence rate for general nonconvex objectives of O(sqrt(1/(AET))) + O(E^2/T) + O(tau^2/T), improving to O(sqrt(1/(AET))) for large T and moderate lag. Experiments on CIFAR-10, CIFAR-100, Fashion-MNIST, and StackOverflow compare the proposed methods against CC-FedSGD and CC-FedAvg.

Significance. If the stated convergence result holds, the paper makes a useful step by combining server-side adaptivity with asynchronous aggregation and heterogeneous local computation in a single framework. The proof is extensive and self-contained, and the rate matches the best known rate for synchronous and asynchronous federated optimization under the uniform-participation assumption. The paper also provides an honest Remark 4.1.2 showing that arbitrary participation introduces an unavoidable Omega(sigma_g^2) term, which is a valuable observation. However, the advertised feature of arbitrary client participation is not covered by Theorem 4.1, and the empirical comparisons do not include the most relevant asynchronous or adaptive baselines. The core theoretical result is narrower than the paper's framing suggests.

major comments (3)
  1. [Theorem 4.1 and Remark 4.1.2] The abstract and Section 1 advertise 'arbitrary client participation' as a central feature, but Theorem 4.1 requires that 'each client is included in S_t with probability A/B uniformly and independently.' This is a specific sampling model, not arbitrary availability. Remark 4.1.2 concedes that without this assumption the rate degrades to O(sqrt(1/(AET))) + O(E^2/T) + O(tau^2/T) + Omega(sigma_g^2), and it constructs a two-client example in which a never-participating client makes the Omega(sigma_g^2) term unavoidable. Consequently, the paper's headline claim of converging with the best-known rate under arbitrary participation is unsupported; the theorem covers only uniformly random participation. This mismatch between the advertised scope and the proven guarantee is load-bearing and should be reconciled, either by restricting the claims or by presenting the arbitrary-participation result as a convergence-to-neighborhood guarantee.
  2. [Appendix B, Theorem 4.1 proof] The statement of Theorem 4.1 says clients are included 'independently,' but Algorithm 2 uses a fixed-size buffer of A clients and the proof in Appendix B uses without-replacement probabilities P{i,j in S_t} = A(A-1)/(B(B-1)). These two models are inconsistent: independent Bernoulli inclusion with probability A/B would give P{i,j in S_t} = A^2/B^2. Since the proof's cancellation step relies on the without-replacement probability, the formal assumption in Theorem 4.1 should be corrected to state without-replacement sampling, or the proof and algorithm should be adjusted to match the stated independence assumption.
  3. [Section 5.2, Tables 1-2, Figures 1-3] The empirical evaluation compares only against CC-FedSGD and CC-FedAvg, which are variants of the proposed framework rather than established asynchronous or adaptive baselines such as FedAsync, FedBuff, AFL, or FedAdam. The claim that the approaches 'consistently outperform the baseline by a large margin' is therefore not yet supported against the most relevant prior work. Additionally, no error bars or multiple-seed results are reported, so the stability of the observed margins is unknown. The experiments should be extended with the relevant baselines and repeated-seed variability to substantiate the empirical contribution.
minor comments (5)
  1. [Algorithm 2] Lines 11-14 in Algorithm 2 mix the update rules for AdaGrad and Adam/AMSGrad without a clear conditional structure; the pseudocode would be clearer if it explicitly showed three separate variants or labeled each line with the corresponding optimizer.
  2. [Corollary 4.1.1] The condition 'T >= A E^5' and tau <= (T/(A E))^{1/4} in Corollary 4.1.1 is stated without intuition; adding a sentence explaining how these regimes relate to the O(tau^2/T) and O(E^2/T) terms would improve readability.
  3. [Section 1 and 3] The phrase 'arbitrary client participation' is used in multiple places (e.g., Section 1 bullet list and Section 3.2) before the uniform-participation assumption of Theorem 4.1 is introduced; this creates a misleading first impression of the theoretical scope. Please qualify the claim in the introduction.
  4. [Appendix B, B.1] The proof uses the notation 1/E_t both as the reciprocal of an average and in sums over clients; the definition in footnote 5 is helpful but the notation is easy to confuse with a simple reciprocal. A distinct symbol, e.g., r_t, would be clearer.
  5. [References] Some references are incomplete or inconsistent (e.g., the ACM reference format on the first page shows a placeholder year and DOI). Please ensure all bibliographic entries are complete.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the convergence analysis is a self-contained derivation from explicit assumptions; the uniform-participation caveat is a scope limitation, not a circular step.

full rationale

The paper's main claim (Theorem 4.1 and Corollary 4.1.1) is a self-contained proof-theoretic derivation. The bound is obtained from Assumptions 1–4 plus bounded delay and the explicitly stated uniform-participation condition; Appendix B supplies the Lyapunov argument, and every constant in the final rate is either a problem parameter (A, E, T, tau, sigma_g, sigma_l) or a standard hyperparameter (eta, eta_l, epsilon). The choices eta_l = Theta(1/sqrt(T)) and eta = Theta(sqrt(A E)) are complexity-optimizing assignments, not fits to data, and the condition B1 eta_l^2 + B2 eta_l <= epsilon^2 is a proof constraint used to cancel error terms, not a target-dependent fitted quantity. The paper itself flags the scope limitation: Remark 4.1.2 concedes that without uniform random participation an unavoidable Omega(sigma_g^2) term appears, and Theorem 4.1 does not cover the advertised 'arbitrary' participation. This is a correctness/scope caveat, not circularity: the theorem does not define arbitrary participation as uniform, and the algorithm description (self-determined participation, buffer size A) is not made equal to the theorem's assumption. Self-citations appear in literature lists but are not load-bearing for the convergence proof; no uniqueness theorem or prior-work ansatz is invoked to force the result. Hence the derivation is not equivalent to its inputs by construction.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The theoretical result rests on standard smoothness, bounded gradient, bounded variance, bounded delay, and uniform random participation assumptions. No data-fitted free parameters enter the convergence proof; the listed experimental hyperparameters are tuned by grid search and are not part of the derivation. No new physical or mathematical entities are postulated; 'Client-Centric FL' is a framing label, not an invented entity.

free parameters (3)
  • Server learning rate eta = Swept over {10^-3, 10^-2.5, ..., 10^1}, best per benchmark
    Used in all experiments; reported curves are the best over this grid, which can inflate apparent performance gains.
  • Adaptivity parameter epsilon = Swept over {10^-4, 10^-3.5, ..., 10^-1}
    Controls the degree of adaptivity and appears in the theorem conditions and bounds; experimental sensitivity is shown.
  • Momentum beta and second-moment gamma = beta = 0.9, gamma = 0.99 default; sensitivity in Figure 3
    Follow FedAdam defaults and are swept in hyperparameter analysis; the theoretical proof treats them as hyperparameters, not fitted constants.
assumptions (6)
  • domain assumption Each local loss f_i has L-Lipschitz continuous gradients and the global objective has a finite optimal value.
    Assumption 1, standard in FL analysis; used throughout the proof to bound the descent of the Lyapunov function.
  • domain assumption The stochastic gradient estimator is unbiased and bounded in norm: E[g] = grad f_i and ||g|| <= G.
    Assumption 2, standard but strong for adaptive methods; used repeatedly to bound momentum, variance, and the adaptive denominator.
  • domain assumption Bounded local variance: E||g - grad f_i||^2 <= sigma_l^2.
    Assumption 3; contributes the sigma_l^2 term in the final convergence bound.
  • domain assumption Bounded global variance: (1/B) sum_i ||grad f_i - grad f||^2 <= sigma_g^2.
    Assumption 4; controls statistical heterogeneity and contributes the sigma_g^2 term in the final bound.
  • domain assumption Maximum random delay is bounded: tau_{t,i} <= tau < infinity.
    Used before Theorem 4.1 to bound the asynchrony and appears in the O(tau^2/T) term.
  • domain assumption Each client is included in S_t with probability A/B uniformly and independently.
    Explicit in Theorem 4.1; this is the load-bearing participation model without which the claimed convergence rate does not hold, as Remark 4.1.2 acknowledges.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Client-Centric Federated Adaptive Optimization." pith.science (2026). https://pith.science/paper/AMZQF4VB

@misc{pith2026250109946,
  author       = {Pith},
  title        = {Pith review of: Client-Centric Federated Adaptive Optimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AMZQF4VB}},
  note         = {Machine review of arXiv:2501.09946}
}
read the original abstract

Federated Learning (FL) is a distributed learning paradigm where clients collaboratively train a model while keeping their own data private. With an increasing scale of clients and models, FL encounters two key challenges, client drift due to a high degree of statistical/system heterogeneity, and lack of adaptivity. However, most existing FL research is based on unrealistic assumptions that virtually ignore system heterogeneity. In this paper, we propose Client-Centric Federated Adaptive Optimization, which is a class of novel federated adaptive optimization approaches. We enable several features in this framework such as arbitrary client participation, asynchronous server aggregation, and heterogeneous local computing, which are ubiquitous in real-world FL systems but are missed in most existing works. We provide a rigorous convergence analysis of our proposed framework for general nonconvex objectives, which is shown to converge with the best-known rate. Extensive experiments show that our approaches consistently outperform the baseline by a large margin across benchmarks.

Figures

Figures reproduced from arXiv: 2501.09946 by the authors.

Figure 1
Figure 1. Training and testing curves for various CC-Federa [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Testing curve for various CC-Federated Adaptive O [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Hyperparameter sensitivity of CC-Federated Adap [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: A toy example of a two-client FL setting. There is a m [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Training/Testing Curves in Figure 2(a), i.e. expe [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: Training/Testing Curves in Figure 2(b), i.e. expe [PITH_FULL_IMAGE:figures/full_fig_p024_6.png]
Figure 7
Figure 7. Figure 7: Training/Testing Curves in Figure 2(c), i.e. expe [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: Training and testing curves for various CC-Federa [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Training and testing curves for various CC-Federa [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]
Figure 10
Figure 10. Figure 10: Training and testing curves for various CC-Feder [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

110 extracted references · 58 canonical work pages

  1. [1]

    Durmus Alp Emre Acar, Yue Zhao, Ramon Matas, Matthew Matt ina, Paul What- mough, and Venkatesh Saligrama. 2021. Federated Learning B ased on Dy- namic Regularization. In International Conference on Learning Representations . https://openreview.net/forum?id=B7v4QMR6Z9w

  2. [2]

    Maruan Al-Shedivat, Jennifer Gillenwater, Eric Xing, a nd Afshin Rostamizadeh

  3. [3]

    Dmitrii Avdiukhin and Shiva Kasiviswanathan. 2021. Fed erated Learning under Arbitrary Communication Patterns. In Proceedings of the 38th Inter- national Conference on Machine Learning (Proceedings of Mac hine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 425–435. https://proceedings.mlr.press/v139/avdiukhin21a.html

  4. [4]

    Runxue Bao, Bin Gu, and Heng Huang. 2020. Fast oscar and ow l regression via safe screening rules. In International conference on machine learning. PMLR, 653–663

  5. [5]

    Runxue Bao, Xidong Wu, Wenhan Xian, and Heng Huang. 2022. Doubly Sparse Asynchronous Learning. In The 31st International Joint Conference on Artificial Intelligence (IJCAI 2022)

  6. [6]

    Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggav i. 2019. Qsparse- Local-SGD: Distributed SGD with Quantization, Sparsificat ion, and Local Com- putations. Curran Associates Inc., Red Hook, NY, USA

  7. [7]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, J ared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Giri sh Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  8. [8]

    PRESTON BUKATY. 2019. The California Consumer Privacy Act (CCPA): An implementation guide . IT Governance Publishing. http://www.jstor.org/stable/j.ctvjghvnn

Show all 110 references
  1. [9]

    Liwei Che, Zewei Long, Jiaqi Wang, Yaqing Wang, Houping Xiao, and Fenglong Ma. 2021. FedTriNet: A pseudo labeling method with three players for federated semi-supervised learning. In 2021 IEEE International Conference on Big Data (Big Data). IEEE, 715–724

  2. [10]

    Liwei Che, Jiaqi Wang, Yao Zhou, and Fenglong Ma. 2023. Multimodal federated learning: A survey. Sensors 23, 15 (2023), 6986

  3. [11]

    Wenlin Chen, Samuel Horváth, and Peter Richtárik. 2020 . Optimal Client Sam- pling for Federated Learning. ArXiv abs/2010.13723 (2020)

  4. [12]

    Xiangyi Chen, Xiaoyun Li, and Ping Li. 2020. Toward Comm unication Efficient Adaptive Gradient Method. In Proceedings of the 2020 ACM-IMS on Foundations of Data Science Conference (Virtual Event, USA) (FODS ’20). Association for Computing Machinery, New York, NY, USA, 11 9–128. ...

  5. [13]

    Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong. 201 9. On the Convergence of A Class of Adam-Type Algorithms for Non-Conv ex Optimization. In International Conference on Learning Representations . https://openreview.net/forum?id=H1x-x309tm

  6. [14]

    Aradhye, Glen Anderson, Gregory S

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shak ed, Tushar Chandra, Hrishikesh B. Aradhye, Glen Anderson, Gregory S. Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, X iaobing Liu, and Hemal Shah. 2016. Wide & Deep Learning for Recom...

  7. [15]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for langu age understanding. arXiv preprint arXiv:1810.04805 (2018)

  8. [16]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Under- standing. In North American Chapter of the Association for Computationa l Lin- guistics. https://api.semanticscholar.org/CorpusID:52967399

  9. [17]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesniko v, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matth ias Min- derer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and N eil Houlsby

  10. [18]

    John Duchi, Elad Hazan, and Yoram Singer. 2011. Adaptiv e Sub- gradient Methods for Online Learning and Stochastic Optimi za- tion. Journal of Machine Learning Research 12, 61 (2011), 2121–2159. http://jmlr.org/papers/v12/duchi11a.html

  11. [19]

    In International Conference on Learning Representations

    An Image is Worth 16x16 Words: Transformers for Image R ecog- nition at Scale. In International Conference on Learning Representations . https://openreview.net/forum?id=YicbFdNTTy

  12. [20]

    Jack Goetz, Kshitiz Malik, Duc Viet Bui, Seungwhan Moon , Honglei Liu, and Anuj Kumar. 2019. Active Federated Learning. ArXiv abs/1909.12641 (2019)

  13. [21]

    European Commission. 2016. Regulation (EU) 2016/679 o f the European Parlia- ment and of the Council of 27 April 2016 on the protection of na tural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (G...

  14. [22]

    Xinran Gu, Kaixuan Huang, Jingzhao Zhang, and Longbo Hu ang

  15. [23]

    Girshick, Pieter Noo rdhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He

    Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noo rdhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He

  16. [24]

    Fengxiang He, Tongliang Liu, and Dacheng Tao. 2019. Con trol Batch Size and Learning Rate to Generalize Well: Theoretical and Empir ical Evidence. In Advances in Neural Information Processing Systems 32 . Curran Associates, Inc., 1143–1152

  17. [25]

    Zhang, Shaoqing Ren, and Jian Sun

    Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2016. D eep Residual Learn- ing for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016), 770–778

  18. [26]

    In Advances in Neural Information Processing Systems , A

    Fast Federated Learning in the Presence of Arbitrary D evice Unavailability. In Advances in Neural Information Processing Systems , A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (Eds.). https://openreview.net/forum?id=1_gaHBaRYt

  19. [27]

    Hamilton, Zhitao Ying, and Jure Leskovec

    William L. Hamilton, Zhitao Ying, and Jure Leskovec. 20 17. Inductive Repre- sentation Learning on Large Graphs. In Neural Information Processing Systems . https://api.semanticscholar.org/CorpusID:4755450

  20. [28]

    Chun-Yin Huang, Kartik Srinivas, Xin Zhang, and Xiaoxi ao Li. 2024. Overcom- ing Data and Model Heterogeneities in Decentralized Federa ted Learning via Synthetic Anchors. arXiv preprint arXiv:2405.11525 (2024)

  21. [29]

    Huang, Z

    G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger. 2017. Densely Connected Convolutional Networks. In 2017 IEEE Conference on Computer Vi- sion and Pattern Recognition (CVPR) . 2261–2269

  22. [30]

    K. He, X. Zhang, S. Ren, and J. Sun. 2016. Deep Residual Le arning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR). 770–778

  23. [31]

    Tzu-Ming Harry Hsu, Qi, and Matthew Brown. 2019. Measur ing the Effects of Non-Identical Data Distribution for Federated Visual Cl assification. ArXiv abs/1909.06335 (2019)

  24. [32]

    Brendan McMahan, Brendan Avent, Auré lien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista A

    Peter Kairouz, H. Brendan McMahan, Brendan Avent, Auré lien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista A. Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Hubert E ichner, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garret...

  25. [33]

    Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebas- tian Stich, and Ananda Theertha Suresh. 2020. Scaffold: Stoc hastic controlled averaging for federated learning. In International Conference on Machine Learn- ing. PMLR, 5132–5143

  26. [34]

    Stanisław Jastrzębski, Zac Kenton, Devansh Arpit, Nic olas Ballas, Asja Fischer, Amos Storkey, and Yoshua Bengio. 2018. Three factors influen cing minima in SGD. https://openreview.net/forum?id=rJma2bZCW

  27. [35]

    Yae Jee Cho, Jianyu Wang, and Gauri Joshi. 2022. Towards Under- standing Biased Client Selection in Federated Learning. In Proceedings of The 25th International Conference on Artificial Intellig ence and Statis- tics (Proceedings of Machine Learning Research, Vol. 151) , Gustau...

  28. [36]

    Alex Krizhevsky. 2009. Learning Multiple Layers of Fea tures from Tiny Images

  29. [37]

    Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018. Vi- sualizing the Loss Landscape of Neural Nets. In Proceedings of the 32nd Interna- tional Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18). Curran Associates Inc., Red ...

  30. [38]

    Prashant Khanduri, Pranay Sharma, Haibo Yang, Mingyi H ong, Jia Liu, Ketan Rajawat, and Pramod Varshney. 2021. Stem: A stochastic two- sided momen- tum algorithm achieving near-optimal sample and communica tion complexi- ties for federated learning. Advances in Neural Informat...

  31. [39]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. CoRR abs/1412.6980 (2015)

  32. [40]

    Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi , Ameet Talwalkar, and Virginia Smith. 2020. Federated Optimization in Hetero geneous Networks. In Proceedings of Machine Learning and Systems , I. Dhillon, D. Papailiopoulos, and V. Sze (Eds.), Vol. 2. 429–450

  33. [41]

    Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi , Ameet Talwalkar, and Virginia Smith. 2020. Federated optimization in hetero geneous networks. Proceedings of Machine Learning and Systems 2 (2020), 429–450

  34. [42]

    Qinbin Li, Yiqun Diao, Quan Chen, and Bingsheng He. 2022 . Federated Learn- ing on Non-IID Data Silos: An Experimental Study. In IEEE International Con- ference on Data Engineering

  35. [43]

    Qinbin Li, Zeyi Wen, Zhaomin Wu, Sixu Hu, Naibo Wang, Yua n Li, Xu Liu, and Bingsheng He. 2021. A survey on federated learning systems: vision, hype and reality for data privacy and protection. IEEE Transactions on Knowledge and Data Engineering (2021). Client-Centric Federate...

  36. [44]

    Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu. 2015 . Asynchronous Paral- lel Stochastic Gradient for Nonconvex Optimization. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS’15). MIT Press, C...

  37. [45]

    Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, X iaodong Liu, Jian- feng Gao, and Jiawei Han. 2020. On the Variance of the Adaptiv e Learn- ing Rate and Beyond. In International Conference on Learning Representations . https://openreview.net/forum?id=rkgz2aEKDr

  38. [46]

    Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Z hi- hua Zhang. 2020. On the Convergence of FedAvg on Non- IID Data. In International Conference on Learning Representations . https://openreview.net/forum?id=HJxNAnVtDS

  39. [47]

    Xinjin Li, Yu Ma, Yangchen Huang, Xingqi Wang, Yuzhen Li n, and Chenxi Zhang. 2024. Integrated Optimization of Large Language Mod els: Synergizing Data Utilization and Compression Techniques. (2024)

  40. [48]

    Liangchen Luo, Yuanhao Xiong, and Yan Liu. 2019. Adapti ve Gradient Methods with Dynamic Bound of Learning Rate. In International Conference on Learning Representations. https://openreview.net/forum?id=Bkg3g2R9FX

  41. [49]

    Yu Mao, Yufei Cui, Tei-Wei Kuo, and Chun Jason Xue. 2022. Accelerating general-purpose lossless compression via simple and scalable parameterization. In Proceedings of the 30th ACM International Conference on Mult imedia. 3205– 3213

  42. [50]

    Weidong Liu, Xiaojun Mao, Xiaofei Zhang, and Xin Zhang. 2024. Robust Per- sonalized Federated Learning with Sparse Penalization. J. Amer. Statist. Assoc. (2024), 1–12

  43. [51]

    Ben London. 2017. A PAC-Bayesian Analysis of Randomize d Learning with Application to Stochastic Gradient Descent. In NIPS

  44. [52]

    Yongsheng Mei, Liangqi Yuan, Dong-Jun Han, Kevin S Chan , Christopher G Brinton, and Tian Lan. 2024. Using Diffusion Models as Genera tive Re- play in Continual Federated Learning–What will Happen? arXiv preprint arXiv:2411.06618 (2024)

  45. [53]

    V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, Ioannis A ntonoglou, Daan Wier- stra, and Martin A. Riedmiller. 2013. Playing Atari with Dee p Reinforcement Learning. ArXiv abs/1312.5602 (2013)

  46. [54]

    Yu Mao, Yufei Cui, Tei-Wei Kuo, and Chun Jason Xue. 2022. Trace: A fast transformer-based general-purpose lossless compressor. In Proceedings of the ACM Web Conference 2022. 1829–1838

  47. [55]

    H. B. McMahan, Eider Moore, Daniel Ramage, Seth Hampson , and Blaise Agüera y Arcas. 2017. Communication-Efficient Learni ng of Deep Net- works from Decentralized Data. In International Conference on Artificial Intelli- gence and Statistics

  48. [56]

    Qi Qi, Youzhi Luo, Zhao Xu, Shuiwang Ji, and Tianbao Yang . 2021. Stochastic optimization of areas under precision-recall curves with provable convergence. Advances in neural information processing systems 34 (2021), 1752–1765

  49. [57]

    Qi Qi, Jiameng Lyu, Er Wei Bai, Tianbao Yang, et al. 2022. Stochastic con- strained dro with a complexity independent of sample size. arXiv preprint arXiv:2210.05740 (2022)

  50. [58]

    Rabbat, Mani Malek, and Dzmitry Huba

    John Nguyen, Kshitiz Malik, Hongyuan Zhan, Ashkan Yous efpour, Michael G. Rabbat, Mani Malek, and Dzmitry Huba. 2021. Federated Learn ing with Buffered Asynchronous Aggregation. In International Conference on Artificial Intelligence and Statistics

  51. [59]

    Takayuki Nishio and Ryo Yonetani. 2018. Client Selecti on for Federated Learn- ing with Heterogeneous Resources in Mobile Edge. ICC 2019 - 2019 IEEE Inter- national Conference on Communications (ICC) (2018), 1–7

  52. [60]

    Mónica Ribero and Haris Vikalo. 2020. Communication-E fficient Federated Learning via Optimal Client Sampling. ArXiv abs/2007.15197 (2020)

  53. [61]

    Yichen Ruan, Xiaoxi Zhang, Shu-Che Liang, and Carlee Jo e-Wong. 2021. To- wards Flexible Device Participation in Federated Learning. In International Con- ference on Artificial Intelligence and Statistics

  54. [62]

    Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachar y Garrett, Keith Rush, Jakub Konečn `y, Sanjiv Kumar, and H Brendan McMahan. 2020. Adaptive fed- erated optimization. arXiv preprint arXiv:2003.00295 (2020)

  55. [63]

    Reddi, Satyen Kale, and Sanjiv Kumar

    Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. 2018. O n the Convergence of Adam and Beyond. In International Conference on Learning Representations . https://openreview.net/forum?id=ryQu7f-RZ

  56. [64]

    Sebastian U. Stich. 2019. Local SGD Converges Fast and C ommu- nicates Little. In International Conference on Learning Representations . https://openreview.net/forum?id=S1g2JnRcFX

  57. [65]

    Jianhui Sun, Mengdi Huai, Kishlay Jha, and Aidong Zhang . 2022. De- mystify Hyperparameters for Stochastic Optimization with Transferable Representations. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22) . A...

  58. [66]

    Karen Simonyan and Andrew Zisserman. 2015. Very Deep Co nvolutional Net- works for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Confer- ence Track Proceedings. http://arxiv.org/abs/1409.1556

  59. [67]

    Smith, Pieter-Jan Kindermans, and Quoc V

    Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le. 2018. Don’t Decay the Learning Rate, Increase the Batch Size. In International Conference on Learning Representations. https://openreview.net/forum?id=B1Yy1BxCZ

  60. [68]

    Jianhui Sun, Ying Yang, Guangxu Xun, and Aidong Zhang. 2 021. A Stagewise Hyperparameter Scheduler to Improve Generalization. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Min ing (Virtual Event, Singapore) (KDD ’21). Association for Computing ...

  61. [69]

    Jianhui Sun, Ying Yang, Guangxu Xun, and Aidong Zhang. 2 023. Scheduling Hyperparameters to Improve Generalization: From Centrali zed SGD to Asyn- chronous SGD. ACM Trans. Knowl. Discov. Data 17, 2, Article 29 (mar 2023), 37 pages. https://doi.org/10.1145/3544782

  62. [70]

    Jianhui Sun, Sanchit Sinha, and Aidong Zhang. 2023. Enh ance Diffusion to Improve Robust Generalization. In Proceedings of the 29th ACM SIGKDD Con- ference on Knowledge Discovery and Data Mining (Long Beach, CA, USA) (KDD ’23). Association for Computing Machinery, New York, NY,...

  63. [71]

    Jianhui Sun, Xidong Wu, Heng Huang, and Aidong Zhang. 20 24. On the role of server momentum in federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 15164–15172

  64. [72]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. At tention is all you need. Advances in neural information processing systems 30 (2017)

  65. [73]

    Haohui Wang, Yuzhen Mao, Yujun Yan, Yaoqing Yang, Jianh ui Sun, Kevin Choi, Balaji Veeramani, Alison Hu, Edward Bowen, Tyler Cody, and D awei Zhou

  66. [74]

    Qiuling Suo, Liuyi Yao, Guangxu Xun, Jianhui Sun, and Ai dong Zhang. 2019. Recurrent Imputation for Multivariate Time Series with Missing Values. In2019 IEEE International Conference on Healthcare Informatics,ICHI 2019, Xi’an, China, June 10-13, 2019 . IEEE, 1–3. https://doi.o...

  67. [75]

    Tijmen Tieleman, Geoffrey Hinton, et al. 2012. Lecture 6 .5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4, 2 (2012), 26–31

  68. [76]

    Charles, Zheng Xu, Gauri Joshi, H

    Jianyu Wang, Zachary B. Charles, Zheng Xu, Gauri Joshi, H. B. McMahan, Blaise Agüera y Arcas, Maruan Al-Shedivat, Galen Andrew, Sa lman Aves- timehr, Katharine Daly, Deepesh Data, Suhas N. Diggavi, Hub ert Eichner, Advait Gadhikar, Zachary Garrett, Antonious M. Girgis, Fil ip ...

  69. [77]

    Vincent Poor

    Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H. Vincent Poor. 2020. Tackling the Objective Inconsistency Problem in Heterogeneous Federated Op- timization. In Proceedings of the 34th International Conference on Neural I nfor- mation Processing Systems (Vancouver, BC, ...

  70. [78]

    Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Micha el Rabbat. 2020. SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum. In International Conference on Learning Representations . https://openreview.net/forum?id=SkxJ8REYPH

  71. [79]

    Han Wang, Jun Tang, Xiaodong Liu, Shanyan Guan, Rong Xie , and Li Song

  72. [80]

    Yujia Wang, Lu Lin, and Jinghui Chen. 2022. Communicati on-Efficient Adap- tive Federated Learning. In Proceedings of the 39th International Conference on Conference’17, July 2017, Washington, DC, USA Jianhui Sun, X idong Wu, Heng Huang, and Aidong Zhang Machine Learning (Procee...

  73. [81]

    Haoyu Wang, Handong Zhao, Yaqing Wang, Tong Yu, Jiuxian g Gu, and Jing Gao. 2022. FedKC: Federated knowledge composition for mult ilingual natural language understanding. In Proceedings of the ACM Web Conference 2022. 1839– 1850

  74. [82]

    Jiawei Wen, Songshan Yang, Chris Wang, Yishen Jiang, an d Runze Li. 2023. Feature-splitting algorithms for ultrahig h di- mensional quantile regression. Journal of Econometrics (2023). https://api.semanticscholar.org/CorpusID:257747996

  75. [83]

    Jiawei Wen, Songshan Yang, and Delin Zhao. 2024. Noncon vex Dantzig selector and its parallel computing algorithm. Statistics and Computing 34, 6 (Sept. 2024), 21 pages. https://doi.org/10.1007/s11222-024-10492-8

  76. [84]

    Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati S rebro, and Benjamin Recht. 2017. The Marginal Value of Adaptive Gradient Method s in Machine Learning. In Advances in Neural Information Processing Systems 30 . Curran Associates, Inc., 4148–4158

  77. [85]

    Shiqiang Wang and Mingyue Ji. 2022. A Unified Analysis of Federated Learning with Arbitrary Client Participation. In Advances in Neural Information Process- ing Systems , Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghy un Cho (Eds.). https://openreview.net/forum?id=q...

  78. [86]

    Xidong Wu, Jianhui Sun, Zhengmian Hu, Aidong Zhang, and Heng Huang

  79. [87]

    Tianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng, Zhengyan g Wang, Jianhui Sun, Qingyu Yin, Hanqing Lu, Suhang Wang, Jingrui He, et al. 2023. Towards uni- versal multi-modal personalization: A language model empo wered generative paradigm. In The Twelfth International Conference ...

  80. [88]

    Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta. 2019. Asynchronous Fed- erated Optimization. ArXiv abs/1903.03934 (2019)

  81. [89]

    Guangxu Xun, Kishlay Jha, Jianhui Sun, and Aidong Zhang . 2020. Correlation Networks for Extreme Multi-label Text Classification. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discover y & Data Mining (2020)

  82. [90]

    Yikai Yan, Chaoyue Niu, Yucheng Ding, Zhenzhe Zheng, Fa n Wu, Guihai Chen, Shaojie Tang, and Zhihua Wu. 2020. Distributed Non-Co nvex Opti- mization with Sublinear Speedup under Intermittent Client Availability. ArXiv abs/2002.07399 (2020)

  83. [91]

    Xidong Wu, Jianhui Sun, Zhengmian Hu, Junyi Li, Aidong Z hang, and Heng Huang. 2023. Federated Conditional Stochastic Optimizati on. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curr...

  84. [92]

    Haibo Yang, Xin Zhang, Prashant Khanduri, and Jia Liu. 2 021. Anarchic Feder- ated Learning. In International Conference on Machine Learning

  85. [93]

    Greenewald, Trong Nghia Hoang, and Yasaman Khazaeni

    Mikhail Yurochkin, Mayank Agarwal, Soumya Shubhra Gho sh, Kristjan H. Greenewald, Trong Nghia Hoang, and Yasaman Khazaeni. 2019. Bayesian Non- parametric Federated Learning of Neural Networks. In International Conference on Machine Learning

  86. [94]

    Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fash ion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithm s. ArXiv abs/1708.07747 (2017)

  87. [95]

    Matthew D. Zeiler. 2012. ADADELTA: An Adaptive Learnin g Rate Method. ArXiv abs/1212.5701 (2012)

  88. [96]

    Xinwei Zhang, Mingyi Hong, Sairaj Dhople, Wotao Yin, an d Yang Liu. 2020. Fedpd: A federated learning framework with optimal rates an d adaptivity to non-iid data. arXiv preprint arXiv:2005.11418 (2020)

  89. [97]

    Zeyu Zhang, Thuy Vu, and Alessandro Moschitti. 2021. Jo int models for answer verification in question answering systems. arXiv preprint arXiv:2107.04217 (2021)

  90. [98]

    Haibo Yang, Minghong Fang, and Jia Liu. 2021. Achieving Lin- ear Speedup with Partial Worker Participation in Non-IID Fe der- ated Learning. In International Conference on Learning Representations . https://openreview.net/forum?id=jDdzh5ul-d

  91. [99]

    Jiayun Zheng and Maggie Makar. 2022. Causally motivate d multi-shortcut iden- tification and removal. Advances in Neural Information Processing Systems 35 (2022), 12800–12812

  92. [100]

    Hanhan Zhou, Tian Lan, and Vaneet Aggarwal. 2022. PAC: Assisted Value Factorization with Counterfactual Predictions in Multi-A gent Reinforcement Learning. Advances in Neural Information Processing Systems (NeurIPS) 36 (2022), 15757–15769

  93. [101]

    Reddi, Devendra Sachan, Saty en Kale, and Sanjiv Kumar

    Manzil Zaheer, Sashank J. Reddi, Devendra Sachan, Saty en Kale, and Sanjiv Kumar. 2018. Adaptive Methods for Nonconvex Optimization. In Proceed- ings of the 32nd International Conference on Neural Informa tion Processing Sys- tems (Montréal, Canada) (NIPS’18). Curran Associate...

  94. [102]

    Hanhan Zhou, Tian Lan, Guru Prasadh Venkataramani, an d Wenbo Ding. 2023. Every parameter matters: Ensuring the convergence of federated learning with dynamic heterogeneous models reduction. Advances in Neural Information Pro- cessing Systems 37 (2023)

  95. [103]

    acm-jdslogo.png

    Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatiko nda, Nicha Dvornek, Xenophon Papademetris, and James Duncan. 2020. AdaBelief Optimizer: Adapt- ing Stepsizes by the Belief in Observed Gradients. Conference on Neural Infor- mation Processing Systems (2020). Client-Centric ...

  96. [105]

    Ce Zheng, Xianpeng Liu, Guo-Jun Qi, and Chen Chen. 2023. Potter: Pooling attention transformer for efficient human mesh recovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit ion. 1611–1620

  97. [108]

    Hanhan Zhou, Tian Lan, Guru Prasadh Venkataramani, an d Wenbo Ding. 2022. Federated learning with online adaptive heterogeneous loc al models. In Work- shop on Federated Learning: Recent Advances and New Challen ges (in Conjunc- tion with NeurIPS 2022)

  98. [2017]

    CoRR abs/1706.02677 (2017)

    Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour. CoRR abs/1706.02677 (2017). arXiv:1706.02677 http://arxiv.o rg/abs/1706.02677

  99. [2021]

    In International Conference on Learning Representations (ICLR)

    Federated Learning via Posterior Averaging: A New Per spective and Practical Algorithms. In International Conference on Learning Representations (ICLR)

  100. [2022]

    In European Conference on Computer Vision

    Ptseformer: Progressive temporal-spatial enhanced transformer towards video object detection. In European Conference on Computer Vision . Springer, 732–747

  101. [2024]

    Advances in Neural Information Processing Systems 36 (2024)

    Solving a class of non-convex minimax optimization in federated learn- ing. Advances in Neural Information Processing Systems 36 (2024)

  102. [2025]

    In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML’24)

    EvoluNet: advancing dynamic non-IID transfer learni ng on graphs. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML’24). JMLR.org, Article 2095, 19 pages

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.