REVIEW 3 major objections 5 minor 3 cited by
HASFL: Heterogeneity-aware Split Federated Learning over Edge Computing Systems
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that per-device batch sizes and model split points, chosen from a new convergence bound, minimize the training time of split federated learning on heterogeneous edge devices.
desk verdict Solid first convergence bound for SFL with per-device batch sizes, but the latency-minimization claim rests on an unverified tightness assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the convergence bound of Theorem 1 together with its inversion in Corollary 1. The bound decomposes the optimization error into three terms: an initial-optimization-gap term $\frac{2\vartheta}{\gamma R}$ that shrinks with more rounds, a mini-batch variance term $\frac{\beta\gamma}{N^2}\sum_i\sum_j \sigma_j^2/b_i$ that shrinks when any device raises its batch size, and a staleness term $4\beta^2\gamma^2I^2\sum_{j=1}^{L_c}G_j^2$ that appears only when client-side sub-models are aggregated every $I>1$ rounds and grows with the client-side split depth $L_c$. Because the variance term depends on the sum of $\sigma_j^2/b_i$ over devices, the bound predicts that batch sizes compensate across devices: a fast device can carry a large batch while a slow device uses a small one without changing the worst-case convergence guarantee. This bound converts the system design problem into the transformed program $P''$, solved by alternating a closed-form batch-size update (Proposition 1, using the Newton-Jacobi method) with a Dinkelbach-based split-point update.
What would settle it
Record, on the paper's CIFAR-10/VGG-16 setting, the round at which the average squared gradient norm first drops below a target $\varepsilon$ and compare it with the value $R$ predicted by Corollary 1 using the same estimated parameters; a ratio far from 1 across different batch-size assignments would show the bound is too loose for the equality assumption to hold.
Extended reading notes
Core claim
Under $\beta$-smoothness of the local losses and bounded per-layer gradient variance and second moments, the paper proves (Theorem 1) that the average squared gradient norm after $R$ rounds satisfies $\frac{1}{R}\sum_{t=1}^R \mathbb{E}\|\nabla f(w^{t-1})\|^2 \le \frac{2\vartheta}{\gamma R} + \frac{\beta\gamma}{N^2}\sum_{i=1}^N \sum_{j=1}^L \frac{\sigma_j^2}{b_i} + \mathbf{1}_{\{I>1\}} 4\beta^2 \gamma^2 I^2 \sum_{j=1}^{L_c} G_j^2$. Inverting this bound (Corollary 1) gives a lower bound on the rounds $R$ needed to reach accuracy $\varepsilon$, and the paper asserts that the optimal schedule is obtained when this lower bound holds with equality. Substituting that expression for $R$ into a per-round latency model that includes client forward/backward passes, activation upload/download, and the periodic client-side model aggregation yields an explicit latency objective, which is then minimized over per-device batch sizes $b_i$ and binary split-point variables $\mu_{i,j}$. The paper claims this is the first convergence bound for SFL with both varied batch sizes and cut layers, and that the resulting heterogeneity-aware schedule is what enables the reported speedups.
Load-bearing premise
The paper assumes the convergence upper bound is tight enough that the round count it predicts is the round count really needed, so minimizing the bound is the same as minimizing training time; if the bound is loose, the optimized batch sizes and split points can lower the surrogate without lowering real latency.
Editorial extensions
If this is right
- If the bound is tight, the optimal schedule assigns faster devices larger batch sizes and slower devices smaller ones, with each device's cut layer chosen to balance smashed-data communication cost against aggregation frequency.
- For aggregation interval $I>1$, the bound predicts that shallower client-side splits accelerate convergence, so the optimizer sends more layers to the server even when that increases activation traffic, stopping where the latency trade-off balances.
- Because the batch-size rule is nearly closed-form, the schedule can be recomputed cheaply each aggregation round, letting HASFL adapt to drifting device speeds and channel rates during training.
- The experiments report at least 4x faster convergence and about 1% higher accuracy than benchmarks that randomize batch sizes, split points, or both.
Reading between the lines
- The paper never verifies its equality assumption: a reader who logs predicted versus actual rounds to reach $\varepsilon$ on the paper's own testbed could check whether the optimized schedule is truly latency-optimal or only optimal for the surrogate.
- The per-layer variance assumption ($\sigma_j^2/b$ per layer, summed over layers) is a strong structural condition; if modern layers with normalization or adaptive optimizers violate it, the batch-size compensation rule could mis-rank devices.
- Because the server-side common sub-model is assumed to synchronize every round for free, the analysis applies to a single edge server; a hierarchical or multi-hop version of SFL would need a modified staleness term.
- A direct stress test is to compare HASFL's end-to-end measured latency against a brute-force grid search over batch sizes and split points on a small model; the size of any gap quantifies how much is lost by treating the upper bound as exact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes HASFL, a heterogeneity-aware split federated learning framework that adaptively controls per-device batch sizes and model split points to minimize the training latency required to reach a target convergence accuracy. The authors derive a convergence upper bound for SFL with heterogeneous batch sizes and cut layers (Lemma 1 and Theorem 1), then use the lower bound on the required number of training rounds obtained from this bound (Corollary 1, Eq. (27)) to formulate a latency minimization problem. The problem is decomposed into batch-size and model-splitting subproblems, solved alternately by a block-coordinate descent algorithm (Algorithm 2). Experiments on CIFAR-10 and CIFAR-100 with VGG-16 and ResNet-18 report faster convergence and higher accuracy than several benchmarks.
Significance. If the optimization-to-latency bridge were valid, the paper would make a useful contribution to resource-efficient split federated learning: it provides the first convergence bound for SFL with jointly varied batch sizes and cut layers, and the proposed system-level optimization is evaluated across a broad set of scenarios. The theoretical derivations in Lemma 1 and Theorem 1 follow standard smoothness and bounded-variance machinery and appear algebraically sound. The experimental study is extensive and consistently shows improvements over the chosen benchmarks. However, the central claim that the proposed optimization minimizes actual training latency is not established, because the optimization is built on an unexamined equality assumption that replaces the convergence bound with an exact round-count prediction; the evaluation does not validate this assumption. The significance of the paper therefore hinges on a load-bearing point that currently lacks support.
major comments (3)
- [Section VI, Eqs. (27) and (43)] The paper converts a lower bound into an exact equality. Corollary 1 establishes only that R >= R_bound(epsilon, b, mu) for the average squared gradient norm to be at most epsilon; it does not establish that the actual number of rounds equals R_bound. The statement in Section VI that 'the objective function is minimized if and only if (27) holds as an equality' is therefore unsupported. Since Eqs. (42)-(45) and the entire subsequent optimization minimize Theta(b, mu) = R_bound * [T_S + T_A/I], the claimed latency optimality is for a surrogate objective. The manuscript provides no tightness analysis, no comparison of predicted R_bound against observed rounds, and no argument that the ranking of configurations by R_bound matches the ranking by actual R. This is load-bearing for the paper's main claim, and the issue cannot be dismissed as a minor technicality because the bound contains worst-case constants and is derived through multiple relaxations in Lemma 1 and Theorem 1.
- [Section VI, Proposition 1] Proposition 1 claims 'The optimal BS decision' is given by Eq. (48), obtained by solving the first-order condition with the Newton-Jacobi method. The proof, however, only demonstrates that for each i', holding all other batch sizes fixed, the objective is initially decreasing and then increasing in b_i' (Eqs. (49)-(50)). This establishes coordinate-wise unimodality, not joint convexity or global optimality of the simultaneous root. The optimal solution under the coupling constraints C4, R3, and R4 may lie on the boundary even when the unconstrained root does not; the correction using kappa_i is heuristic. The remark's claim that exhaustive search over 3^N combinations identifies the global optimum is also inconsistent with the algorithm actually deployed, which solves the nonlinear system and applies a one-time correction. Consequently, the guarantee that P1 is solved exactly in the BCD algorithm is not proven.
- [Section VII-B, Figs. 5-6] The empirical evaluation does not test the key theoretical bridge. Corollary 1 concerns the average squared gradient norm condition (Eq. (26)), but the convergence criterion in the experiments is a test-accuracy plateau (e.g., 'test accuracy increases by less than 0.02 percent across five consecutive training rounds'). Therefore, the measured speedups in Figs. 5-6 cannot validate the assertion that Eq. (27) accurately predicts the number of rounds required to reach epsilon. The paper should report the actual versus predicted round counts, or measure the gradient-norm criterion directly, to support the equality assumption that underlies the optimization.
minor comments (5)
- [Section II] There is a typo in the related-work discussion: 'incentive mechanisms to to balance the training load' should read 'incentive mechanisms to balance the training load'.
- [Section VI, Eq. (40)] The approximation ceil(R/I) ≈ R/I is used without an error bound. Since the optimization can, in principle, select small R values, the relative error of this approximation should be quantified or justified.
- [Section VI, Proposition 1 remark] The remark states that the optimal solution can be obtained by exhaustive search over 3^N combinations, but the described algorithm solves the nonlinear system and performs a correction step. These two statements describe different procedures, and the paper should clarify which one is implemented and what guarantees each provides.
- [Algorithm 2 and Section VI] Algorithm 2 terminates when the objective change is below epsilon, but no convergence analysis of the BCD scheme is provided. This is acceptable for a heuristic, but the text should avoid implying that the final solution is globally optimal.
- [Figures 2 and 3] The x-axis in Figs. 2(a) and 3(a) is labeled 'Training Rounds' while the captions say 'epochs'; the terminology should be made consistent.
Circularity Check
No circularity: the optimization is built on a derived (and possibly loose) bound, but the bound is not fitted to, nor defined by, the predicted outcome.
full rationale
Theorem 1 is derived directly from Assumptions 1 and 2 using a standard smoothness/descent argument, and the proof is given in full, so the convergence bound is not imported from prior work by construction. Corollary 1 is a rearrangement of Theorem 1's upper bound into a lower bound on the number of rounds R. Section VI substitutes this lower bound into the latency expression and solves the resulting problem P'/P''. The statement that 'the objective function is minimized if and only if (27) holds as an equality' is a tightness/surrogacy assumption: if the bound is loose in a configuration-dependent way, the optimized solution may minimize the surrogate rather than actual wall-clock time. That is a validity concern, not circularity. The bound is not fitted to the experimental outcome, and no equation in the derivation chain is defined in terms of the quantity it is supposed to predict. The constants beta, sigma_j^2, and G_j^2 are estimated following external prior work [24], not tuned to make HASFL's claims true. The self-citations [21] and [23] appear in the introduction and in the justification for per-round server-side common sub-model aggregation, but the convergence proof and the optimization are self-contained and do not rest on those papers. No uniqueness theorem is imported from the authors. The experimental comparisons use external datasets and benchmarks, so the central claim has independent content.
Assumptions & free parameters
free parameters (3)
- β, σ_j², G_j² (loss smoothness and gradient statistics) =
estimated following the approach in [24]
- Learning rate γ =
5e-4 in simulations
- Client-side aggregation interval I =
15 in simulations
assumptions (6)
- domain assumption Each local loss function f_i is β-smooth.
- domain assumption Stochastic gradients have bounded per-layer variance σ_j²/b and bounded second moments Σ G_j².
- domain assumption Stochastic gradients are unbiased estimates of the true gradients.
- domain assumption Per-round training latency is the max over devices of compute+communication times plus the serial server time, and compute/communication times scale linearly with batch size.
- ad hoc to paper The convergence bound in Corollary 1 can be used as an equality to predict the required number of training rounds R.
- domain assumption Device memory constraint C4 is modeled as a sum of activation, gradient, optimizer state, and model parameter sizes.
Cite this review
Pith. "Pith review of HASFL: Heterogeneity-aware Split Federated Learning over Edge Computing Systems." pith.science (2026). https://pith.science/paper/NZZ4EQIC
@misc{pith2026250608426,
author = {Pith},
title = {Pith review of: HASFL: Heterogeneity-aware Split Federated Learning over Edge Computing Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/NZZ4EQIC}},
note = {Machine review of arXiv:2506.08426}
}
read the original abstract
Split federated learning (SFL) has emerged as a promising paradigm to democratize machine learning (ML) on edge devices by enabling layer-wise model partitioning. However, existing SFL approaches suffer significantly from the straggler effect due to the heterogeneous capabilities of edge devices. To address the fundamental challenge, we propose adaptively controlling batch sizes (BSs) and model splitting (MS) for edge devices to overcome resource heterogeneity. We first derive a tight convergence bound of SFL that quantifies the impact of varied BSs and MS on learning performance. Based on the convergence bound, we propose HASFL, a heterogeneity-aware SFL framework capable of adaptively controlling BS and MS to balance communication-computing latency and training convergence in heterogeneous edge networks. Extensive experiments with various datasets validate the effectiveness of HASFL and demonstrate its superiority over state-of-the-art benchmarks.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 3 Pith papers
-
RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
RRTO identifies static inference operator sequences from CUDA call logs alone and replays them on an edge GPU, cutting transparent-offloading communication to 11 RPCs per inference instead of thousands, with performan...
-
A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation
SpaceVerse jointly decides where to run vision-language inference in LEO satellite networks and compresses task-irrelevant image regions before downlink, improving accuracy and cutting latency versus baselines.
-
PHandover: Parallel Handover in Mobile Satellite Network
A parallel, plan-based handover using a new Satellite Synchronized Function cuts LEO satellite handover latency to about 9 ms on average in an emulated prototype.
Reference graph
Works this paper leans on
-
[24]
Adaptive Federated Learning in Resource Constrained Edge Computing Systems,
S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and K. Chan, “Adaptive Federated Learning in Resource Constrained Edge Computing Systems,” IEEE J. Sel. Areas Commun. , vol. 37, no. 6, pp. 1205–1221, Mar. 2019
work page 2019
-
[1]
Optimiz- ing Parameter Mixing Under Constrained Communications in Parallel Federated Learning,
X. Liu, Z. Yan, Y . Zhou, D. Wu, X. Chen, and J. H. Wang, “Optimiz- ing Parameter Mixing Under Constrained Communications in Parallel Federated Learning,” IEEE/ACM Trans. Networking, vol. 31, no. 6, pp. 2640–2652, Dec. 2023
work page 2023
-
[2]
Z. Lin, Y . Zhang, Z. Chen, Z. Fang, X. Chen, P. Vepakomma, W. Ni, J. Luo, and Y . Gao, “HSplitLoRA: A Heterogeneous Split Parameter- Efficient Fine-Tuning Framework for Large Language Models,” arXiv preprint arXiv:2505.02795, 2025
arXiv 2025
-
[3]
Actions at the Edge: Jointly Optimizing the Resources in Multi-access Edge Computing,
Y . Deng, X. Chen, G. Zhu, Y . Fang, Z. Chen, and X. Deng, “Actions at the Edge: Jointly Optimizing the Resources in Multi-access Edge Computing,” IEEE Wireless Commun., vol. 29, no. 2, pp. 192–198, Apr. 2022
work page 2022
-
[4]
Federated Learning over Multihop Wireless Networks with In-network Aggregation,
X. Chen, G. Zhu, Y . Deng, and Y . Fang, “Federated Learning over Multihop Wireless Networks with In-network Aggregation,”IEEE Trans. Wireless Commun., vol. 21, no. 6, pp. 4622–4634, Apr. 2022
work page 2022
-
[5]
Communication-efficient Learning of Deep Networks From Decentral- ized Data,
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. Arcas, “Communication-efficient Learning of Deep Networks From Decentral- ized Data,” in Proc. AISTATS, Apr. 2017
work page 2017
-
[6]
Federated Learning: Strategies for Improving Communica- tion Efficiency,
J. Kone ˇcn`y, H. B. McMahan, F. X. Yu, P. Richt ´arik, A. T. Suresh, and D. Bacon, “Federated Learning: Strategies for Improving Communica- tion Efficiency,” arXiv preprint arXiv:1610.05492 , Oct. 2016
arXiv 2016
-
[7]
FedSN: A Federated Learning Framework over Heterogeneous LEO Satellite Networks,
Z. Lin, Z. Chen, Z. Fang, X. Chen, X. Wang, and Y . Gao, “FedSN: A Federated Learning Framework over Heterogeneous LEO Satellite Networks,” IEEE Trans. Mobile Comput. , 2024
work page 2024
Show all 55 references
-
[8]
Accelerating Federated Learning with Model Segmentation for Edge Networks,
M. Hu, J. Zhang, X. Wang, S. Liu, and Z. Lin, “Accelerating Federated Learning with Model Segmentation for Edge Networks,” IEEE Trans. Green Commun. Netw., 2024
2024
-
[9]
Automated Federated Pipeline for Parameter-efficient Fine-tuning of Large Lan- guage Models,
Z. Fang, Z. Lin, Z. Chen, X. Chen, Y . Gao, and Y . Fang, “Automated Federated Pipeline for Parameter-efficient Fine-tuning of Large Lan- guage Models,” arXiv preprint arXiv:2404.06448 , 2024
2024 arXiv
-
[10]
Leo-Split: A Semi-Supervised Split Learning Framework over LEO Satellite Networks,
Z. Lin, Y . Zhang, Z. Chen, Z. Fang, C. Wu, X. Chen, Y . Gao, and J. Luo, “Leo-Split: A Semi-Supervised Split Learning Framework over LEO Satellite Networks,” arXiv preprint arXiv:2501.01293 , 2025
2025 arXiv
-
[11]
Gemini: A Family of Highly Capable Multimodal Models,
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican et al. , “Gemini: A Family of Highly Capable Multimodal Models,” arXiv preprint arXiv:2312.11805, Dec. 2023
2023 arXiv
-
[12]
Split Learning for Health: Distributed Deep Learning without Sharing Raw Patient Data,
P. Vepakomma, O. Gupta, T. Swedish, and R. Raskar, “Split Learning for Health: Distributed Deep Learning without Sharing Raw Patient Data,” arXiv preprint arXiv:1812.00564 , Dec. 2018
2018 arXiv
-
[13]
Pipelining Split Learning in Multi-hop Edge Networks,
W. Wei, Z. Lin, T. Li, X. Li, and X. Chen, “Pipelining Split Learning in Multi-hop Edge Networks,” arXiv preprint arXiv:2505.04368 , 2025
2025
-
[14]
Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,
Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, and K. Huang, “Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,” IEEE Commun. Mag. , 2023
2023
-
[15]
Optimal Resource Allocation for U-Shaped Parallel Split Learning,
S. Lyu, Z. Lin, G. Qu, X. Chen, X. Huang, and P. Li, “Optimal Resource Allocation for U-Shaped Parallel Split Learning,” in Proc. IEEE Globecom Wkshps , 2023, pp. 197–202
2023
-
[16]
Splitfed: When Federated Learning Meets Split Learning,
C. Thapa, P. C. M. Arachchige, S. Camtepe, and L. Sun, “Splitfed: When Federated Learning Meets Split Learning,” in Proc. AAAI, Feb. 2022
2022
-
[17]
Distributed Learning in Wireless Networks: Recent Progress and Future Challenges,
M. Chen, D. G ¨und¨uz, K. Huang, W. Saad, M. Bennis, A. V . Feljan, and H. V . Poor, “Distributed Learning in Wireless Networks: Recent Progress and Future Challenges,” IEEE J. Sel. Areas Commun. , Dec. 2021
2021
-
[18]
Time-sensitive Learning For Heterogeneous Federated Edge Intelligence,
Y . Xiao, X. Zhang, Y . Li, G. Shi, M. Krunz, D. N. Nguyen, and D. T. Hoang, “Time-sensitive Learning For Heterogeneous Federated Edge Intelligence,” IEEE Trans. Mobile Comput. , vol. 23, no. 2, pp. 1382– 1400, 2023
2023
-
[19]
Speeding Up Distributed Machine Learning Using Codes,
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran, “Speeding Up Distributed Machine Learning Using Codes,” IEEE Trans. Inf. Theory, vol. 64, no. 3, pp. 1514–1529, Mar. 2017
2017
-
[20]
Split learning over Wireless Networks: Parallel Design and Resource Management,
W. Wu, M. Li, K. Qu, C. Zhou, X. Shen, W. Zhuang, X. Li, and W. Shi, “Split learning over Wireless Networks: Parallel Design and Resource Management,” IEEE J. Sel. Areas Commun. , vol. 41, no. 4, pp. 1051– 1066, Feb. 2023
2023
-
[21]
AdaptsFL: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,
Z. Lin, G. Qu, W. Wei, X. Chen, and K. K. Leung, “AdaptsFL: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,” arXiv preprint arXiv:2403.13101 , 2024
2024 arXiv
-
[22]
Game-Theoretic Joint Incentive and Cut Layer Selection Mechanism in Split Federated Learning,
J. Lee, J. Cho, W. Lee, M. Seif, and H. V . Poor, “Game-Theoretic Joint Incentive and Cut Layer Selection Mechanism in Split Federated Learning,” arXiv preprint arXiv:2412.07813 , 2024
2024 arXiv
-
[23]
Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,
Z. Lin, W. Wei, Z. Chen, C.-T. Lam, X. Chen, Y . Gao, and J. Luo, “Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,” IEEE Trans. Mobile Comput. , Apr. 2025
2025
-
[25]
Adaptive Batchsize Selection and Gradient Compression For Wireless Federated Learning,
S. Liu, G. Yu, R. Yin, J. Yuan, and F. Qu, “Adaptive Batchsize Selection and Gradient Compression For Wireless Federated Learning,” in Proc. GLOBECOM, 2020
2020
-
[26]
Adaptive Batch Size For Federated Learning in Resource-constrained Edge Computing,
Z. Ma, Y . Xu, H. Xu, Z. Meng, L. Huang, and Y . Xue, “Adaptive Batch Size For Federated Learning in Resource-constrained Edge Computing,” IEEE Trans. Mobile Comput. , vol. 22, no. 1, pp. 37–53, Jan. 2021
2021
-
[27]
DYNAMITE: Dynamic Interplay of Mini-batch Size and Aggregation Frequency For Federated Learning with Static and Streaming Datasets,
W. Liu, X. Zhang, J. Duan, C. Joe-Wong, Z. Zhou, and X. Chen, “DYNAMITE: Dynamic Interplay of Mini-batch Size and Aggregation Frequency For Federated Learning with Static and Streaming Datasets,” IEEE Trans. Mobile Comput. , vol. 23, no. 7, pp. 7664–7679, 2023
2023
-
[28]
To Talk or to Work: Dynamic Batch Sizes Assisted Time Efficient Federated Learning over Future Mobile Edge Devices,
D. Shi, L. Li, M. Wu, M. Shu, R. Yu, M. Pan, and Z. Han, “To Talk or to Work: Dynamic Batch Sizes Assisted Time Efficient Federated Learning over Future Mobile Edge Devices,” IEEE Trans. Wireless Commun. , vol. 21, no. 12, pp. 11 038–11 050, Dec. 2022
2022
-
[29]
Don’t Decay the Learning Rate, Increase the Batch Size,
S. L. Smith, P.-J. Kindermans, and Q. V . Le, “Don’t Decay the Learning Rate, Increase the Batch Size,” in Proc. ICLR, Feb. 2018
2018
-
[30]
On the Computation and Communication Complexity of Parallel SGD with Dynamic Batch Sizes for Stochastic Non-convex Optimization,
H. Yu and R. Jin, “On the Computation and Communication Complexity of Parallel SGD with Dynamic Batch Sizes for Stochastic Non-convex Optimization,” in Proc. ICML, Jun. 2019
2019
-
[31]
Efficient Parallel Split Learning over Resource-constrained Wireless Edge Networks,
Z. Lin, G. Zhu, Y . Deng, X. Chen, Y . Gao, K. Huang, and Y . Fang, “Efficient Parallel Split Learning over Resource-constrained Wireless Edge Networks,” IEEE Trans. Mobile Comput. , 2024
2024
-
[32]
Convergence Analysis of Split Federated Learning on Heterogeneous Data,
P. Han, C. Huang, G. Tian, M. Tang, and X. Liu, “Convergence Analysis of Split Federated Learning on Heterogeneous Data,” in Proc. NIPS , 2025
2025
-
[33]
Unleashing the Tiger: Inference Attacks on Split Learning,
D. Pasquini, G. Ateniese, and M. Bernaschi, “Unleashing the Tiger: Inference Attacks on Split Learning,” in Proc. CCS, Nov. 2021
2021
-
[34]
On the Convergence Properties of A K-step Averaging Stochastic Gradient Descent Algorithm for Nonconvex Optimization,
Karimireddy, Sai Praneeth and Kale, Satyen and Mohri, Mehryar and Reddi, Sashank and Stich, Sebastian and Suresh, Ananda Theertha, “On the Convergence Properties of A K-step Averaging Stochastic Gradient Descent Algorithm for Nonconvex Optimization,” in Proc. IJCAI, Jul. 2018
2018
-
[35]
On the Linear Speedup Analysis of Communication Efficient Momentum SGD For Distributed Non-convex Optimization,
H. Yu, R. Jin, and S. Yang, “On the Linear Speedup Analysis of Communication Efficient Momentum SGD For Distributed Non-convex Optimization,” in Proc. ICML, Jun. 2019, pp. 7184–7193
2019
-
[36]
Scaffold: Stochastic Controlled Averaging for Federated Learn- ing,
S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh, “Scaffold: Stochastic Controlled Averaging for Federated Learn- ing,” in Proc. ICLR, Apr. 2020
2020
-
[37]
Communication-efficient Algorithms for Statistical Optimization,
Y . Zhang, M. J. Wainwright, and J. C. Duchi, “Communication-efficient Algorithms for Statistical Optimization,” in Proc. NIPS, Jun. 2012
2012
-
[38]
Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,” in Proc. NIPS, Jun. 2017
2017
-
[39]
Perturbed Iterate Analysis for Asynchronous Stochastic Optimization,
H. Mania, X. Pan, D. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan, “Perturbed Iterate Analysis for Asynchronous Stochastic Optimization,” SIAM J. Optim., vol. 27, no. 4, pp. 2202–2229, Jan. 2017
2017
-
[40]
Don’t Use Large Mini- batches, Use Local SGD,
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi, “Don’t Use Large Mini- batches, Use Local SGD,” in Proc. ICLR, Dec. 2019
2019
-
[41]
QSFL: Two-Level Communication-Efficient Federated Learning on Mobile Edge Devices,
L. Yi, G. Wang, X. Wang, and X. Liu, “QSFL: Two-Level Communication-Efficient Federated Learning on Mobile Edge Devices,” IEEE Trans. Serv. Comput. , 2024
2024
-
[42]
Joint Device Scheduling and Resource Allocation for Latency Constrained Wireless Federated Learning,
W. Shi, S. Zhou, Z. Niu, M. Jiang, and L. Geng, “Joint Device Scheduling and Resource Allocation for Latency Constrained Wireless Federated Learning,” IEEE Trans. Wireless Commun., vol. 20, no. 1, pp. 453–467, Sep. 2020
2020
-
[43]
Federated-learning-based Client Scheduling for Low-latency Wireless Communications,
W. Xia, W. Wen, K.-K. Wong, T. Q. Quek, J. Zhang, and H. Zhu, “Federated-learning-based Client Scheduling for Low-latency Wireless Communications,” IEEE Wireless Commun. , vol. 28, no. 2, pp. 32–38, Apr. 2021
2021
-
[44]
Horus: Interference-aware and Prediction-based Scheduling in Deep Learning Systems,
G. Yeung, D. Borowiec, R. Yang, A. Friday, R. Harper, and P. Garraghan, “Horus: Interference-aware and Prediction-based Scheduling in Deep Learning Systems,” IEEE Trans. Parallel Distrib. Syst. , vol. 33, no. 1, pp. 88–100, May. 2021
2021
-
[45]
Pipeline Network Simulation Calculation based on Improved Newton Jacobian Iterative Method,
Z. Geng, C. Wei, Y . Han, Q. Wei, and Z. Ouyang, “Pipeline Network Simulation Calculation based on Improved Newton Jacobian Iterative Method,” in Proc. AICS, Jul. 2019
2019
-
[46]
On Nonlinear Fractional Programming,
W. Dinkelbach, “On Nonlinear Fractional Programming,” Manage. Sci., vol. 13, no. 7, pp. 492–498, Mar. 1967. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 16
1967
-
[47]
A Reformulation-linearization Method for the Global Optimization of Large-scale Mixed-Integer Linear Fractional Programming Problems and Cyclic Scheduling aApplication,
D. Yue and F. You, “A Reformulation-linearization Method for the Global Optimization of Large-scale Mixed-Integer Linear Fractional Programming Problems and Cyclic Scheduling aApplication,” in Proc. ACC, Jun. 2013
2013
-
[48]
Extensions of Dinkelbach’s Algorithm for Solving Non-linear Fractional Programming Problems,
R. G. R ´odenas, M. L. L ´opez, and D. Verastegui, “Extensions of Dinkelbach’s Algorithm for Solving Non-linear Fractional Programming Problems,” Top, vol. 7, pp. 33–70, Jun. 1999
1999
-
[49]
Convergence of A Block Coordinate Descent Method for Nondifferentiable Minimization,
P. Tseng, “Convergence of A Block Coordinate Descent Method for Nondifferentiable Minimization,” J Optim Theory Appl , vol. 109, pp. 475–494, Jun. 2001
2001
-
[50]
Learning Multiple Layers of Features From Tiny Images,
A. Krizhevsky, G. Hinton et al., “Learning Multiple Layers of Features From Tiny Images,” Tech. Rep., Apr. 2009
2009
-
[51]
Broadband analog aggregation for low-latency federated edge learning,
G. Zhu, Y . Wang, and K. Huang, “Broadband analog aggregation for low-latency federated edge learning,” IEEE Trans. Wireless Commun. , vol. 19, no. 1, pp. 491–506, Oct. 2019
2019
-
[52]
Energy Efficient Federated Learning over Wireless Communication Networks,
Z. Yang, M. Chen, W. Saad, C. S. Hong, and M. Shikh-Bahaei, “Energy Efficient Federated Learning over Wireless Communication Networks,” IEEE Trans. Wireless Commun. , vol. 20, no. 3, pp. 1935–1949, Nov. 2020
1935
-
[53]
Very Deep Convolutional Networks for Large-scale Image Recognition,
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-scale Image Recognition,” in Proc. ICLR, 2015
2015
-
[54]
Deep Residual Learning for Image Recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proc. CVPR, Jun. 2016, pp. 770–778
2016
-
[55]
CoopFL: Accelerating Federated Learning with DNN Partitioning and Offloading in Hetero- geneous Edge Computing,
Z. Wang, H. Xu, Y . Xu, Z. Jiang, and J. Liu, “CoopFL: Accelerating Federated Learning with DNN Partitioning and Offloading in Hetero- geneous Edge Computing,” Comput. Netw., vol. 220, p. 109490, Jan. 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.