REVIEW 4 major objections 6 minor 60 references
FedGAT: A Privacy-Preserving Federated Approximation Algorithm for Graph Attention Networks
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read FedGAT claims that the attention scores of a Graph Attention Network can be replaced by a precomputable Chebyshev polynomial expansion, so federated GAT training needs only one pre-training communication round.
desk verdict Clever single-layer trick, but the one-round claim collapses for the multi-layer GATs actually used in the experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is a truncated Chebyshev series for the attention score. The paper defines $x_{ij} = b_1^T h_i + b_2^T h_j$ and expands $\exp(\psi(x_{ij}))$ as $\sum_{n=0}^p q_n x_{ij}^n$. It then builds, for each node $i$, idempotent matrices $U_j$ from a set of orthonormal vectors, so that the weighted sum $D_i = \sum_{j \in N_i} x_{ij} U_j$ obeys $D_i^n = \sum_{j \in N_i} x_{ij}^n U_j$. With auxiliary vectors $K_{1i}$ and $K_{2i}$, this turns the graph-neighbourhood sums $E_i^{(n)}$ and $F_i^{(n)}$ into matrix products $K^T D^n K$, and the matrices inside $D_i$ are separated into parameter-independent pieces $M_{1i}(s)$ and $M_{2i}(s)$ that can be precomputed and transmitted once.
What would settle it
On a small graph, run the pre-training round and train a single-layer FedGAT, comparing the approximate attention coefficients to the centralized GAT. If the maximum error does not shrink as the Chebyshev degree $p$ increases, the error-bound theorem fails. Separately, try training a 3-layer FedGAT while enforcing strict client isolation, so no client ever receives another client's post-first-layer embeddings; if the updates cannot be computed or accuracy drops to the edge-dropping baseline, the one-round claim does not hold for more than one layer.
Extended reading notes
Core claim
The central claim is that a GAT update can be approximated in a federated setting without dropping cross-client edges and without per-round feature exchange. The paper writes the attention score $e_{ij} = \exp(\psi(x_{ij}))$ as a truncated power series in $x_{ij}$, and then shows the sums $E_i^{(n)}$ and $F_i^{(n)}$ needed for the update can be recovered from four quantities ($M_{1i}(s)$, $M_{2i}(s)$, $K_{1i}$, $K_{2i}$) that depend only on fixed input features and a private orthonormal construction. Because those quantities do not depend on the learnable parameters, they are shared exactly once before training; after that, clients compute approximate GAT updates locally and only exchange model parameters. Theorems bound the error in attention coefficients and embeddings, showing the error shrinks with polynomial degree and propagates across layers at a rate controlled by the Lipschitz constants of the activations.
Load-bearing premise
For more than one GAT layer, the method assumes that embeddings produced after the first layer can be viewed by nodes on other clients, and the paper gives no protocol or communication cost for sharing those changing embeddings.
Editorial extensions
If this is right
- Federated GAT training becomes communication-feasible: feature-derived information is exchanged in a single pre-training round, and all later rounds exchange only model parameters.
- Cross-client edges can be retained, so accuracy does not degrade with the number of clients or with non-iid label distributions the way edge-dropping baselines do.
- The approximation error is tunable: increasing the Chebyshev degree decreases attention-score error, and the propagation bounds say the final embedding error stays controlled for the shallow GATs used in practice.
- Privacy is preserved in aggregate form: the shared objects reveal only neighbourhood sums, and the algorithm drops a cross-client neighbor when it would be the only one, avoiding direct feature recovery.
- Computational and communication costs grow with the maximum node degree ($O(K B^L d B^2)$ communication), making the method suited to sparse graphs; an appendix variant lowers the per-node cost at the price of weaker privacy for special features.
Reading between the lines
- The single-round trick likely extends beyond GATs to any attention layer whose score is exp of a bilinear form over fixed input features, so the same precomputation idea could apply to first layers of transformers in federated settings.
- For more than one layer, the paper's assumption that post-first-layer embeddings are freely visible is a second implicit communication round (or a trusted shared memory); a complete protocol would need to specify how those embeddings are exchanged before the one-round claim covers multilayer GATs.
- The aggregate-only privacy guarantee is heuristic, not cryptographic; pairing the pre-training exchange with secure aggregation or homomorphic encryption, which the paper names as a future direction, would convert it into a computational guarantee.
- Because FedGAT keeps all cross-client edges, its robustness to data heterogeneity is better explained by edge retention than by the federated averaging scheme; an ablation that randomly removes cross-client edges could separate the two effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FedGAT, a federated training algorithm for Graph Attention Networks (GATs) on graphs partitioned across clients. The method approximates the attention scores exp(ψ(x_ij)) by a truncated Chebyshev series, re-expresses the result as a power series in x_ij, and pre-computes parameter-independent aggregate matrices (M1, M2, K1, K2) during a single pre-training communication round. These aggregates then allow clients to compute approximate GAT updates without exchanging node features during training. The paper provides a communication overhead analysis, approximation error bounds across layers, a heuristic privacy analysis, and experiments on Cora, Citeseer, and Pubmed showing accuracy close to centralized GAT and above FedGCN and DistGAT.
Significance. The single-layer construction is original and addresses a real obstacle in federated GAT training: attention scores depend on features at both endpoints of cross-client edges and change every training round. If the one-communication-round claim were fully supported, the paper would be a meaningful advance, since existing methods either drop cross-client edges or incur per-round communication. The paper includes hand-written proofs for the approximation error and a communication complexity theorem, and the empirical results are promising. However, the multi-layer extension is not compatible with the one-round claim as written, and the Chebyshev error bounds are applied outside their stated domain. These issues concern the paper's central contributions and need to be resolved before the claims can be accepted.
major comments (4)
- [Section 4, 'FedGAT for Multiple GAT Layers'; Algorithm 2; Theorem 1] The one-round communication claim is unsupported for multi-layer GATs. The paper assumes 'nodes are permitted to view the embeddings of any other node generated after the first GAT layer, including nodes on other clients,' but for L>1 the regular GAT update (Eqs. 1-3) used in Algorithm 2 requires, for each cross-client edge (i,j), the current embedding h_j^{(l-1)}, which is a function of the current global parameters and changes every training round. No protocol for sharing these embeddings is specified, and Theorem 1 counts only the pre-training communication. Since Appendix C evaluates a 2-layer GAT, the experiments exercise exactly the setting in which the one-round claim is not established.
- [Section 4, Eq. (5); Section 5, Theorem 2] The Chebyshev approximation and its error bound are stated for a function on [-1,1], but the argument x_ij = b_1^T h_i + b_2^T h_j is a learned, potentially unbounded scalar that depends on the current parameters and features. The paper does not restrict x_ij to [-1,1], does not rescale or clamp it, and does not re-derive the Chebyshev coefficients or error bounds for the actual range of x_ij. Consequently, Theorems 2-5 do not apply to the algorithm as implemented, and the 'provable bounds' claim is not supported.
- [Section 5, 'Privacy Analysis of FedGAT'] The privacy argument is informal and does not match the strong 'privacy-preserving' claim in the title and abstract. The analysis shows that certain products of K1_i and K2_i recover only aggregate sums of features, but it does not define an adversary model, does not account for auxiliary information or repeated queries, and does not quantify the leakage through the M2_i(s) matrices when the random U_j masks are unknown to the observer. Section 7 correctly identifies 'theoretical privacy guarantees' as future work, which conflicts with the unqualified privacy claim made earlier in the paper.
- [Section 5, Theorems 3-5; Appendix E] The error propagation results assume a Lipschitz, monotone activation ψ (Assumption 4), but Theorem 2 requires f to be k-times differentiable with f^{(k)} of bounded variation. Standard GAT activations such as LeakyReLU are not differentiable at zero, so the composition exp(ψ(x)) does not satisfy the differentiability condition needed for the stated Chebyshev convergence rate. The paper should either state which activation functions the error bounds actually cover or derive bounds that do not rely on higher-order differentiability.
minor comments (6)
- [Appendix A, Algorithm 1] In the definition of M2_i(s), the symbol 'A_j' appears instead of the earlier-defined 'U_j'; this makes the algorithm ambiguous.
- [Section 6, 'Methods Compared'] The text says 'FedCGN (Yao et al. 2023a)' but the correct name is FedGCN; please correct this typo.
- [Section 4, after Eq. (14)] The expression 'b2 + hj(s)' should presumably read 'b2(s) + hj(s)' or 'b2(s)hj(s)'; as written, it is unclear.
- [Appendix E, proof of Claim 2] In the chain of inequalities, the term 'ˆα_ij' appears inside a sum over k; this should likely be 'ˆα_ik' or similar, and the derivation should be cleaned up.
- [Section 5, 'Communication Overhead'] The symbol B is used for both the maximum node degree in Section 5 and the background appendix title; this is a minor notational collision that could confuse readers.
- [Conclusion] The conclusion states FedGAT is 'one of the first' algorithms while the introduction says 'to the best of our knowledge, this is the first work'; please make these claims consistent.
Circularity Check
No circular derivation: the Chebyshev approximation and error bounds are externally grounded; the multi-layer one-round gap is a correctness limitation, not a circular reduction.
full rationale
The paper's central claim is that FedGAT approximates the GAT attention scores e_ij = exp(psi(x_ij)) by a truncated Chebyshev series and pre-communicates feature-and-graph quantities that are independent of the learnable parameters. This is not circular: the coefficients q_n come from the known activation function via Chebyshev expansion, and the error bounds (Theorems 2-5) are derived from Trefethen's external approximation theorem plus Lipschitz assumptions, not from the accuracy values the paper predicts. The precomputed matrices M1i(s), M2i(s), K1i, K2i are functions of node features and graph structure only; they do not encode the centralized GAT's accuracy, so comparing FedGAT to a centralized GAT is an external benchmark rather than an identity. The only self-citation is FedGCN (Yao et al. 2023a), which the paper cites as inspiration and uses as a baseline; FedGAT's polynomial-expansion mechanism is not imported from FedGCN, so this citation is not load-bearing. A genuine limitation appears in Section 4: for multiple GAT layers, FedGAT assumes 'nodes are permitted to view the embeddings of any other node generated after the first GAT layer, including nodes on other clients,' and the paper later concedes it is 'not optimized to handle GATs with several layers (more than 2).' Since layer-1 embeddings change with every training round, this assumption would require additional per-round communication for the 2-layer GATs used in the experiments. That is a soundness/completeness gap in the one-round claim, but it is not a circular reduction of any equation to its input. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is merely relabeled. I therefore find no significant circularity.
Assumptions & free parameters
free parameters (2)
- Chebyshev approximation degree p =
16 in all reported experiments (degree sensitivity 8-16 in Figure 5)
- Chebyshev approximation interval =
not specified in paper
assumptions (4)
- domain assumption The attention score function exp(psi(x)) is representable on the relevant range of x_ij by a Chebyshev series whose truncated error obeys Theorem 2.
- domain assumption For L>1, nodes may view the first-layer embeddings of any other node, including nodes on other clients, without additional communication cost.
- domain assumption Activation functions are Lipschitz continuous and monotone; parameters and features have bounded norms (Assumptions 2-4).
- standard math Existence of orthonormal vectors u1j, u2j in R^{2deg(i)} with the idempotence and orthogonality properties of U_j.
Cite this review
Pith. "Pith review of FedGAT: A Privacy-Preserving Federated Approximation Algorithm for Graph Attention Networks." pith.science (2026). https://pith.science/paper/MHEHUHFO
@misc{pith2026241216144,
author = {Pith},
title = {Pith review of: FedGAT: A Privacy-Preserving Federated Approximation Algorithm for Graph Attention Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/MHEHUHFO}},
note = {Machine review of arXiv:2412.16144}
}
read the original abstract
Federated training methods have gained popularity for graph learning with applications including friendship graphs of social media sites and customer-merchant interaction graphs of huge online marketplaces. However, privacy regulations often require locally generated data to be stored on local clients. The graph is then naturally partitioned across clients, with no client permitted access to information stored on another. Cross-client edges arise naturally in such cases and present an interesting challenge to federated training methods, as training a graph model at one client requires feature information of nodes on the other end of cross-client edges. Attempting to retain such edges often incurs significant communication overhead, and dropping them altogether reduces model performance. In simpler models such as Graph Convolutional Networks, this can be fixed by communicating a limited amount of feature information across clients before training, but GATs (Graph Attention Networks) require additional information that cannot be pre-communicated, as it changes from training round to round. We introduce the Federated Graph Attention Network (FedGAT) algorithm for semi-supervised node classification, which approximates the behavior of GATs with provable bounds on the approximation error. FedGAT requires only one pre-training communication round, significantly reducing the communication overhead for federated GAT training. We then analyze the error in the approximation and examine the communication overhead and computational complexity of the algorithm. Experiments show that FedGAT achieves nearly the same accuracy as a GAT model in a centralised setting, and its performance is robust to the number of clients as well as data distribution.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Albrecht, J. P. 2016. How the GDPR will change the world. Eur. Data Prot. L. Rev., 2: 287
work page 2016
-
[4]
B.; Patel, S.; Ramage, D.; Segal, A.; and Seth, K
Bonawitz, K.; Ivanov, V.; Kreuter, B.; Marcedone, A.; McMahan, H. B.; Patel, S.; Ramage, D.; Segal, A.; and Seth, K. 2017. Practical secure aggregation for privacy-preserving machine learning. In proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 1175--1191
work page 2017
-
[5]
Boyd, S.; Parikh, N.; and Chu, E. 2011. Distributed Optimization and Statistical Learning Via the Alternating Direction Method of Multipliers. Foundations and trends in machine learning. Now Publishers. ISBN 9781601984609
work page 2011
-
[6]
Brody, S.; Alon, U.; and Yahav, E. 2022. How Attentive are Graph Attention Networks? arXiv:2105.14491
arXiv 2022
-
[7]
M.; Bruna, J.; Cohen, T.; and Veličković, P
Bronstein, M. M.; Bruna, J.; Cohen, T.; and Veličković, P. 2021. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. arXiv:2104.13478
arXiv 2021
-
[8]
M.; Bruna, J.; LeCun, Y.; Szlam, A.; and Vandergheynst, P
Bronstein, M. M.; Bruna, J.; LeCun, Y.; Szlam, A.; and Vandergheynst, P. 2017. Geometric Deep Learning: Going beyond Euclidean data. IEEE Signal Processing Magazine, 34(4): 18–42
work page 2017
Show all 60 references
-
[9]
Chen, C.; Hu, W.; Xu, Z.; and Zheng, Z. 2021 a . FedGL: Federated Graph Learning Framework with Global Self-Supervision. arXiv:2105.03170
2021 arXiv
-
[10]
Chen, F.; Li, P.; Miyazaki, T.; and Wu, C. 2021 b . FedGraph: Federated Graph Learning with Intelligent Sampling. arXiv:2111.01370
2021 arXiv
-
[11]
Chen, G.; and Liu, Z.-P. 2022. Graph attention network for link prediction of gene regulations from single-cell RNA-sequencing data . Bioinformatics, 38(19): 4522--4529
2022
-
[12]
Clement, P. R. 1953. THE CHEBYSHEV APPROXIMATION METHOD. Quarterly of Applied Mathematics, 11(2): 167--183
1953
-
[13]
M.; and Mahdavi, M
Deng, Y.; Kamani, M. M.; and Mahdavi, M. 2020. Adaptive personalized federated learning. arXiv preprint arXiv:2003.13461
2020 arXiv
-
[14]
Descloux, J. 1963. Approximations in L\^ P and Chebyshev approximations. Journal of the Society for Industrial and Applied Mathematics, 11(4): 1017--1026
1963
-
[15]
Drusvyatskiy, D. 2017. The proximal point method revisited. arXiv:1712.06038
2017 arXiv
-
[16]
M.; and Sidford, A
Frostig, R.; Ge, R.; Kakade, S. M.; and Sidford, A. 2015. Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization. arXiv:1506.07512
2015 arXiv
-
[17]
Gabay, D.; and Mercier, B. 1976. A dual algorithm for the solution of nonlinear variational problems via finite element approximation. Computers & Mathematics with Applications, 2(1): 17--40
1976
-
[18]
Gandhi, S.; and Iyer, A. P. 2021. P3: Distributed Deep Graph Learning at Scale. In 15th USENIX Symposium on Operating Systems Design and Implementation ( OSDI 21) , 551--568. USENIX Association. ISBN 978-1-939133-22-9
2021
-
[19]
Glowinski, R.; and Marroco, A. 1975. Sur l'approximation, par éléments finis d'ordre un, et la résolution, par pénalisation-dualité d'une classe de problèmes de Dirichlet non linéaires. ESAIM: Mathematical Modelling and Numerical Analysis - Modélisation Mathématique et Analyse...
1975
-
[20]
L.; Ying, R.; and Leskovec, J
Hamilton, W. L.; Ying, R.; and Leskovec, J. 2018. Inductive Representation Learning on Large Graphs. arXiv:1706.02216
2018 arXiv
-
[21]
S.; Rong, Y.; Zhao, P.; Huang, J.; Annavaram, M.; and Avestimehr, S
He, C.; Balasubramanian, K.; Ceyani, E.; Yang, C.; Xie, H.; Sun, L.; He, L.; Yang, L.; Yu, P. S.; Rong, Y.; Zhao, P.; Huang, J.; Annavaram, M.; and Avestimehr, S. 2021. FedGraphNN: A Federated Learning System and Benchmark for Graph Neural Networks. arXiv:2104.07145
2021 arXiv
-
[22]
Henaff, M.; Bruna, J.; and LeCun, Y. 2015. Deep Convolutional Networks on Graph-Structured Data. arXiv:1506.05163
2015 arXiv
-
[23]
Hernández, M. 2001. Chebyshev's approximation algorithms and applications. Computers & Mathematics with Applications, 41(3): 433--445
2001
-
[24]
H.; Qi, H.; and Brown, M
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019. Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification. arXiv:1909.06335
2019 arXiv
-
[25]
Hu, W.; Fey, M.; Zitnik, M.; Dong, Y.; Ren, H.; Liu, B.; Catasta, M.; and Leskovec, J. 2021. Open Graph Benchmark: Datasets for Machine Learning on Graphs. arXiv:2005.00687
2021 arXiv
-
[26]
Hu, Z.; Dong, Y.; Wang, K.; and Sun, Y. 2020. Heterogeneous Graph Transformer. arXiv:2003.01332
2020 arXiv
-
[27]
Kafash, B.; Delavarkhalafi, A.; and Karbassi, S. 2012. Application of Chebyshev polynomials to derive efficient algorithms for the solution of optimal control problems. Scientia Iranica, 19(3): 795--805
2012
-
[28]
P.; and Ba, J
Kingma, D. P.; and Ba, J. 2017. Adam: A Method for Stochastic Optimization. arXiv:1412.6980
2017 arXiv
-
[29]
N.; and Welling, M
Kipf, T. N.; and Welling, M. 2017. Semi-Supervised Classification with Graph Convolutional Networks. arXiv:1609.02907
2017 arXiv
-
[30]
Kosaraju, V.; Sadeghian, A.; Mart\' n-Mart\' n, R.; Reid, I.; Rezatofighi, H.; and Savarese, S. 2019. Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks. In Wallach, H.; Larochelle, H.; Beygelzimer, A.; d Alch\' e -Buc, F.; Fox, E.; ...
2019
-
[31]
Li, G.; Müller, M.; Thabet, A.; and Ghanem, B. 2019. DeepGCNs: Can GCNs Go as Deep as CNNs? arXiv:1904.03751
2019 arXiv
-
[32]
K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V
Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020 a . Federated Optimization in Heterogeneous Networks. arXiv:1812.06127
2020 arXiv
-
[33]
Li, X.; Huang, K.; Yang, W.; Wang, S.; and Zhang, Z. 2020 b . On the Convergence of FedAvg on Non-IID Data. arXiv:1907.02189
2020 arXiv
-
[34]
Luo, Z.-Q.; and Tseng, P. 1993. On the Convergence Rate of Dual Ascent Methods for Linearly Constrained Convex Minimization. Mathematics of Operations Research, 18(4): 846--867
1993
-
[35]
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, 1273--1282. PMLR
2017
-
[36]
B.; Moore, E.; Ramage, D.; and y Arcas, B
McMahan, H. B.; Moore, E.; Ramage, D.; and y Arcas, B. A. 2016. Federated learning of deep networks using model averaging. arXiv preprint arXiv:1602.05629, 2(2)
2016 arXiv
-
[37]
Mendieta, M.; Yang, T.; Wang, P.; Lee, M.; Ding, Z.; and Chen, C. 2022. Local Learning Matters: Rethinking Data Heterogeneity in Federated Learning. arXiv:2111.14213
2022 arXiv
-
[38]
Reddi, S.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush, K.; Kone c n \`y , J.; Kumar, S.; and McMahan, H. B. 2020. Adaptive federated optimization. arXiv preprint arXiv:2003.00295
2020 arXiv
-
[39]
Rivlin, T. 2020. Chebyshev Polynomials. Dover Books on Mathematics. Dover Publications. ISBN 9780486842332
2020
-
[40]
Scardapane, S.; Spinelli, I.; and Lorenzo, P. D. 2021. Distributed Training of Graph Convolutional Networks. IEEE Transactions on Signal and Information Processing over Networks, 7: 87–100
2021
-
[41]
L.; and Bach, F
Schmidt, M.; Roux, N. L.; and Bach, F. 2011. Convergence Rates of Inexact Proximal-Gradient Methods for Convex Optimization. arXiv:1109.2415
2011 arXiv
-
[42]
Shao, Y.; Li, H.; Gu, X.; Yin, H.; Li, Y.; Miao, X.; Zhang, W.; Cui, B.; and Chen, L. 2023. Distributed Graph Neural Network Training: A Survey. arXiv:2211.00216
2023 arXiv
-
[43]
Song, W.; Xiao, Z.; Wang, Y.; Charlin, L.; Zhang, M.; and Tang, J. 2019. Session-Based Social Recommendation via Dynamic Graph Attention Networks. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, WSDM '19, 555–563. New York, NY, USA: As...
2019
-
[44]
L.; Watanabe, Y.; Loyola, P.; Klyashtorny, D.; Ludwig, H.; and Bhaskaran, K
Suzumura, T.; Zhou, Y.; Baracaldo, N.; Ye, G.; Houck, K.; Kawahara, R.; Anwar, A.; Stavarache, L. L.; Watanabe, Y.; Loyola, P.; Klyashtorny, D.; Ludwig, H.; and Bhaskaran, K. 2019. Towards Federated Graph Learning for Collaborative Financial Crimes Detection. arXiv:1909.12946
2019 arXiv
-
[45]
Trefethen, L. N. 1981. Rational Chebyshev approximation on the unit disk. Numerische Mathematik, 37: 297--320
1981
-
[46]
Trefethen, L. N. 2019. Approximation Theory and Approximation Practice, Extended Edition. Philadelphia, PA: Society for Industrial and Applied Mathematics
2019
-
[47]
Tseng, P. 1990. Dual Ascent Methods for Problems with Strictly Convex Costs and Linear Constraints: A Unified Approach. SIAM Journal on Control and Optimization, 28(1): 214--242
1990
-
[48]
N.; Kaiser, L
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017. Attention is All you Need. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Infor...
2017
-
[49]
Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; and Bengio, Y. 2018. Graph Attention Networks. arXiv:1710.10903
2018 arXiv
-
[50]
S.; and Lin, Y
Wan, C.; Li, Y.; Li, A.; Kim, N. S.; and Lin, Y. 2022. BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node Sampling. arXiv:2203.10983
2022 arXiv
-
[51]
Wang, X.; He, X.; Cao, Y.; Liu, M.; and Chua, T.-S. 2019. KGAT: Knowledge Graph Attention Network for Recommendation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD '19, 950–958. New York, NY, USA: Association for Compu...
2019
-
[52]
Yao, Y.; Jin, W.; Ravi, S.; and Joe-Wong, C. 2023 a . FedGCN: Convergence-Communication Tradeoffs in Federated Training of Graph Convolutional Networks. arXiv:2201.12433
2023 arXiv
-
[53]
M.; Cheng, Z.; Chen, L.; Joe-Wong, C.; and Liu, T
Yao, Y.; Kamani, M. M.; Cheng, Z.; Chen, L.; Joe-Wong, C.; and Liu, T. 2023 b . FedRule: Federated Rule Recommendation System with Graph Neural Networks. In Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation, IoTDI '23, 197–208. New York...
2023
-
[54]
Zhang, C.; Xie, Y.; Bai, H.; Yu, B.; Li, W.; and Gao, Y. 2021 a . A survey on federated learning. Knowledge-Based Systems, 216: 106775
2021
-
[55]
Zhang, J.; Wu, Q.; Zhang, J.; Shen, C.; and Lu, J. 2019. Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention Networks. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2951--2959. Los Alamitos, CA, USA: IEEE Compu...
2019
-
[56]
Zhang, K.; Yang, C.; Li, X.; Sun, L.; and Yiu, S. M. 2021 b . Subgraph Federated Learning with Missing Neighbor Generation. arXiv:2106.13430
2021 arXiv
-
[57]
Zheng, C.; Fan, X.; Wang, C.; and Qi, J. 2019. GMAN: A Graph Multi-Attention Network for Traffic Prediction. arXiv:1911.08415
2019 arXiv
-
[58]
Zheng, D.; Ma, C.; Wang, M.; Zhou, J.; Su, Q.; Song, X.; Gan, Q.; Zhang, Z.; and Karypis, G. 2020 a . DistDGL: Distributed graph neural network training for billion-scale graphs. In IA\^3 2020 10th Workshop
2020
-
[59]
Zheng, D.; Ma, C.; Wang, M.; Zhou, J.; Su, Q.; Song, X.; Gan, Q.; Zhang, Z.; and Karypis, G. 2020 b . DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. In 2020 IEEE/ACM 10th Workshop on Irregular Applications: Architectures and Algorithms (IA3), 36--44
2020
-
[60]
Zheng, L.; Zhou, J.; Chen, C.; Wu, B.; Wang, L.; and Zhang, B. 2020 c . ASFGNN: Automated Separated-Federated Graph Neural Network. arXiv:2011.03248
2020 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.