REVIEW 3 major objections 6 minor 127 references
GRAMA: Adaptive Graph Autoregressive Moving Average Models
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read GRAMA turns static graph features into a learnable sequence, making any GNN backbone equivalent to a state-space model while preserving permutation equivariance.
desk verdict Useful equivariant ARMA-on-sequences architecture with strong empirical gains, but the paper's theoretical foundation for oversquashing is asserted rather than proven, and the stability theorems do not match the deployed normalization. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the GRAMA block: an ARMA($p,q$) recurrence $f^{(\ell+L)}=\sum_{i=1}^{p}\phi_i f^{(\ell+L-i)}+\sum_{j=1}^{q}\theta_j\delta^{(\ell+L-j)}+\delta^{(\ell+L)}$, where the residual $\delta^{(\ell+L)}$ is produced by a permutation-equivariant GNN backbone applied without nonlinearity. The coefficients $\{\phi_i\}$ and $\{\theta_j\}$ are not free parameters but are predicted per-input by a softmax-free multi-head attention layer over the pooled state and residual sequences, giving a selective mechanism analogous to Mamba. The block's key property is its equivalent state matrix $A$ in the companion form of Theorem 4.1; the paper shows that controlling the roots of $P(\lambda)$ controls both stability and how many sequence steps information can propagate.
What would settle it
Train GRAMA on a long-range task (e.g., the feature-transfer benchmark) and then at inference time randomly permute the order of the $L$ embedded copies that form the input sequence. If accuracy is unchanged by the permutation, the sequential ARMA structure is not the mechanism driving the gain. Alternatively, replace the $L$ learned embeddings with $L$ identical copies of the same feature vector: if GRAMA retains its improvement, the gain comes from depth and residual connections rather than from the ARMA recurrence.
Extended reading notes
Core claim
The central discovery is that the static-to-sequential transformation defined in Equation (2)—stacking the same node features through L different embeddings—lets a standard ARMA recurrence act as a drop-in module for any GNN backbone, and that this module is exactly a linear SSM on the constructed sequence. Theorem 4.1 establishes the ARMA–SSM equivalence in the graph setting; Lemma 4.2 and Theorem 4.3 tie stability to the spectral radius of the companion state matrix; Theorem 4.4 states that placing the roots of $P(\lambda)=\lambda^p-\sum_{j=1}^p\phi_j\lambda^{p-j}$ close to the unit circle extends the propagation range. The paper's claim is that this controlled, learnable recurrence mitigates oversquashing and improves long-range interaction modeling, with experiments on feature transfer, graph property prediction, the Long Range Graph Benchmark, MalNet-Tiny, and heterophilic node classification supporting the claim.
Load-bearing premise
The method assumes that stacking differently embedded copies of the same static node features into a sequence creates a meaningful temporal structure for the ARMA model to exploit; the theoretical analysis characterizes stability and decay along that synthetic sequence but never proves that this controls oversquashing caused by graph bottlenecks.
Editorial extensions
If this is right
- GRAMA can be attached to any permutation-equivariant GNN backbone—MPNN or graph transformer—and the resulting model is equivalent to a stack of graph-informed SSMs.
- By initializing the autoregressive coefficients to keep the roots of $P(\lambda)$ near the unit circle, practitioners get a principled way to extend the propagation range of a graph network.
- The selective attention mechanism makes the ARMA coefficients input-dependent, so the model can choose different dynamics for different graphs or different feature states.
- Because the sequence length $L$ and block count $S$ control effective depth, GRAMA depth can be scaled without the performance collapse seen in deep MPNNs.
Reading between the lines
- The 'time' axis of GRAMA is an artifact of construction: the L sequence steps are differently embedded copies of the same static features, so the ARMA dynamics act on a dimension the model itself creates. The paper's theory characterizes decay along that constructed dimension, not a proven bound on oversquashing caused by graph bottlenecks; the oversquashing benefit rests on the empirical transfer
- A direct test of the mechanism would shuffle the order of the L embeddings at inference: if GRAMA's gains are truly from the sequential ARMA structure, permuting the sequence should degrade performance; if the gains are from the enriched feature set and residual depth, the order should matter little.
- Because the model already treats graphs as sequences of feature states, GRAMA is naturally positioned for spatio-temporal graph data; the authors report preliminary positive results on three such benchmarks, a direction they flag as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GRAMA transforms a static graph into a sequence by stacking L different MLP embeddings of the same node features (Eq. (2)), then applies learnable AR and MA recurrences whose innovation term is supplied by a GNN backbone (Eq. (6)). The ARMA coefficients are generated by a graph-adaptive attention mechanism (Eqs. (19)–(24)). The paper claims an equivalence between GRAMA and graph SSMs, stability and long-range propagation theorems (Section 4), and reports experiments on 14 datasets using GCN, GatedGCN, and GPS backbones, with consistent gains over the backbones and competitive state-of-the-art performance.
Significance. If the empirical results are taken at face value, GRAMA is a useful drop-in module for equivariant GNN backbones: it preserves permutation equivariance, is relatively cheap (Appendix E), and shows consistent improvements across long-range and heterophilic benchmarks. The paper also includes ablations on selective vs. naive ARMA coefficients and on depth/width, which are helpful. The claimed theoretical foundation, however, is weaker than advertised: Theorems 4.1–4.4 concern the linear recurrence along the constructed sequence index and do not establish a graph-level oversquashing result, and the deployed coefficient normalization violates the strict stability hypothesis of Theorem 4.4. These gaps are fixable but require reworking the theory or reframing the claims.
major comments (3)
- [Section 4, Eq. (6), Lemma B.1] The theoretical results analyze a linear recurrence along the artificially constructed sequence index, not along the graph. Lemma B.1 and Theorem 4.4 bound the decay of powers of the state matrix of that recurrence; they contain no graph ingredient beyond the generic update δ^(ℓ+L) = GNN(f^(ℓ+L−1); G). Consequently, no statement is proved about how information crosses graph bottlenecks or about oversquashing as defined by Alon & Yahav (2021) and Di Giovanni et al. (2023). The Section 1 bullet claiming a ‘theoretical foundation’ for addressing oversquashing therefore overstates what Section 4 establishes; either a direct graph-level bound (e.g., in terms of effective resistance, cuts, or Jacobian sensitivity along graph paths) must be added, or the claim should be reframed as long-range propagation along the GRAMA sequence.
- [Section 3.2, Eq. (23), and Section 4, Theorems 4.3–4.4] The normalization in Eq. (23) enforces Σ_j cAR_j = 1. Since in all experiments p = q = R = L (Section 3.1 and Appendix D.6), the characteristic polynomial P(λ)=λ^p − Σ_j cAR_j λ^(p−j) satisfies P(1)=0. Theorem 4.4 assumes all roots of P are strictly inside the unit circle, so the deployed parameterization violates the theorem’s hypothesis; the root at λ=1 also means the system is on the stability boundary, not strictly stable. Moreover Eq. (23) does not enforce Σ_j |cAR_j| ≤ 1, so Theorem 4.3 gives no stability guarantee for the learned coefficients. The authors should either change the normalization, restrict the coefficients, or restate the theorems with hypotheses that the model actually satisfies.
- [Section 4, Theorem 4.4] As stated, Theorem 4.4 is qualitative: the measure of ‘closeness’ to the unit circle and the notion of ‘range propagation’ are not defined, and the conclusion is not a quantitative bound. Given that the central contribution is a theoretical foundation, the paper should replace this heuristic statement with a precise bound relating the spectrum of A to the number of steps over which information can propagate (e.g., a decay rate or mixing time bound). This is also important because the overall model inserts nonlinearities between blocks (Eq. (8)), so the linear-SSM analysis applies only within a block and its scope should be stated explicitly.
minor comments (6)
- [Abstract and Section 1] The phrase ‘permutation invariance’ is used where permutation equivariance is meant; these are different properties and the text should be corrected throughout.
- [Appendix G, Table 17] The summary states that GRAMA is better than state-of-the-art methods, but Table 17 reports negative improvements over the best baseline on Peptides-func (−0.57 AP) and Amazon-ratings (−0.46 Acc); please qualify these claims or report confidence intervals.
- [References] The bibliography contains several apparent errors, e.g., ‘Albert Gu, Yilun Fu, and Edward J. Liu. Mamba: A flexible mechanism...’ appears to be a garbled citation, and the entry for Topping et al. (2022) is titled ‘Understanding over-smoothing in graph neural networks’ instead of the actual oversquashing paper. Please clean up the reference list.
- [Appendix B.1, proof of Theorem 4.1] In the SSM-to-ARMA direction, the proof sets p = t and identifies the autoregressive coefficients with the p entries of CA^p without precisely specifying the indexing of the initial state; please make this derivation rigorous or cite a standard textbook treatment.
- [Appendix D, code availability] Appendix D states that code ‘will be openly released upon acceptance’; for reproducibility, please provide a repository link or a detailed pseudocode description in the current version.
- [Section 5, Tables 1–4] For most benchmarks the baselines are not matched in the number of GNN calls, skip connections, or parameter count, so the gains could partly come from added depth and residual connections rather than from the ARMA/SSM mechanism; a same-depth residual GNN control would strengthen the attribution, as Table 7 does for one dataset.
Circularity Check
No significant circularity: the SSM equivalence, stability, and spectral-decay results are standard external facts re-derived from explicit assumptions; the graph-level oversquashing claim is under-supported but not circular.
full rationale
GRAMA's derivation chain is self-contained: Eq. (2) defines the constructed sequence, Eq. (6) defines the ARMA-plus-GNN recurrence, and Theorems 4.1-4.4 with Lemma B.1 are proved from the companion-matrix SSM realization in Eqs. (11)-(14) and classical polynomial root bounds. These are standard linear-systems facts re-derived under stated assumptions, not outputs fitted from GRAMA's own predictions, so no prediction reduces to a fitted value. The self-citations in the paper (Gravina et al. 2023/2024a for A-DGN/SWAN baselines, Eliasof et al. 2024a and Mantri et al. 2024 for graph-adaptive design comments) are used as baselines or supportive context, not as load-bearing justification, and no uniqueness theorem is imported from the authors' prior work. The skeptical concern that Section 4 proves only sequence-index range and never bounds propagation across graph bottlenecks is a soundness and scope gap, not a circularity: the paper asserts the bridge between sequence decay and oversquashing rather than deriving it. Similarly, the sum-to-one normalization in Eq. (23) may put a root at lambda equal to one and so violate Theorem 4.4's strict inside-unit-circle hypothesis, but that is an applicability issue for the stated theory, not a definitional reduction of the result to its inputs. Empirical comparisons are against external benchmarks and external baselines, so the central experimental claims are not circular.
Assumptions & free parameters
free parameters (5)
- Sequence length L =
1 to 50 depending on task, chosen by grid search
- Number of GRAMA blocks S =
1 to 4 depending on task
- Hidden dimension d =
10 to 256 depending on task
- ARMA orders p and q and recurrence steps R =
p=q=R=L in all experiments
- Learned ARMA coefficients and network weights =
Fitted by gradient descent
assumptions (5)
- standard math Every ARMA(p,q) process has an equivalent linear state space representation and vice versa.
- standard math A linear recurrence is stable if and only if the spectral radius of its state matrix is at most 1.
- standard math Long-range dependence in a linear SSM is controlled by decay of powers of the state matrix.
- ad hoc to paper The L stacked embeddings of the same static features form a sequence with meaningful ARMA structure.
- domain assumption Permutation equivariance of the backbone GNN is preserved by the sequence construction and attention-pooled coefficients.
Cite this review
Pith. "Pith review of GRAMA: Adaptive Graph Autoregressive Moving Average Models." pith.science (2026). https://pith.science/paper/YSSCKJZV
@misc{pith2026250112732,
author = {Pith},
title = {Pith review of: GRAMA: Adaptive Graph Autoregressive Moving Average Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/YSSCKJZV}},
note = {Machine review of arXiv:2501.12732}
}
read the original abstract
Graph State Space Models (SSMs) have recently been introduced to enhance Graph Neural Networks (GNNs) in modeling long-range interactions. Despite their success, existing methods either compromise on permutation equivariance or limit their focus to pairwise interactions rather than sequences. Building on the connection between Autoregressive Moving Average (ARMA) and SSM, in this paper, we introduce GRAMA, a Graph Adaptive method based on a learnable Autoregressive Moving Average (ARMA) framework that addresses these limitations. By transforming from static to sequential graph data, GRAMA leverages the strengths of the ARMA framework, while preserving permutation equivariance. Moreover, GRAMA incorporates a selective attention mechanism for dynamic learning of ARMA coefficients, enabling efficient and flexible long-range information propagation. We also establish theoretical connections between GRAMA and Selective SSMs, providing insights into its ability to capture long-range dependencies. Extensive experiments on 14 synthetic and real-world datasets demonstrate that GRAMA consistently outperforms backbone models and performs competitively with state-of-the-art methods.
Figures
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Mixhop: Higher-order graph convolutional architectures via sparsified neighborhood mixing
Sami Abu-El-Haija, Bryan Perozzi, Amol Kapoor, Nazanin Alipourfard, Kristina Lerman, Hrayr Harutyunyan, Greg Ver Steeg, and Aram Galstyan. Mixhop: Higher-order graph convolutional architectures via sparsified neighborhood mixing . In International Conference on Machine Learning, pp.\ 21--29. PMLR, 2019
2019
-
[3]
On the bottleneck of graph neural networks and its practical implications
Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=i80OPhOCVH2
2021
-
[4]
State Space Models: A Unifying Framework
Masao Aoki. State Space Models: A Unifying Framework. Springer, Berlin, Germany, 2013. ISBN 978-3-642-35040-6
2013
-
[5]
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio. Unitary evolution recurrent neural networks. In International conference on machine learning, pp.\ 1120--1128. PMLR, 2016
2016
-
[6]
Accurate prediction of protein structures and interactions using a three-track neural network
Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373 0 (6557): 0 871--876, 2021
2021
-
[7]
A3T-GCN: Attention Temporal Graph Convolutional Network for Traffic Forecasting
Jiandong Bai, Jiawei Zhu, Yujiao Song, Ling Zhao, Zhixiang Hou, Ronghua Du, and Haifeng Li. A3T-GCN: Attention Temporal Graph Convolutional Network for Traffic Forecasting . ISPRS International Journal of Geo-Information, 10 0 (7), 2021. ISSN 2220-9964. doi:10.3390/ijgi10070485
-
[8]
o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \
Maximilian Beck, Korbinian P \"o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \"u nter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024
arXiv 2024
Show all 127 references
-
[9]
Graph Mamba: Towards Learning on Graphs with State Space Models , 2024
Ali Behrouz and Farnoosh Hashemi. Graph Mamba: Towards Learning on Graphs with State Space Models , 2024. URL https://arxiv.org/abs/2402.08678
2024 arXiv
-
[10]
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks, 5 0 (2): 0 157--166, 1994
1994
-
[11]
Graph neural networks with convolutional arma filters
Filippo Maria Bianchi, Simone Scardapane, Lorenzo Livi, and Cesare Alippi. Graph neural networks with convolutional arma filters. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42 0 (8): 0 2208--2222, 2019
2019
-
[12]
Beyond low-frequency information in graph convolutional networks
Deyu Bo, Xiao Wang, Chuan Shi, and Huawei Shen. Beyond low-frequency information in graph convolutional networks. Proceedings of the AAAI Conference on Artificial Intelligence, 35 0 (5): 0 3950--3957, May 2021. doi:10.1609/aaai.v35i5.16514. URL https://ojs.aaai.org/index.php/A...
2021 doi
-
[13]
Improving graph neural network expressivity via subgraph isomorphism counting
Giorgos Bouritsas, Fabrizio Frasca, Stefanos Zafeiriou, and Michael M Bronstein. Improving graph neural network expressivity via subgraph isomorphism counting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (1): 0 657--668, 2022
2022
-
[14]
Time Series Analysis: Forecasting and Control
George EP Box, Gwilym M Jenkins, and Gregory C Reinsel. Time Series Analysis: Forecasting and Control. Holden-Day, 1970
1970
-
[15]
Residual Gated Graph ConvNets
Xavier Bresson and Thomas Laurent. Residual Gated Graph ConvNets . arXiv preprint arXiv:1711.07553, 2018
2018 arXiv
-
[16]
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veli c kovi \'c . Geometric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021
2021 arXiv
-
[17]
A note on over-smoothing for graph neural networks
Chen Cai and Yusu Wang. A note on over-smoothing for graph neural networks. arXiv preprint arXiv:2006.13318, 2020
2006 arXiv
-
[18]
GRAND : Graph neural diffusion
Benjamin Paul Chamberlain, James Rowbottom, Maria Gorinova, Stefan Webb, Emanuele Rossi, and Michael M Bronstein. GRAND : Graph neural diffusion. In International Conference on Machine Learning (ICML), pp.\ 1407--1418. PMLR, 2021
2021
-
[19]
Simple and Deep Graph Convolutional Networks
Ming Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding, and Yaliang Li. Simple and Deep Graph Convolutional Networks . In Hal Daumé III and Aarti Singh (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Resear...
2020
-
[20]
Neural ordinary differential equations
Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, pp.\ 6571--6583, 2018
2018
-
[21]
Adaptive universal generalized pagerank graph neural network
Eli Chien, Jianhao Peng, Pan Li, and Olgica Milenkovic. Adaptive universal generalized pagerank graph neural network. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=n6jl7fLxrP
2021
-
[22]
Gread: Graph neural reaction-diffusion networks
Jeongwhan Choi, Seoyoung Hong, Noseong Park, and Sung-Bae Cho. Gread: Graph neural reaction-diffusion networks. In ICML, 2023
2023
-
[23]
Beletsky, Konrad M
Krzysztof Choromanski, Marcin Kuczynski, Jacek Cieszkowski, Paul L. Beletsky, Konrad M. Smith, Wojciech Gajewski, Gabriel De Masson, Tomasz Z. Broniatowski, Antonina B. Gorny, Leszek M. Kaczmarek, and Stanislaw K. Andrzejewski. Performers: A new approach to scaling transformer...
2020 arXiv
-
[24]
From block-toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked transformers
Krzysztof Choromanski, Han Lin, Haoxian Chen, Tianyi Zhang, Arijit Sehanobish, Valerii Likhosherstov, Jack Parker-Holder, Tamas Sarlos, Adrian Weller, and Thomas Weingarten. From block-toeplitz matrices to differential equations on graphs: towards a general theory for scalable...
2022
-
[25]
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al. Griffin: Mixing gated linear recurrences with local attention for efficient language models. arXiv preprint ...
2024 arXiv
-
[26]
The arma model in state space form
Piet de Jong and Jeremy Penzer. The arma model in state space form. Statistics & probability letters, 70 0 (1): 0 119--125, 2004
2004
-
[27]
Polynormer: Polynomial-expressive graph transformer in linear time
Chenhui Deng, Zichao Yue, and Zhiru Zhang. Polynormer: Polynomial-expressive graph transformer in linear time. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=hmv1LpNfXa
2024
-
[28]
On over-squashing in message passing neural networks: the impact of width, depth, and topology
Francesco Di Giovanni, Lorenzo Giusti, Federico Barbero, Giulia Luise, Pietro Li\` o , and Michael Bronstein. On over-squashing in message passing neural networks: the impact of width, depth, and topology. In Proceedings of the 40th International Conference on Machine Learning...
2023
-
[29]
Gbk-gnn: Gated bi-kernel graph neural networks for modeling both homophily and heterophily
Lun Du, Xiaozhou Shi, Qiang Fu, Xiaojun Ma, Hengyu Liu, Shi Han, and Dongmei Zhang. Gbk-gnn: Gated bi-kernel graph neural networks for modeling both homophily and heterophily. In Proceedings of the ACM Web Conference 2022, WWW '22, pp.\ 1550–1558, New York, NY, USA, 2022. Asso...
2022
-
[30]
Dwivedi and X
V. Dwivedi and X. Bresson. Benchmarking graph transformers. Journal of Machine Learning Research, 23: 0 1--32, 2022
2022
-
[31]
A Generalization of Transformer Networks to Graphs
Vijay Prakash Dwivedi and Xavier Bresson. A Generalization of Transformer Networks to Graphs . AAAI Workshop on Deep Learning on Graphs: Methods and Applications, 2021
2021
-
[32]
Graph neural networks with learnable structural and positional representations
Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Graph neural networks with learnable structural and positional representations. In International Conference on Learning Representations, 2022 a . URL https://openreview.net/forum?id=wTTjnvGphYj
2022
-
[33]
Long Range Graph Benchmark
Vijay Prakash Dwivedi, Ladislav Ramp\' a s ek, Michael Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, and Dominique Beaini. Long Range Graph Benchmark . In Advances in Neural Information Processing Systems, volume 35, pp.\ 22326--22340. Curran Associates, Inc., 2022 b
2022
-
[34]
Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson
Vijay Prakash Dwivedi, Chaitanya K. Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks. Journal of Machine Learning Research, 24 0 (43): 0 1--48, 2023. URL http://jmlr.org/papers/v24/22-0567.html
2023
-
[35]
Benchmarking graph neural networks
Vishwajeet Dwivedi, Xavier Bresson, and Lior Wolf. Benchmarking graph neural networks. Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[36]
PDE-GCN : Novel architectures for graph neural networks motivated by partial differential equations
Moshe Eliasof, Eldad Haber, and Eran Treister. PDE-GCN : Novel architectures for graph neural networks motivated by partial differential equations. Advances in Neural Information Processing Systems, 34: 0 3836--3849, 2021
2021
-
[37]
Granola: Adaptive normalization for graph neural networks
Moshe Eliasof, Beatrice Bevilacqua, Carola-Bibiane Sch \"o nlieb, and Haggai Maron. Granola: Adaptive normalization for graph neural networks. arXiv preprint arXiv:2404.13344, 2024 a
2024 arXiv
-
[38]
On the temporal domain of differential equation inspired graph neural networks
Moshe Eliasof, Eldad Haber, Eran Treister, and Carola-Bibiane B Sch \"o nlieb. On the temporal domain of differential equation inspired graph neural networks. In International Conference on Artificial Intelligence and Statistics, pp.\ 1792--1800. PMLR, 2024 b
2024
-
[39]
Bronstein, and Ismail Ilkan Ceylan
Ben Finkelshtein, Xingyue Huang, Michael M. Bronstein, and Ismail Ilkan Ceylan. Cooperative Graph Neural Networks . In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=ZQcqXCuoxD
2024
-
[40]
A large-scale database for graph representation learning
Scott Freitas, Yuxiao Dong, Joshua Neil, and Duen Horng Chau. A large-scale database for graph representation learning. In J. Vanschoren and S. Yeung (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021
2021
-
[41]
S4: Structured state space for scalable and efficient sequence modeling
Deng Fu, Yao Ma, and Yao Qian. S4: Structured state space for scalable and efficient sequence modeling. Proceedings of the 41st International Conference on Machine Learning (ICML), 2023
2023
-
[42]
Diffusion Improves Graph Learning
Johannes Gasteiger, Stefan Wei enberger, and Stephan G\" u nnemann. Diffusion Improves Graph Learning . In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019
2019
-
[43]
Anti-Symmetric DGN: a stable architecture for Deep Graph Networks
Alessio Gravina, Davide Bacciu, and Claudio Gallicchio. Anti-Symmetric DGN: a stable architecture for Deep Graph Networks . In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=J3Y7cgZOOS
2023
-
[44]
Tackling Oversquashing by Global and Local Non-Dissipativity
Alessio Gravina, Moshe Eliasof, Claudio Gallicchio, Davide Bacciu, and Carola-Bibiane Sch\"onlieb. Tackling Oversquashing by Global and Local Non-Dissipativity . arXiv preprint arXiv:2405.01009, 2024 a
2024 arXiv
-
[45]
Temporal graph odes for irregularly-sampled time series
Alessio Gravina, Daniele Zambon, Davide Bacciu, and Cesare Alippi. Temporal graph odes for irregularly-sampled time series. In Kate Larson (ed.), Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24 , pp.\ 4025--4034. Internationa...
2024 doi
-
[46]
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher R \'e . Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021 a
2021 arXiv
-
[47]
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R \'e . Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34: 0 572--585, 2021 b
2021
-
[48]
Albert Gu, Yilun Fu, and Edward J. Liu. Mamba: A flexible mechanism for long-range dependencies in state space models. NeurIPS, 2023
2023
-
[49]
Siany Gu, Xin Yu, and Viktor K. Chao. Structured state space models for efficient sequence modeling. Proceedings of the 38th International Conference on Machine Learning (ICML), 2021 c
2021
-
[50]
Drew: Dynamically rewired message passing with delay
Benjamin Gutteridge, Xiaowen Dong, Michael M Bronstein, and Francesco Di Giovanni. Drew: Dynamically rewired message passing with delay . In International Conference on Machine Learning, pp.\ 12252--12267. PMLR, 2023
2023
-
[51]
Hamilton
James D. Hamilton. Time Series Analysis. Princeton University Press, 1994 a
1994
-
[52]
State-space models
James D Hamilton. State-space models. Handbook of econometrics, 4: 0 3039--3080, 1994 b
1994
-
[53]
Hamilton, Rex Ying, and Jure Leskovec
William L. Hamilton, Rex Ying, and Jure Leskovec. Inductive representation learning on large graphs . In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, pp.\ 1025–1035. Curran Associates Inc., 2017. ISBN 9781510860964
2017
-
[54]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 770--778, 2016
2016
-
[55]
Bounding the roots of polynomials
Holly P Hirst and Wade T Macey. Bounding the roots of polynomials. The College Mathematics Journal, 28 0 (4): 0 292--295, 1997
1997
-
[56]
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, J \"u rgen Schmidhuber, et al. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
2001
-
[57]
Matrix analysis
Roger A Horn and Charles R Johnson. Matrix analysis. Cambridge university press, 2012
2012
-
[58]
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Informat...
2020
-
[59]
Strategies for Pre-training Graph Neural Networks
Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. Strategies for Pre-training Graph Neural Networks . In International Conference on Learning Representations, 2020 b . URL https://openreview.net/forum?id=HJlWWJSFDH
2020
-
[60]
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 4700--4708, 2017
2017
-
[61]
What can we learn from state space models for machine learning on graphs? arXiv preprint arXiv:2406.05815, 2024
Yinan Huang, Siqi Miao, and Pan Li. What can we learn from state space models for machine learning on graphs? arXiv preprint arXiv:2406.05815, 2024
2024 arXiv
-
[62]
Autoregressive moving average graph filtering
Elvin Isufi, Andreas Loukas, Andrea Simonetto, and Geert Leus. Autoregressive moving average graph filtering. IEEE Transactions on Signal Processing, 65 0 (2): 0 274--288, 2016
2016
-
[63]
Unleashing the potential of fractional calculus in graph neural networks with FROND
Qiyu Kang, Kai Zhao, Qinxu Ding, Feng Ji, Xuhao Li, Wenfei Liang, Yang Song, and Wee Peng Tay. Unleashing the potential of fractional calculus in graph neural networks with FROND . In The Twelfth International Conference on Learning Representations, 2024. URL https://openrevie...
2024
-
[64]
Banerjee, and Guido Montufar
Kedar Karhadkar, Pradeep Kr. Banerjee, and Guido Montufar. Fo SR : First-order spectral rewiring for addressing oversquashing in GNN s. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=3YjQfCLdrzz
2023
-
[65]
Hassan K. Khalil. Nonlinear Systems. Prentice Hall, Upper Saddle River, NJ, 3rd edition, 2002. ISBN 978-0130673893
2002
-
[66]
A review of graph neural networks: concepts, architectures, techniques, challenges, datasets, applications, and future directions
Bharti Khemani, Shruti Patil, Ketan Kotecha, and Sudeep Tanwar. A review of graph neural networks: concepts, architectures, techniques, challenges, datasets, applications, and future directions. Journal of Big Data, 11 0 (1): 0 18, 2024
2024
-
[67]
Kipf and M
T. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. Proceedings of the International Conference on Learning Representations, 2016
2016
-
[68]
Bayan Bruss, and Tom Goldstein
Kezhi Kong, Jiuhai Chen, John Kirchenbauer, Renkun Ni, C. Bayan Bruss, and Tom Goldstein. GOAT : A global transformer on large-scale graphs. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.), Proceedings of the 40t...
2023
-
[69]
Rethinking graph transformers with spectral attention
Devin Kreuzer, Dominique Beaini, Will Hamilton, Vincent L \'e tourneau, and Prudencio Tossou. Rethinking graph transformers with spectral attention. Advances in Neural Information Processing Systems, 34: 0 21618--21629, 2021 a
2021
-
[70]
Kreuzer et al
M. Kreuzer et al. Positional encodings in graph transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43: 0 345--356, 2021 b
2021
-
[71]
Sven Kreuzer, Michael Reiner, and Stefan D. D. De Villiers. Sant: Structural attention networks for graphs. Proceedings of the 38th International Conference on Machine Learning (ICML), 2021 c
2021
-
[72]
Finding global homophily in graph neural networks when meeting heterophily
Xiang Li, Renyu Zhu, Yao Cheng, Caihua Shan, Siqiang Luo, Dongsheng Li, and Weining Qian. Finding global homophily in graph neural networks when meeting heterophily. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.), Proceedi...
2022
-
[73]
Toloker Graph: Interaction of Crowd Annotators , February 2023
Daniil Likhobaba, Nikita Pavlichenko, and Dmitry Ustalov. Toloker Graph: Interaction of Crowd Annotators , February 2023. URL https://doi.org/10.5281/zenodo.7620796
2023 doi
-
[74]
Mamba: Beyond long sequences
Chao Liu, Hongdong Li, and Tao Xu. Mamba: Beyond long sequences. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[75]
Li, Jian Tang, Guy Wolf, and Stefanie Jegelka
Sitao Luan, Chenqing Hua, Qincheng Lu, Liheng Ma, Lirong Wu, Xinyu Wang, Minkai Xu, Xiao-Wen Chang, Doina Precup, Rex Ying, Stan Z. Li, Jian Tang, Guy Wolf, and Stefanie Jegelka. The heterophilic graph learning handbook: Benchmarks, models, theoretical analysis, applications a...
2024 arXiv
-
[76]
Digraf: Diffeomorphic graph-adaptive activation function
Krishna Sri Ipsit Mantri, Xinzhi Wang, Carola-Bibiane Sch \"o nlieb, Bruno Ribeiro, Beatrice Bevilacqua, and Moshe Eliasof. Digraf: Diffeomorphic graph-adaptive activation function. arXiv preprint arXiv:2407.02013, 2024
2024 arXiv
-
[77]
A fractional graph laplacian approach to oversmoothing
Sohir Maskey, Raffaele Paolino, Aras Bacho, and Gitta Kutyniok. A fractional graph laplacian approach to oversmoothing. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=kS7ED7eE74
2023
-
[78]
Simplifying approach to node classification in graph neural networks
Sunil Kumar Maurya, Xin Liu, and Tsuyoshi Murata. Simplifying approach to node classification in graph neural networks. Journal of Computational Science, 62: 0 101695, 2022. ISSN 1877-7503. doi:https://doi.org/10.1016/j.jocs.2022.101695. URL https://www.sciencedirect.com/scien...
2022
-
[79]
Weisfeiler and leman go neural: Higher-order graph neural networks
Christopher Morris, Martin Ritzert, Matthias Fey, William L Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and leman go neural: Higher-order graph neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp.\ 4602--...
2019
-
[80]
Attending to graph transformers
Luis M \"u ller, Mikhail Galkin, Christopher Morris, and Ladislav Ramp \'a s ek. Attending to graph transformers. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=HhbqHBBrfZ
2024
-
[81]
Nguyen et al
H. Nguyen et al. Efficiency in state space models for sequence learning. Journal of Machine Learning Research, 24: 0 3678--3690, 2023
2023
-
[82]
Revisiting graph neural networks: All we have is low-pass filters
Hoang Nt and Takanori Maehara. Revisiting graph neural networks: All we have is low-pass filters. arXiv preprint arXiv:1905.09550, 2019
1905 arXiv
-
[83]
Graph neural networks exponentially lose expressive power for node classification
Kenta Oono and Taiji Suzuki. Graph neural networks exponentially lose expressive power for node classification. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=S1ldO2EFPr
2020
-
[84]
On the universality of linear recurrences followed by nonlinear projections
Antonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu, and Samuel L Smith. On the universality of linear recurrences followed by nonlinear projections. arXiv preprint arXiv:2307.11888, 2023 a
2023 arXiv
-
[85]
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. Resurrecting recurrent neural networks for long sequences. In International Conference on Machine Learning, pp.\ 26670--26698. PMLR, 2023 b
2023
-
[86]
Permutation equivariant layers for higher order interactions
Horace Pan and Risi Kondor. Permutation equivariant layers for higher order interactions. In International Conference on Artificial Intelligence and Statistics, pp.\ 5987--6001. PMLR, 2022
2022
-
[87]
On the difficulty of training recurrent neural networks
R Pascanu. On the difficulty of training recurrent neural networks. arXiv preprint arXiv:1211.5063, 2013
2013 arXiv
-
[88]
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019
2019
-
[89]
Geom-gcn: Geometric graph convolutional networks
Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei, and Bo Yang. Geom-gcn: Geometric graph convolutional networks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=S1e2agrFvS
2020
-
[90]
A critical look at the evaluation of GNN s under heterophily: Are we really making progress? In The Eleventh International Conference on Learning Representations, 2023
Oleg Platonov, Denis Kuznedelev, Michael Diskin, Artem Babenko, and Liudmila Prokhorenkova. A critical look at the evaluation of GNN s under heterophily: Are we really making progress? In The Eleventh International Conference on Learning Representations, 2023. URL https://open...
2023
-
[91]
Graph neural ordinary differential equations
Michael Poli, Stefano Massaroli, Junyoung Park, Atsushi Yamashita, Hajime Asama, and Jinkyoo Park. Graph neural ordinary differential equations. arXiv preprint arXiv:1911.07532, 2019
1911 arXiv
-
[92]
Recipe for a General, Powerful, Scalable Graph Transformer
Ladislav Ramp\' a s ek, Mikhail Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini. Recipe for a General, Powerful, Scalable Graph Transformer . Advances in Neural Information Processing Systems, 35, 2022
2022
-
[93]
Pytorch geometric temporal: Spatiotemporal signal processing with neural machine learning models
Benedek Rozemberczki, Paul Scherer, Yixuan He, George Panagopoulos, Alexander Riedel, Maria Astefanoaei, Oliver Kiss, Ferenc Beres, Guzm\' a n L\' o pez, Nicolas Collignon, and Rik Sarkar. Pytorch geometric temporal: Spatiotemporal signal processing with neural machine learnin...
2021
-
[94]
Graph-coupled oscillator networks
T Konstantin Rusch, Ben Chamberlain, James Rowbottom, Siddhartha Mishra, and Michael Bronstein. Graph-coupled oscillator networks. In International Conference on Machine Learning, pp.\ 18888--18909. PMLR, 2022
2022
-
[95]
Konstantin Rusch, Michael M
T. Konstantin Rusch, Michael M. Bronstein, and Siddhartha Mishra. A Survey on Oversmoothing in Graph Neural Networks . arXiv preprint arXiv:2303.10993, 2023
2023 arXiv
-
[96]
Deep neural networks motivated by partial differential equations
Lars Ruthotto and Eldad Haber. Deep neural networks motivated by partial differential equations. Journal of Mathematical Imaging and Vision, 62: 0 352--364, 2020
2020
-
[97]
Theoretical guarantees for permutation-equivariant quantum neural networks
Louis Schatzki, Martin Larocca, Quynh T Nguyen, Frederic Sauvage, and Marco Cerezo. Theoretical guarantees for permutation-equivariant quantum neural networks. npj Quantum Information, 10 0 (1): 0 12, 2024
2024
-
[98]
Masked label prediction: Unified message passing model for semi-supervised classification
Yunsheng Shi, Zhengjie Huang, Shikun Feng, Hui Zhong, Wenjing Wang, and Yu Sun. Masked label prediction: Unified message passing model for semi-supervised classification. In Zhi-Hua Zhou (ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intellige...
2021 doi
-
[99]
Rahmani, and Marzieh Aghaei
Behzad Shirzad, Amir M. Rahmani, and Marzieh Aghaei. Exphormer: Sparse attention for graphs. Proceedings of the 40th International Conference on Machine Learning (ICML), 2023
2023
-
[100]
Applied nonlinear control, volume 199
Jean-Jacques E Slotine, Weiping Li, et al. Applied nonlinear control, volume 199. Prentice hall Englewood Cliffs, NJ, 1991
1991
-
[101]
Where did the gap go? reassessing the long-range graph benchmark
Jan T \"o nshoff, Martin Ritzert, Eran Rosenbluth, and Martin Grohe. Where did the gap go? reassessing the long-range graph benchmark. In The Second Learning on Graphs Conference, 2023. URL https://openreview.net/forum?id=rIUjwxc5lj
2023
-
[102]
Understanding over-smoothing in graph neural networks
Matthew Topping, Sebastian Ruder, and Chris Dyer. Understanding over-smoothing in graph neural networks. Proceedings of the 39th International Conference on Machine Learning (ICML), 2022
2022
-
[103]
Vaswani et al
A. Vaswani et al. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017
2017
-
[104]
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJXMpikCZ
2018
-
[105]
Graph-mamba: Towards long-range graph sequence modeling with selective state spaces
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang. Graph-mamba: Towards long-range graph sequence modeling with selective state spaces. arXiv preprint arXiv:2402.00789, 2024 a
2024 arXiv
-
[106]
The heterophilic snowflake hypothesis: Training and empowering gnns for heterophilic graphs
Kun Wang, Guibin Zhang, Xinnan Zhang, Junfeng Fang, Xun Wu, Guohao Li, Shirui Pan, Wei Huang, and Yuxuan Liang. The heterophilic snowflake hypothesis: Training and empowering gnns for heterophilic graphs. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery ...
2024
-
[107]
How powerful are spectral graph neural networks
Xiyuan Wang and Muhan Zhang. How powerful are spectral graph neural networks. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings ...
2022
-
[108]
Dissecting the Diffusion Process in Linear Graph Convolutional Networks
Yifei Wang, Yisen Wang, Jiansheng Yang, and Zhouchen Lin. Dissecting the Diffusion Process in Linear Graph Convolutional Networks . In Advances in Neural Information Processing Systems, volume 34, pp.\ 5758--5769. Curran Associates, Inc., 2021
2021
-
[109]
ACMP : Allen-cahn message passing with attractive and repulsive forces for graph neural networks
Yuelin Wang, Kai Yi, Xinliang Liu, Yu Guang Wang, and Shi Jin. ACMP : Allen-cahn message passing with attractive and repulsive forces for graph neural networks. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=4fZc_79Lrqs
2023
-
[110]
P. Whittle. Hypothesis Testing in Time Series Analysis. Statistics / Uppsala universitet. Almqvist & Wiksells boktr., 1951. ISBN 9780598919823
1951
-
[111]
Kdlgt: A linear graph transformer framework via kernel decomposition approach
Yi Wu, Yanyang Xu, Wenhao Zhu, Guojie Song, Zhouchen Lin, Liang Wang, and Shaoguo Liu. Kdlgt: A linear graph transformer framework via kernel decomposition approach. In IJCAI, pp.\ 2370--2378, 2023
2023
-
[112]
Representation learning on graphs with jumping knowledge networks
Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. In Jennifer Dy and Andreas Krause (eds.), Proceedings of the 35th International Conference on Machine Learni...
2018
-
[113]
How powerful are graph neural networks? In International Conference on Learning Representations, 2019
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=ryGs6iA5Km
2019
-
[114]
Cohen, and Ruslan Salakhutdinov
Zhilin Yang, William W. Cohen, and Ruslan Salakhutdinov. Revisiting semi-supervised learning with graph embeddings. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML'16, pp.\ 40–48. JMLR.org, 2016
2016
-
[115]
Graphormer: A transformer for graphs
Zhitao Ying and Jure Leskovec. Graphormer: A transformer for graphs. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021
2021
-
[116]
A review of recurrent neural networks: Lstm cells and network architectures
Yong Yu, Xiaosheng Si, Changhua Hu, and Jianxun Zhang. A review of recurrent neural networks: Lstm cells and network architectures. Neural computation, 31 0 (7): 0 1235--1270, 2019
2019
-
[117]
Yun et al
S. Yun et al. Graph transformers for long-range dependencies. IEEE Transactions on Neural Networks, 30: 0 2451--2462, 2019
2019
-
[118]
H., Lihong Wang, S
Manzil Zaheer, Guru prasad G. H., Lihong Wang, S. V. K. N. L. Wang, Yujia Li, Jakub Konečný, Shalmali Joshi, Danqi Chen, Jennifer R. R., Zhenyu Zhang, Shalini Devaraj, and Srinivas Narayanan. Bigbird: Transformers for longer sequences. Proceedings of the 37th International Con...
2020 arXiv
-
[119]
Linked dynamic graph cnn: Learning on point cloud via linking hierarchical features
Kuangen Zhang, Ming Hao, Jing Wang, Clarence W de Silva, and Chenglong Fu. Linked dynamic graph cnn: Learning on point cloud via linking hierarchical features. arXiv preprint arXiv:1904.10014, 2019
1904 arXiv
-
[120]
K. Zhao, Q. Kang, Y. Song, R. She, S. Wang, and W. P. Tay. Graph neural convection-diffusion with heterophily. In Proc. International Joint Conference on Artificial Intelligence, Macao, China, Aug 2023
2023
-
[121]
T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction
Ling Zhao, Yujiao Song, Chao Zhang, Yu Liu, Pu Wang, Tao Lin, Min Deng, and Haifeng Li. T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction . IEEE Transactions on Intelligent Transportation Systems, 21 0 (9): 0 3848--3858, 2020. doi:10.1109/TITS.2019.2935152
2020
-
[122]
Beyond homophily in graph neural networks: Current limitations and effective designs
Jiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann, Leman Akoglu, and Danai Koutra. Beyond homophily in graph neural networks: Current limitations and effective designs. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Pro...
2020
-
[123]
Rossi, Anup Rao, Tung Mai, Nedim Lipka, Nesreen K
Jiong Zhu, Ryan A. Rossi, Anup Rao, Tung Mai, Nedim Lipka, Nesreen K. Ahmed, and Danai Koutra. Graph neural networks with heterophily. Proceedings of the AAAI Conference on Artificial Intelligence, 35 0 (12): 0 11168--11176, May 2021. doi:10.1609/aaai.v35i12.17332. URL https:/...
2021 doi
-
[124]
Ordinary differential equations on graph networks
Juntang Zhuang, Nicha Dvornek, Xiaoxiao Li, and James S Duncan. Ordinary differential equations on graph networks. 2020
2020
-
[125]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[126]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[127]
and ``0
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
1949
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.