A new distributed framework for graph transformer training auto-selects parallel strategies and optimizes sparse operations to deliver up to 6x speedup on 8 GPUs and 78% memory reduction.
Exphormer: Sparse Transformers for Graphs
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Transductive Sharpening adds an entropy-minimization term on unlabeled-node predictions to the training objective for graph node classification.
SigGate-GT adds a per-head sigmoid gate to graph transformer attention outputs, relaxing the softmax convex-combination constraint to reduce over-smoothing and improve stability at ~1% parameter overhead.
citing papers explorer
-
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
A new distributed framework for graph transformer training auto-selects parallel strategies and optimizes sparse operations to deliver up to 6x speedup on 8 GPUs and 78% memory reduction.
-
Graph Transductive Sharpening: Leveraging Unlabeled Predictions in Node Classification
Transductive Sharpening adds an entropy-minimization term on unlabeled-node predictions to the training objective for graph node classification.
-
Capacity-Controlled Global Attention for Graph Transformers
SigGate-GT adds a per-head sigmoid gate to graph transformer attention outputs, relaxing the softmax convex-combination constraint to reduce over-smoothing and improve stability at ~1% parameter overhead.