Pith. sign in

REVIEW 9 cited by

Do Transformers Really Perform Bad for Graph Representation?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05234 v5 pith:DKGX6GQ5 submitted 2021-06-09 cs.LG cs.AI

Do Transformers Really Perform Bad for Graph Representation?

classification cs.LG cs.AI
keywords graphgraphormerencodingrepresentationstructuraltransformerarchitectureinformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GLACIER: Rethinking Mass Spectrum Prediction as an Object Detection Problem

    cs.LG 2026-06 unverdicted novelty 7.0

    GLACIER is a single-stage transformer model treating MS/MS fragmentation as subgraph detection on molecular graphs, reporting 70.0% Top-1 accuracy on MassSpecGym and 8x speedup over prior two-stage methods.

  2. Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks

    cs.NI 2026-05 conditional novelty 7.0

    Graph transformer RL for dynamic RMSA supports up to 13% more traffic than benchmarks on networks up to 143 nodes and 362 links.

  3. Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks

    cs.NI 2026-05 unverdicted novelty 7.0

    A graph transformer with RL stabilizations is the first to exceed benchmarks for dynamic RMSA, supporting up to 13% more traffic load on networks up to 143 nodes.

  4. Graph Neural Networks for the Graphical Bootstrap

    hep-th 2026-07 conditional novelty 6.0

    GNNs and graph transformers classify vanishing coefficients on millions of N=4 SYM f-graphs, generalizing to larger n with 99.996% ROC AUC and pruning up to 85.5% of redundant d-graphs.

  5. GCCM: Enhancing Generative Graph Prediction via Contrastive Consistency Model

    cs.AI 2026-05 unverdicted novelty 6.0

    GCCM prevents shortcut collapse in consistency models for graph prediction by using contrastive negative pairs and input feature perturbation, leading to better performance than deterministic baselines.

  6. Deep sequence models tend to memorize geometrically; it is unclear why

    cs.LG 2025-10 unverdicted novelty 6.0

    Deep sequence models develop geometric memory in embeddings that encodes novel global relationships, transforming l-fold composition tasks into 1-step navigation via a natural spectral bias connected to Node2Vec.

  7. QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

    cs.AI 2026-07 conditional novelty 5.0

    QLPO resamples GRPO training groups to favor short correct and long incorrect responses, cutting reasoning length substantially while keeping accuracy roughly unchanged.

  8. Closed-Loop Molecular Design with Calibrated Deference

    cs.CE 2026-05 unverdicted novelty 5.0

    CLIO agent applies calibrated deference in closed-loop AORFB negolyte design, achieving 90 mV redox potential gain with restored reversibility after hypothesis-driven redesign from phosphonate to sulfonate.

  9. How Embeddings Shape Graph Neural Networks: Classical vs Quantum-Oriented Node Representations

    cs.LG 2026-04 unverdicted novelty 5.0

    Quantum-oriented embeddings deliver consistent gains on structure-driven graph datasets while classical baselines perform adequately on attribute-limited social graphs, under identical training pipelines across five T...