Pith. sign in

REVIEW 19 cited by

Open Graph Benchmark: Datasets for Machine Learning on Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00687 v7 pith:UC5OZ7SX submitted 2020-05-02 cs.LG cs.SIstat.ML

Open Graph Benchmark: Datasets for Machine Learning on Graphs

classification cs.LG cs.SIstat.ML
keywords datasetsgraphbenchmarkdataevaluationgraphscodedataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present the Open Graph Benchmark (OGB), a diverse set of challenging and realistic benchmark datasets to facilitate scalable, robust, and reproducible graph machine learning (ML) research. OGB datasets are large-scale, encompass multiple important graph ML tasks, and cover a diverse range of domains, ranging from social and information networks to biological networks, molecular graphs, source code ASTs, and knowledge graphs. For each dataset, we provide a unified evaluation protocol using meaningful application-specific data splits and evaluation metrics. In addition to building the datasets, we also perform extensive benchmark experiments for each dataset. Our experiments suggest that OGB datasets present significant challenges of scalability to large-scale graphs and out-of-distribution generalization under realistic data splits, indicating fruitful opportunities for future research. Finally, OGB provides an automated end-to-end graph ML pipeline that simplifies and standardizes the process of graph data loading, experimental setup, and model evaluation. OGB will be regularly updated and welcomes inputs from the community. OGB datasets as well as data loaders, evaluation scripts, baseline code, and leaderboards are publicly available at https://ogb.stanford.edu .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces

    cs.LG 2026-06 unverdicted novelty 7.0

    GraphDR-LinUCB projects contextual bandit arms onto a graph's low-frequency eigenspace to obtain the first Õ(k√T) regret bound under approximate smoothness, with a spectral predictor Γ_k that matches outcomes on five ...

  2. Learning Dynamic Stability Landscapes in Synchronization Networks

    cs.LG 2026-05 unverdicted novelty 7.0

    Introduces graph-to-image prediction of per-node dynamic stability landscapes in oscillator networks from topology, releases two 10k-graph datasets, and shows GNN-CNN models achieve good accuracy with cross-size gener...

  3. Relevant Walk Search for Explaining Graph Neural Networks

    cs.LG 2026-05 unverdicted novelty 7.0

    Polynomial-time max-product algorithms for exact (neuron-level) and approximate (node-level) top-K relevant walk search in GNN-LRP explanations.

  4. Evaluating LLMs on Large-Scale Graph Property Estimation via Random Walks

    cs.LG 2026-05 unverdicted novelty 7.0

    EstGraph benchmark evaluates LLMs on estimating properties of very large graphs from random-walk samples that fit in context limits.

  5. TRAVELFRAUDBENCH: A Configurable Evaluation Framework for GNN Fraud Ring Detection in Travel Networks

    cs.LG 2026-04 unverdicted novelty 7.0

    TravelFraudBench is a new configurable benchmark for GNN-based fraud ring detection in travel networks, simulating star, clique, and chain topologies and showing GraphSAGE outperforming MLP baselines on AUC and ring recovery.

  6. How Attentive are Graph Attention Networks?

    cs.LG 2021-05 conditional novelty 7.0

    GAT uses static attention where neighbor rankings ignore the query node and thus cannot express some graph problems; GATv2 enables dynamic attention and outperforms GAT on 11 OGB and other benchmarks with equal parameters.

  7. T3R: Deeper Test-Time Adaptation for Graph Neural Networks via Gradient Rotation

    cs.LG 2026-06 unverdicted novelty 6.0

    T3R applies multiple Rotograd matrices and a rotation technique to create surrogate gradients, enabling deeper test-time adaptation in GNNs and yielding 0.172 MAE reduction plus 9.37% relative gains on OGB benchmarks.

  8. Efficient Higher-order Subgraph Attribution via Message Passing

    cs.LG 2026-05 unverdicted novelty 6.0

    Message-passing algorithms compute GNN-LRP subgraph attributions in linear time w.r.t. network depth by exploiting the distributive property.

  9. H3: A Healthcare Three-Hop Index for Physician Referral Network Prediction

    cs.SI 2026-05 unverdicted novelty 6.0

    H3 is a new three-hop index that predicts physician referrals using normalized indirect pathways and outperforms heuristics and neural nets on Medicare shared-patient data in both within-period and cross-period settings.

  10. GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning

    cs.LG 2026-05 unverdicted novelty 6.0

    GraphSculptor builds efficient pre-training coresets for graph self-supervised learning using combined structural and semantic diversity metrics, achieving 99.6% performance with 10% of the data.

  11. Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3

    cs.DC 2026-04 unverdicted novelty 6.0

    Cerebras CS-3 achieves up to 100x speedup over CPU for SpMM and 20x for SDDMM at 90% sparsity, with performance improving for larger matrices, but becomes slower than CPU beyond 99% sparsity.

  12. Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training

    cs.LG 2026-04 unverdicted novelty 6.0

    ScaleGNN uses communication-free sampling and 4D parallelism to scale mini-batch GNN training to 2048 GPUs, achieving 3.5x speedup over prior state-of-the-art on ogbn-products.

  13. SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication

    cs.DC 2025-12 unverdicted novelty 6.0

    SHIRO achieves geometric mean speedups of 221.5x to 8.8x over four baselines in distributed SpMM on up to 128 GPUs by exploiting sparsity patterns and two-tier network topologies.

  14. Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks

    cs.LG 2019-09 unverdicted novelty 6.0

    DGL is a graph-centric library that optimizes GNNs via generalized sparse tensor operations, transparent graph-based optimizations, and framework-neutral design, claiming superior speed and memory use over other GNN f...

  15. DeXposure-Claw: An Agentic System for DeFi Risk Supervision

    cs.AI 2026-06 unverdicted novelty 5.0

    Presents DeXposure-Claw, an agentic supervision system that combines DeXposure-FM forecasts, deterministic monitors, and DeXposure-Bench evaluation on five years of DeFi data.

  16. DeXposure-Claw: An Agentic System for DeFi Risk Supervision

    cs.AI 2026-06 unverdicted novelty 5.0

    DeXposure-Claw combines a graph time-series foundation model for forecasting DeFi networks with rule-based monitors and data-health gates to emit regulator-aligned risk tickets, evaluated via a new six-axis benchmark ...

  17. Handling Feature Heterogeneity with Learnable Graph Patches

    cs.LG 2026-06 unverdicted novelty 5.0

    Learnable graph patches enable domain-agnostic pre-training of graph models by decomposing heterogeneous graphs into transferable semantic units via patch encoders and aggregators.

  18. On Efficient Scaling of GNNs via IO-Aware Layers Implementations

    cs.LG 2026-05 unverdicted novelty 5.0

    IO-aware GPU kernels for SpMM convolutions, degree-aware reductions, and fused attention layers deliver median speedups of 1.6-2.6x (up to 10x) and memory reductions up to 76x over DGL/PyG baselines on realistic graphs.

  19. Fast and Featureless Node Representation Learning with Partial Pairwise Supervision

    cs.LG 2026-05 unverdicted novelty 5.0

    Contrastive FUSE learns node embeddings from partial pairwise supervision and structural signals alone by optimizing a spectral contrastive objective with a lightweight modularity approximation, yielding competitive p...