Pith. sign in

REVIEW 14 cited by

OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.09430 v3 pith:ZXWBTKDH submitted 2021-03-17 cs.LG

classification cs.LG
keywords graphdatasetslarge-scaleogb-lsclearningbaselinechallengededicated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Enabling effective and efficient machine learning (ML) over large-scale graph data (e.g., graphs with billions of edges) can have a great impact on both industrial and scientific applications. However, existing efforts to advance large-scale graph ML have been largely limited by the lack of a suitable public benchmark. Here we present OGB Large-Scale Challenge (OGB-LSC), a collection of three real-world datasets for facilitating the advancements in large-scale graph ML. The OGB-LSC datasets are orders of magnitude larger than existing ones, covering three core graph learning tasks -- link prediction, graph regression, and node classification. Furthermore, we provide dedicated baseline experiments, scaling up expressive graph ML models to the massive datasets. We show that expressive models significantly outperform simple scalable baselines, indicating an opportunity for dedicated efforts to further improve graph ML at scale. Moreover, OGB-LSC datasets were deployed at ACM KDD Cup 2021 and attracted more than 500 team registrations globally, during which significant performance improvements were made by a variety of innovative techniques. We summarize the common techniques used by the winning solutions and highlight the current best practices in large-scale graph ML. Finally, we describe how we have updated the datasets after the KDD Cup to further facilitate research advances. The OGB-LSC datasets, baseline code, and all the information about the KDD Cup are available at https://ogb.stanford.edu/docs/lsc/ .

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    GFFMERGE formulates GNN force field merging as a convex embedding-alignment problem with an analytical solution, recovering near joint-training performance on MD17, MD22, LiPS20 and other benchmarks while delivering 5...

  2. Fused Gromov-Wasserstein Distance with Feature Selection

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Fused Gromov-Wasserstein distances are extended with feature selection via Lasso/Ridge regularization or simplex-constrained weights, yielding theoretical bounds, metric properties, and an alternating minimization algorithm.

  3. The limits of bio-molecular modeling with large language models : a cross-scale evaluation

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    LLMs perform adequately on bio-molecular classification tasks but remain weak on regression, with hybrid architectures outperforming others on long sequences and fine-tuning hurting generalization.

  4. DSBD: Dual-Aligned Structural Basis Distillation for Graph Domain Adaptation

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    DSBD distills a dual-aligned structural basis to adapt GNNs across graphs with structural distribution shifts, outperforming prior methods on benchmarks.

  5. DisRFM: Polar Riemannian Flow Matching for Structure-Preserving Graph Domain Adaptation

    cs.LG 2026-01 unverdicted novelty 7.0 of 10

    DisRFM uses polar Riemannian flow matching on constant-curvature manifolds to align graph domains while preserving label-relevant topology via radial Wasserstein and angular confidence matching.

  6. Cross-Resolution Semantic Learning for Graph Domain Adaptation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    CReSL improves graph domain adaptation by learning cross-resolution source-to-target routing and grafting target representations toward source class prototypes.

  7. GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    GraspLLM extracts dataset-agnostic structural patterns via motif contrastive learning and aligns contextual subgraphs to LLM tokens, outperforming prior LLM-based methods on TAGs especially in zero-shot settings.

  8. Agentic Fusion of Large Atomic and Language Models to Accelerate Superconductor Discovery

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    An agentic framework fusing large atomic and language models rediscovers 66 known superconductors and guides experimental verification of four new ones with transition temperatures from 2.5 K to 6.5 K.

  9. Pretraining a Foundation Model for Small-Molecule Natural Products

    q-bio.QM 2025-03 unverdicted novelty 6.0 of 10

    NaFM is a pretrained foundation model for natural products using scaffold-focused contrastive learning and masked graph objectives that achieves SOTA on taxonomy classification, gene/microbial analysis, and virtual sc...

  10. Handling Feature Heterogeneity with Learnable Graph Patches

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Learnable graph patches enable domain-agnostic pre-training of graph models by decomposing heterogeneous graphs into transferable semantic units via patch encoders and aggregators.

  11. Safe-Subspace Pseudo-Label Refinement for Source-Free Graph Domain Adaptation

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    S²PLR identifies a safe subspace for reliable pseudo-labels in source-free graph domain adaptation using semantic committee signals and structural contrastive verification, then applies noise-tolerant regularization t...

  12. Graph Hierarchical Recurrence for Long-Range Generalization

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    GHR uses hierarchical recurrence on pooled graph abstractions to improve long-range dependency capture and out-of-range generalization while using far fewer parameters than existing models.

  13. Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction

    cs.CL 2026-05 unverdicted novelty 4.0 of 10

    GA-S2S integrates T5 with RGAT to jointly process text and k-hop subgraph topology for knowledge graph link prediction, reporting up to 19% relative accuracy gain over seq2seq baselines on CoDEx.

  14. AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation

    cs.LG 2026-02 unverdicted novelty 4.0 of 10

    AdvSynGNN uses multi-resolution structural synthesis, contrastive objectives, an adaptive transformer, and an adversarial propagation engine with residual label correction to improve node-level predictions on challeng...

Pith tools