REVIEW 14 cited by
OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs
read the original abstract
Enabling effective and efficient machine learning (ML) over large-scale graph data (e.g., graphs with billions of edges) can have a great impact on both industrial and scientific applications. However, existing efforts to advance large-scale graph ML have been largely limited by the lack of a suitable public benchmark. Here we present OGB Large-Scale Challenge (OGB-LSC), a collection of three real-world datasets for facilitating the advancements in large-scale graph ML. The OGB-LSC datasets are orders of magnitude larger than existing ones, covering three core graph learning tasks -- link prediction, graph regression, and node classification. Furthermore, we provide dedicated baseline experiments, scaling up expressive graph ML models to the massive datasets. We show that expressive models significantly outperform simple scalable baselines, indicating an opportunity for dedicated efforts to further improve graph ML at scale. Moreover, OGB-LSC datasets were deployed at ACM KDD Cup 2021 and attracted more than 500 team registrations globally, during which significant performance improvements were made by a variety of innovative techniques. We summarize the common techniques used by the winning solutions and highlight the current best practices in large-scale graph ML. Finally, we describe how we have updated the datasets after the KDD Cup to further facilitate research advances. The OGB-LSC datasets, baseline code, and all the information about the KDD Cup are available at https://ogb.stanford.edu/docs/lsc/ .
Forward citations
Cited by 14 Pith papers
-
GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond
GFFMERGE formulates GNN force field merging as a convex embedding-alignment problem with an analytical solution, recovering near joint-training performance on MD17, MD22, LiPS20 and other benchmarks while delivering 5...
-
Fused Gromov-Wasserstein Distance with Feature Selection
Fused Gromov-Wasserstein distances are extended with feature selection via Lasso/Ridge regularization or simplex-constrained weights, yielding theoretical bounds, metric properties, and an alternating minimization algorithm.
-
The limits of bio-molecular modeling with large language models : a cross-scale evaluation
LLMs perform adequately on bio-molecular classification tasks but remain weak on regression, with hybrid architectures outperforming others on long sequences and fine-tuning hurting generalization.
-
DSBD: Dual-Aligned Structural Basis Distillation for Graph Domain Adaptation
DSBD distills a dual-aligned structural basis to adapt GNNs across graphs with structural distribution shifts, outperforming prior methods on benchmarks.
-
DisRFM: Polar Riemannian Flow Matching for Structure-Preserving Graph Domain Adaptation
DisRFM uses polar Riemannian flow matching on constant-curvature manifolds to align graph domains while preserving label-relevant topology via radial Wasserstein and angular confidence matching.
-
Cross-Resolution Semantic Learning for Graph Domain Adaptation
CReSL improves graph domain adaptation by learning cross-resolution source-to-target routing and grafting target representations toward source class prototypes.
-
GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs
GraspLLM extracts dataset-agnostic structural patterns via motif contrastive learning and aligns contextual subgraphs to LLM tokens, outperforming prior LLM-based methods on TAGs especially in zero-shot settings.
-
Agentic Fusion of Large Atomic and Language Models to Accelerate Superconductor Discovery
An agentic framework fusing large atomic and language models rediscovers 66 known superconductors and guides experimental verification of four new ones with transition temperatures from 2.5 K to 6.5 K.
-
Pretraining a Foundation Model for Small-Molecule Natural Products
NaFM is a pretrained foundation model for natural products using scaffold-focused contrastive learning and masked graph objectives that achieves SOTA on taxonomy classification, gene/microbial analysis, and virtual sc...
-
Handling Feature Heterogeneity with Learnable Graph Patches
Learnable graph patches enable domain-agnostic pre-training of graph models by decomposing heterogeneous graphs into transferable semantic units via patch encoders and aggregators.
-
Safe-Subspace Pseudo-Label Refinement for Source-Free Graph Domain Adaptation
S²PLR identifies a safe subspace for reliable pseudo-labels in source-free graph domain adaptation using semantic committee signals and structural contrastive verification, then applies noise-tolerant regularization t...
-
Graph Hierarchical Recurrence for Long-Range Generalization
GHR uses hierarchical recurrence on pooled graph abstractions to improve long-range dependency capture and out-of-range generalization while using far fewer parameters than existing models.
-
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
GA-S2S integrates T5 with RGAT to jointly process text and k-hop subgraph topology for knowledge graph link prediction, reporting up to 19% relative accuracy gain over seq2seq baselines on CoDEx.
-
AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation
AdvSynGNN uses multi-resolution structural synthesis, contrastive objectives, an adaptive transformer, and an adversarial propagation engine with residual label correction to improve node-level predictions on challeng...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.