Pith. sign in

REVIEW 5 cited by

Understanding Isomorphism Bias in Graph Data Sets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12091 v2 pith:OUIV3RAP submitted 2019-10-26 cs.LG cs.SIstat.ML

classification cs.LGcs.SIstat.ML
keywords graphdataisomorphismsetsbiasmodelsclassificationused
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years there has been a rapid increase in classification methods on graph structured data. Both in graph kernels and graph neural networks, one of the implicit assumptions of successful state-of-the-art models was that incorporating graph isomorphism features into the architecture leads to better empirical performance. However, as we discover in this work, commonly used data sets for graph classification have repeating instances which cause the problem of isomorphism bias, i.e. artificially increasing the accuracy of the models by memorizing target information from the training set. This prevents fair competition of the algorithms and raises a question of the validity of the obtained results. We analyze 54 data sets, previously extensively used for graph-related tasks, on the existence of isomorphism bias, give a set of recommendations to machine learning practitioners to properly set up their models, and open source new data sets for the future experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CryptGNN: Enabling Secure Inference for Graph Neural Networks

    cs.CR 2025-09 reject novelty 6.0 of 10

    CryptGNN lets a client run a third-party GNN model on sensitive graph data using secret-shared computations across multiple cloud servers, keeping both data and model private.

  2. Robust Graph Learning Against Adversarial Evasion Attacks via Prior-Free Diffusion-Based Structure Purification

    cs.LG 2025-02 conditional novelty 6.0 of 10

    DiffSP purifies attacked graphs by learning the clean graph distribution with a discrete diffusion model and reports consistent accuracy improvements over baselines on nine datasets and nine evasion attacks.

  3. Predicting Steady-State Behavior in Complex Networks with Graph Neural Networks

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A GCN/GAT framework predicts IPR values of the principal eigenvector and classifies networks into delocalized, weakly localized, and strongly localized states with roughly 95 percent accuracy on synthetic test networks.

  4. KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural Networks

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Swapping attentive GNN score mappings for a single-layer Kolmogorov-Arnold Network improves benchmark performance and, on a specially constructed input matrix, provably achieves zero maximum ranking error.

  5. Graph Counterfactual Explainable AI via Latent Space Traversal

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Counterfactual graph explanations are generated by gradient descent in the latent space of a permutation-equivariant graph VAE, steering the graph's encoding to the opposite class.

Pith tools