Pith. sign in

REVIEW 1 cited by

Transformers are efficient hierarchical chemical graph learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01704 v1 pith:P5JSNQ5H submitted 2023-10-02 cs.LG

classification cs.LG
keywords graphtransformersapproachchemicalsubformercomputationallearningstructures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers, adapted from natural language processing, are emerging as a leading approach for graph representation learning. Contemporary graph transformers often treat nodes or edges as separate tokens. This approach leads to computational challenges for even moderately-sized graphs due to the quadratic scaling of self-attention complexity with token count. In this paper, we introduce SubFormer, a graph transformer that operates on subgraphs that aggregate information by a message-passing mechanism. This approach reduces the number of tokens and enhances learning long-range interactions. We demonstrate SubFormer on benchmarks for predicting molecular properties from chemical structures and show that it is competitive with state-of-the-art graph transformers at a fraction of the computational cost, with training times on the order of minutes on a consumer-grade graphics card. We interpret the attention weights in terms of chemical structures. We show that SubFormer exhibits limited over-smoothing and avoids over-squashing, which is prevalent in traditional graph neural networks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BSA: Ball Sparse Attention for Large-scale Geometries

    cs.LG 2025-06 conditional novelty 4.0 of 10

    BSA combines Native Sparse Attention with ball-tree neighborhoods to give transformers a global view of 3D point sets at lower compute cost.

Pith tools