Pith. sign in

REVIEW 6 cited by

Curve Your Attention: Mixed-Curvature Transformers for Graph Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.04082 v1 pith:Z6IINXFO submitted 2023-09-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords graphtransformerscurvaturemodelnon-euclideanattentiongeometrylearn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world graphs naturally exhibit hierarchical or cyclical structures that are unfit for the typical Euclidean space. While there exist graph neural networks that leverage hyperbolic or spherical spaces to learn representations that embed such structures more accurately, these methods are confined under the message-passing paradigm, making the models vulnerable against side-effects such as oversmoothing and oversquashing. More recent work have proposed global attention-based graph Transformers that can easily model long-range interactions, but their extensions towards non-Euclidean geometry are yet unexplored. To bridge this gap, we propose Fully Product-Stereographic Transformer, a generalization of Transformers towards operating entirely on the product of constant curvature spaces. When combined with tokenized graph Transformers, our model can learn the curvature appropriate for the input graph in an end-to-end fashion, without the need of additional tuning on different curvature initializations. We also provide a kernelized approach to non-Euclidean attention, which enables our model to run in time and memory cost linear to the number of nodes and edges while respecting the underlying geometry. Experiments on graph reconstruction and node classification demonstrate the benefits of generalizing Transformers to the non-Euclidean domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Riemannian Geometry for Pre-trained Language Model Embeddings

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Aggregating per-token pullback metrics via the Fréchet mean on the SPD manifold outperforms Euclidean mean pooling for sentence classification, with most of the gain attributable to geometric aggregation rather than l...

  2. HyPCV-Former: Hyperbolic Spatio-Temporal Transformer for 3D Point Cloud Video Anomaly Detection

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    HyPCV-Former embeds point cloud video features in Lorentzian hyperbolic space and uses hyperbolic attention to improve video anomaly detection on two benchmarks.

  3. Spectro-Riemannian Graph Neural Networks

    cs.LG 2025-02 conditional novelty 6.0 of 10

    CUSP is a graph neural network that combines Ollivier-Ricci curvature with spectral filters on a product of hyperbolic, spherical, and Euclidean spaces, claiming SOTA results on eight node and link prediction benchmarks.

  4. GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts

    cs.LG 2024-12 conditional novelty 6.0 of 10

    GraphMoRE learns per-node mixtures of Riemannian expert embedding spaces, reducing distortion on graphs that mix tree-like, cycle-like, and flat substructures.

  5. Geometric Machine Learning on EEG Signals

    cs.LG 2025-02 reject novelty 5.0 of 10

    An EEG pipeline combining transformer-based denoising with graph Ricci flow and a GCN reports 0.97 accuracy for digit versus non-digit thought classification, but without baselines or code.

  6. Removing Neural Signal Artifacts with Autoencoder-Targeted Adversarial Transformers (AT-AT)

    cs.LG 2025-02 conditional novelty 5.0 of 10

    An autoencoder-gated adversarial transformer denoises EEG-EMG mixtures with reconstruction accuracy comparable to larger published models at a fraction of the model size.

Pith tools