Pith. sign in

REVIEW 7 cited by

Pre-training Molecular Graph Representation with 3D Geometry

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07728 v2 pith:GAP6YEFT submitted 2021-10-07 cs.LG cs.CVeess.IVq-bio.QM

classification cs.LGcs.CVeess.IVq-bio.QM
keywords graphmoleculargraphmvpgeometriclearningrepresentationgeometryinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Molecular graph representation learning is a fundamental problem in modern drug and material discovery. Molecular graphs are typically modeled by their 2D topological structures, but it has been recently discovered that 3D geometric information plays a more vital role in predicting molecular functionalities. However, the lack of 3D information in real-world scenarios has significantly impeded the learning of geometric graph representation. To cope with this challenge, we propose the Graph Multi-View Pre-training (GraphMVP) framework where self-supervised learning (SSL) is performed by leveraging the correspondence and consistency between 2D topological structures and 3D geometric views. GraphMVP effectively learns a 2D molecular graph encoder that is enhanced by richer and more discriminative 3D geometry. We further provide theoretical insights to justify the effectiveness of GraphMVP. Finally, comprehensive experiments show that GraphMVP can consistently outperform existing graph SSL methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 154 citations worldwide. Full citation record

  1. ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density

    cs.LG 2026-08 conditional novelty 6.0 of 10

    ED-DiT pretrains a diffusion transformer on electron-density point clouds with a physical electron-number constraint, and the resulting encoder outperforms scratch models across six molecular tasks.

  2. Khan-GCL: Kolmogorov-Arnold Network Based Graph Contrastive Learning with Hard Negatives

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Khan-GCL combines KAN encoders with coefficient-based critical feature identification to generate hard negatives and reports state-of-the-art graph classification results.

  3. ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Corrupted, ambiguous context semantics, not the presence of [MASK] symbols, drive MLM accuracy loss; expanding each [MASK] into multiple modeled states mitigates this.

  4. Graph Generative Pre-trained Transformer

    cs.LG 2025-01 conditional novelty 6.0 of 10

    G2PT represents graphs as node-then-edge token sequences and learns them with GPT-style next-token prediction, matching or beating diffusion baselines on seven graph and molecule datasets.

  5. CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs

    q-bio.QM 2025-08 conditional novelty 5.0 of 10

    Cross-view prefix resampling, guided by the LLM's SMILES encoding, lets a Galactica-based model exploit molecular graphs and images at low context cost, improving captioning, IUPAC naming, and property prediction.

  6. Dual-Modality Representation Learning for Molecular Property Prediction

    cs.LG 2025-01 conditional novelty 4.0 of 10

    DMCA combines GAT and ChemBERTa embeddings via cross-attention and reports the best classification rank and second-best regression average on eight MoleculeNet datasets.

  7. Computational Protein Science in the Era of Large Language Models (LLMs)

    cs.CE 2025-01 conditional novelty 3.0 of 10

    A survey that categorizes protein language models by the knowledge they learn and reviews their applications, with no new experimental results.

Pith tools