MolSight injects molecular graph topology and image-derived SVG annotations into a vision-language model, reporting state-of-the-art results on image-to-SMILES, captioning, descriptor, and bioactivity tasks.
arXiv preprint arXiv:2406.14021(2024)
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
FARM adds atomic-level functional group annotations to create FG-enhanced SMILES and FG graphs, trains them with masked language modeling and GNNs plus contrastive alignment, and reports state-of-the-art results on 8 of 13 MoleculeNet tasks.
QUIET is a hierarchical RVQ-based graph tokenizer with a learned level-weighting gate; it improves several benchmarks but not consistently against the strongest baselines.
citing papers explorer
-
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
MolSight injects molecular graph topology and image-derived SVG annotations into a vision-language model, reporting state-of-the-art results on image-to-SMILES, captioning, descriptor, and bioactivity tasks.
-
FARM: Enhancing Molecular Representations with Functional Group Awareness
FARM adds atomic-level functional group annotations to create FG-enhanced SMILES and FG graphs, trains them with masked language modeling and GNNs plus contrastive alignment, and reports state-of-the-art results on 8 of 13 MoleculeNet tasks.
-
A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
QUIET is a hierarchical RVQ-based graph tokenizer with a learned level-weighting gate; it improves several benchmarks but not consistently against the strongest baselines.