MolTextNet is a 2.5 million molecule-text dataset whose GPT-4o-mini descriptions are grounded in ChEMBL35; downstream gains are reported but may be confounded by label leakage.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-bio.BM 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
MolTextNet is a 2.5 million molecule-text dataset whose GPT-4o-mini descriptions are grounded in ChEMBL35; downstream gains are reported but may be confounded by label leakage.