Rewriting existing molecule annotations with LLMs to create a more diverse training set (LaChEBI-20) lets a small T5-based model beat larger molecule-language models on generation and captioning.
The ogbg-molbace dataset provides quantitative (IC50) and qualitative (bi- nary label) binding results for a set of inhibitors of human b-secretase 1 (BACE-1)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Automatic Annotation Augmentation Boosts Translation between Molecules and Natural Language
Rewriting existing molecule annotations with LLMs to create a more diverse training set (LaChEBI-20) lets a small T5-based model beat larger molecule-language models on generation and captioning.