Pith. sign in

REVIEW 1 cited by

Can Large Language Models Understand Molecules?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00024 v3 pith:6Y2ZOC66 submitted 2024-01-05 q-bio.BM cs.AIcs.CLcs.LG

classification q-bio.BMcs.AIcs.CLcs.LG
keywords smilesmolecularmodelsllmspredictionpre-trainedtasksembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of cheminformatics, particularly in understanding Simplified Molecular Input Line Entry System (SMILES), a standard method for representing chemical structures. These LLMs also have the ability to decode SMILES strings into vector representations. Method: We investigate the performance of GPT and LLaMA compared to pre-trained models on SMILES in embedding SMILES strings on downstream tasks, focusing on two key applications: molecular property prediction and drug-drug interaction prediction. Results: We find that SMILES embeddings generated using LLaMA outperform those from GPT in both molecular property and DDI prediction tasks. Notably, LLaMA-based SMILES embeddings show results comparable to pre-trained models on SMILES in molecular prediction tasks and outperform the pre-trained models for the DDI prediction tasks. Conclusion: The performance of LLMs in generating SMILES embeddings shows great potential for further investigation of these models for molecular embedding. We hope our study bridges the gap between LLMs and molecular embedding, motivating additional research into the potential of LLMs in the molecular representation field. GitHub: https://github.com/sshaghayeghs/LLaMA-VS-GPT

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging neural network interatomic potentials for a foundation model of chemistry

    cond-mat.mtrl-sci 2025-06 conditional novelty 5.0 of 10

    Using embeddings from a pretrained neural network interatomic potential as features for small machine learning models gives competitive or better property predictions than end-to-end deep networks, especially with lim...

Pith tools