Pith. sign in

REVIEW 2 cited by

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.02810 v2 pith:VLMG63SJ submitted 2025-02-05 cs.LG cs.AIphysics.chem-phq-bio.BM

classification cs.LGcs.AIphysics.chem-phq-bio.BM
keywords moleculargeneralistgraphllmsmultimodalpredictioninformationmol-llm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in large language models (LLMs) have led to models that tackle diverse molecular tasks, such as chemical reaction prediction and molecular property prediction. Large-scale molecular instruction-tuning datasets have enabled sequence-only (e.g., SMILES or SELFIES) generalist molecular LLMs, and researchers are now exploring multimodal approaches that incorporate molecular structural information for further gains. However, a genuinely multimodal, generalist LLM that covers a broad spectrum of molecular tasks has yet to be fully investigated. We observe that naive next token prediction training ignores graph-structural information, limiting an LLM's ability to exploit molecular graphs. To address this, we propose (i) Molecular structure Preference Optimization (MolPO), which facilitates graph usage by optimizing preferences between pairs of correct and perturbed molecular structures, and (ii) an advanced graph encoder with a tailored pre-training strategy to improve the effect of graph utilization by MolPO. Building on these contributions, we introduce Mol-LLM, the first multimodal generalist model that (a) handles a broad spectrum of molecular tasks among molecular LLMs, (b) explicitly leverages molecular-structure information, and (c) takes advantage of extensive instruction tuning. Mol-LLM attains state-of-the-art or comparable results across the most comprehensive molecular-LLM benchmark-even on out-of-distribution datasets for reaction and property prediction, where it surpasses prior generalist molecular LLMs by a large margin.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?

    cs.AI 2025-06 conditional novelty 7.0 of 10

    A new benchmark called ToxiMol evaluates how well 43 multimodal LLMs can edit toxic molecules into structurally similar, non-toxic, drug-like candidates; the best model succeeds on 43.3% of tasks.

  2. DrugGen 2: A disease-aware language model for enhancing drug discovery

    q-bio.QM 2026-07 conditional novelty 6.0 of 10

    A GPT-2 model fine-tuned with disease MeSH + protein sequence inputs and GRPO rewards produces more unique, valid, drug-like, high-PLAPT-affinity ligands than DrugGPT or DrugGen on five diabetic-nephropathy targets.

Pith tools