Pith. sign in

REVIEW 6 cited by

ChemMLLM: Chemical Multimodal Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16326 v2 pith:WF5A5HOQ submitted 2025-05-22 cs.LG

classification cs.LG
keywords chemmllmchemicalmultimodallanguagelargemllmstasksacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified chemical multimodal large language model for molecule understanding and generation. Also, we design five multimodal tasks across text, molecular SMILES strings, and image, and curate the datasets. We benchmark ChemMLLM against a range of general leading MLLMs and Chemical LLMs on these tasks. Experimental results show that ChemMLLM achieves superior performance across all evaluated tasks. For example, in molecule image optimization task, ChemMLLM outperforms the best baseline (GPT-4o) by 116.75\% (4.27 vs 1.97 property improvement). The code is publicly available at https://github.com/bbsbz/ChemMLLM.git.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MolRecBench-Wild: A Real-World Benchmark for Optical Chemical Structure Recognition

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    MolRecBench-Wild reveals that 18 existing OCSR models suffer severe performance drops on complex real-world academic molecular images compared with prior patent benchmarks.

  2. MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MolSight injects molecular graph topology and image-derived SVG annotations into a vision-language model, reporting state-of-the-art results on image-to-SMILES, captioning, descriptor, and bioactivity tasks.

  3. LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

    physics.chem-ph 2026-02 conditional novelty 6.0 of 10

    LatentChem reasons in continuous latent space for chemistry, achieving a 59.88% non-tie win rate over explicit CoT on ChemCoTBench with a 10.84x average reduction in reasoning overhead.

  4. Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning

    cs.CE 2025-07 conditional novelty 6.0 of 10

    RetroDFM-R, a ChemDFM-based LLM trained with reasoning distillation and reinforcement learning, reaches 65.0% top-1 retrosynthesis accuracy on USPTO-50K.

  5. ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge

    cs.CE 2025-07 unverdicted novelty 5.0 of 10

    ChemDFM-R is a chemical reasoning LLM trained via a four-stage pipeline on the ChemFG dataset of functional-group annotations for molecules and reactions, reaching performance comparable to or better than commercial m...

  6. MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

    cs.CV 2026-07 unverdicted novelty 4.0 of 10

    MolSight integrates a Molecular Topology Module and Molecular Grounding Module into VLMs to enhance molecular image understanding and claims to outperform prior models on chemical visual tasks.

Pith tools