Pith. sign in

REVIEW 20 cited by

LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09391 v4 pith:X5GEEHNO submitted 2024-02-14 cs.AI cs.CEcs.CL

LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

classification cs.AI cs.CEcs.CL
keywords chemistrytasksllmscomprehensivedatasetlanguagegpt-4high-quality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Chemistry plays a crucial role in many domains, such as drug discovery and material science. While large language models (LLMs) such as GPT-4 exhibit remarkable capabilities on natural language processing tasks, existing research indicates that their performance on chemistry tasks is discouragingly low. In this paper, however, we demonstrate that our developed LLMs can achieve very strong results on a comprehensive set of chemistry tasks, outperforming the most advanced GPT-4 and Claude 3 Opus by a substantial margin. To accomplish this, we propose SMolInstruct, a large-scale, comprehensive, and high-quality dataset for instruction tuning. It contains 14 selected chemistry tasks and over three million samples, laying a solid foundation for training and evaluating LLMs for chemistry. Using SMolInstruct, we fine-tune a set of open-source LLMs, among which, we find that Mistral serves as the best base model for chemistry tasks. Our analysis further demonstrates the critical role of the proposed dataset in driving the performance improvements.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SupraBench: A Benchmark for Supramolecular Chemistry

    cs.LG 2026-06 unverdicted novelty 7.0

    SupraBench introduces four core tasks and a curated corpus to benchmark LLMs on host-guest chemistry reasoning, showing substantial remaining headroom and task-specific failure modes.

  2. Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression

    cs.LG 2026-05 unverdicted novelty 7.0

    Distribution-Aware Reward optimizes LLM regression by treating rollouts as empirical predictive distributions and rewarding marginal improvements in CRPS quality rather than point accuracy alone.

  3. FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization

    cs.LG 2026-05 unverdicted novelty 7.0

    FORGE reformulates molecular optimization as context-aware fragment ranking and replacement using mined low-to-high edit pairs, outperforming larger language models and graph methods on standard benchmarks.

  4. Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

    cs.AI 2026-05 unverdicted novelty 7.0

    LLM agents reach only 50.6% accuracy on chemical cost estimation within 25% error even with tools, dropping with noise due to parsing, pack selection, and tool-use failures.

  5. The limits of bio-molecular modeling with large language models : a cross-scale evaluation

    cs.LG 2026-04 unverdicted novelty 7.0

    LLMs perform adequately on bio-molecular classification tasks but remain weak on regression, with hybrid architectures outperforming others on long sequences and fine-tuning hurting generalization.

  6. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.5

    A 52.6B-token multi-domain biology pretraining corpus with tool enrichment and new binding/localization instructions doubles a fixed base LLM's matched biology-eval score with little language forgetting.

  7. MemSFT: Mitigating Alignment Tax with an External Parametric Memory

    cs.LG 2026-07 conditional novelty 6.0

    MemSFT attaches a retriever-imitating 8B memory plus a word-level router to frozen Qwen3 backbones, boosting domain scores by ~36 points while holding general-benchmark averages essentially flat, where full SFT loses ...

  8. Monkey King Bang: A Unified Scientific Multimodal Foundation Model

    cs.LG 2026-07 conditional novelty 6.0

    A shared-backbone multimodal model integrating six scientific modalities with native-output decoders shows competitive understanding and generation across biology, chemistry, weather, and medical imaging.

  9. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.0

    TheBioCollection, a 52.6B-token unified biology corpus with tool-computed text and new instruction tasks, raises a fixed 16B LLM's score on its matched biology eval from 0.223 to 0.499 (2.24×).

  10. MolDA: Molecular Understanding and Generation via Large Language Diffusion Model

    cs.AI 2026-04 unverdicted novelty 6.0

    MolDA is a multimodal molecular model that uses a discrete large language diffusion backbone plus a hybrid graph encoder to achieve better global coherence and validity than autoregressive approaches.

  11. MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

    cs.LG 2026-06 unverdicted novelty 5.0

    MolE-RAG is a training-free RAG framework that augments LLMs with literature, molecular context, and structural analogs to improve performance on nine molecular property prediction tasks.

  12. SciCore-Mol: Augmenting Large Language Models with Pluggable Molecular Cognition Modules

    cs.AI 2026-05 unverdicted novelty 5.0

    SciCore-Mol augments LLMs with three integrated modules for molecular perception, latent diffusion generation, and reaction reasoning, claiming an 8B open model competes with or exceeds proprietary systems on chemical tasks.

  13. Bolek: A Multimodal Language Model for Molecular Reasoning

    cs.LG 2026-05 unverdicted novelty 5.0

    Bolek injects Morgan fingerprint embeddings into an instruction-tuned text model, then fine-tunes on molecular alignment and synthetic chain-of-thought tasks to improve performance and grounding on 15 TDC binary class...

  14. Heterogeneous Scientific Foundation Model Collaboration

    cs.AI 2026-04 unverdicted novelty 5.0

    Eywa enables language-based agentic AI systems to collaborate with specialized scientific foundation models for improved performance on structured data tasks.

  15. ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge

    cs.CE 2025-07 unverdicted novelty 5.0

    ChemDFM-R is a chemical reasoning LLM trained via a four-stage pipeline on the ChemFG dataset of functional-group annotations for molecules and reactions, reaching performance comparable to or better than commercial m...

  16. Molecular Lead Optimization via Agentic Tool Planning

    cs.LG 2026-05 unverdicted novelty 4.0

    TRACE is a trajectory-aware LLM agent that treats molecular tool selection as sequential decision-making to achieve higher success rates and larger ADMET improvements than one-step baselines on optimization tasks.

  17. OpenCompass: A Universal Evaluation Platform for Large Language Models

    cs.CL 2026-05 conditional novelty 4.0

    OpenCompass is a modular, high-concurrency platform for unified LLM evaluation across knowledge, reasoning, code, and other domains with support for rule-based, LLM-as-judge, and cascaded evaluators.

  18. Regression with Large Language Models for Materials and Molecular Property Prediction

    cond-mat.mtrl-sci 2024-09 unverdicted novelty 4.0

    Fine-tuned LLaMA 3 achieves regression performance on QM9 molecular properties and 28 materials properties from composition strings that rivals random forests but is 5-10x worse than specialized models using atomic co...

  19. SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration

    physics.chem-ph 2024-09 unverdicted novelty 4.0

    SmileyLlama is an LLM transformed via SFT and DPO to generate valid novel drug-like molecules with user-specified properties and optimized 3D conformations for high binding affinity.

  20. OpenCompass: A Universal Evaluation Platform for Large Language Models

    cs.CL 2026-05 unverdicted novelty 3.0

    OpenCompass is presented as a one-stop, scalable, high-concurrency LLM evaluation platform with modular architecture supporting multiple domains and evaluator types.