Pith. sign in

REVIEW 3 cited by

DARWIN 1.5: Large Language Models as Materials Science Adapted Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.11970 v3 pith:F36R2QR7 submitted 2024-12-16 cs.CL

classification cs.CL
keywords materialsmaterialacrossdarwindescriptorslanguagescienceapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

    cond-mat.mtrl-sci 2025-10 conditional novelty 7.0 of 10

    A new CIF-editing benchmark shows LLMs succeed on simple structure edits but fail on most spatial transformations, especially rotations.

  2. ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    ChemPile is an open 75-billion-token, multimodal chemical dataset spanning education, papers, property tables, code, images, and reasoning traces, released for training chemical foundation models.

  3. TopoMAS: Large Language Model Driven Topological Materials Multiagent System

    cond-mat.mtrl-sci 2025-07 conditional novelty 4.0 of 10

    TopoMAS is a multi-agent LLM framework that automates retrieval, generation, and first-principles validation for topological materials, reporting 94.55% accuracy with a lightweight Qwen2.5-72B model.

Pith tools