MolRecBench-Wild reveals that 18 existing OCSR models suffer severe performance drops on complex real-world academic molecular images compared with prior patent benchmarks.
Self-Referencing Embedded Strings (SELFIES): A 100% Robust Molecular String Representation
8 Pith papers cite this work, alongside 596 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
other 1polarities
unclear 1representative citing papers
SCPT creates similarity-constrained preference triplets from scaffolds to train LLMs as conditional molecular editors that improve properties while keeping scaffolds intact.
On chemistry SMILES with a fixed 165-token base, BPE and Unigram-LM produce near-disjoint vocabularies (Jaccard ≤0.161) and Unigram-LM emits 29–41% more tokens across 22 matched conditions.
Structured text representations like CML and MolJSON outperform SMILES variants on structural tasks while IUPAC dominates semantic tasks such as molecule retrieval across all tested LLMs.
UniField fuses discrete atomic graphs with continuous electron density fields via RBF guidance in an SE(3)-equivariant multimodal model, reporting new SOTA results on QM9-ED, QMugs-ED, and ED5-OE benchmarks with gains up to 37%.
PolyFusionAgent integrates a multimodal polymer foundation model with a literature-grounded AI agent for property prediction and inverse design of novel polymers.
Hyformer jointly models molecule generation and property prediction via alternating attention and joint pre-training, showing synergistic gains in conditional sampling, OOD prediction, and a drug design case for antimicrobial peptides.
Fine-tuned LLaMA 3 achieves regression performance on QM9 molecular properties and 28 materials properties from composition strings that rivals random forests but is 5-10x worse than specialized models using atomic coordinates.
citing papers explorer
-
MolRecBench-Wild: A Real-World Benchmark for Optical Chemical Structure Recognition
MolRecBench-Wild reveals that 18 existing OCSR models suffer severe performance drops on complex real-world academic molecular images compared with prior patent benchmarks.
-
Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models
SCPT creates similarity-constrained preference triplets from scaffolds to train LLMs as conditional molecular editors that improve properties while keeping scaffolds intact.
-
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES
On chemistry SMILES with a fixed 165-token base, BPE and Unigram-LM produce near-disjoint vocabularies (Jaccard ≤0.161) and Unigram-LM emits 29–41% more tokens across 22 matched conditions.
-
Rethinking Molecular Text Representations for LLMs: An Empirical Study
Structured text representations like CML and MolJSON outperform SMILES variants on structural tasks while IUPAC dominates semantic tasks such as molecule retrieval across all tested LLMs.
-
UniField: RBF-Guided Electron Density Fusion for Enhanced Molecular Representations
UniField fuses discrete atomic graphs with continuous electron density fields via RBF guidance in an SE(3)-equivariant multimodal model, reporting new SOTA results on QM9-ED, QMugs-ED, and ED5-OE benchmarks with gains up to 37%.
-
PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design
PolyFusionAgent integrates a multimodal polymer foundation model with a literature-grounded AI agent for property prediction and inverse design of novel polymers.
-
Synergistic Benefits of Joint Molecule Generation and Property Prediction
Hyformer jointly models molecule generation and property prediction via alternating attention and joint pre-training, showing synergistic gains in conditional sampling, OOD prediction, and a drug design case for antimicrobial peptides.
-
Regression with Large Language Models for Materials and Molecular Property Prediction
Fine-tuned LLaMA 3 achieves regression performance on QM9 molecular properties and 28 materials properties from composition strings that rivals random forests but is 5-10x worse than specialized models using atomic coordinates.