A VAE-LSTM model generates valid SMILES molecules conditioned on gene expression profiles, reportedly outperforming prior omics-based generators on validity and Tanimoto similarity.
Hierarchical Structure Enhances the Convergence and Generalizability of Linear Molecular Representation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Language models demonstrate fundamental abilities in syntax, semantics, and reasoning, though their performance often depends significantly on the inputs they process. This study introduces TSIS (Simplified TSID) and its variants:TSISD (TSIS with Depth-First Search), TSISO (TSIS in Order), and TSISR (TSIS in Random), as integral components of the t-SMILES framework. These additions complete the framework's design, providing diverse approaches to molecular representation. Through comprehensive analysis and experiments employing deep generative models, including GPT, diffusion models, and reinforcement learning, the findings reveal that the hierarchical structure of t-SMILES is more straightforward to parse than initially anticipated. Furthermore, t-SMILES consistently outperforms other linear representations such as SMILES, SELFIES, and SAFE, demonstrating superior convergence speed and enhanced generalization capabilities.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
De Novo Generation of Hit-like Molecules from Gene Expression Profiles via Deep Learning
A VAE-LSTM model generates valid SMILES molecules conditioned on gene expression profiles, reportedly outperforming prior omics-based generators on validity and Tanimoto similarity.