REVIEW 3 cited by
Saturn: Sample-efficient Generative Molecular Design using Memory Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative molecular design for drug discovery has very recently achieved a wave of experimental validation, with language-based backbones being the most common architectures employed. The most important factor for downstream success is whether an in silico oracle is well correlated with the desired end-point. To this end, current methods use cheaper proxy oracles with higher throughput before evaluating the most promising subset with high-fidelity oracles. The ability to directly optimize high-fidelity oracles would greatly enhance generative design and be expected to improve hit rates. However, current models are not efficient enough to consider such a prospect, exemplifying the sample efficiency problem. In this work, we introduce Saturn, which leverages the Augmented Memory algorithm and demonstrates the first application of the Mamba architecture for generative molecular design. We elucidate how experience replay with data augmentation improves sample efficiency and how Mamba synergistically exploits this mechanism. Saturn outperforms 22 models on multi-parameter optimization tasks relevant to drug discovery and may possess sufficient sample efficiency to consider the prospect of directly optimizing high-fidelity oracles.
Forward citations
Cited by 3 Pith papers
-
Multi-Objective Molecular Generation with Frequency-Controlled Evolutionary Dynamics
SpectralMol evolves Fourier-coefficient matrices of molecules via NSGA-II to perform multi-objective optimization, showing competitive benchmark results and advantages on a ClpP docking task.
-
NovoMolGen: Rethinking Molecular Language Model Pretraining
A 1.5-billion-molecule pretrained transformer family, NovoMolGen, sets new state-of-the-art results in de novo and goal-directed molecule generation, and shows pretraining loss correlates only weakly with downstream g...
-
A Survey of Mamba
The paper consolidates existing research on Mamba models, their architecture variants, adaptations to different data modalities, and applications across domains.
Discussion (0). Sign in to comment.