REVIEW 3 major objections 6 minor 165 references
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper introduces ToxiMol, the first benchmark for molecular toxicity repair, and shows that the best of 43 multimodal language models succeeds on only 43.3% of repair attempts.
desk verdict ToxiMol is a genuine first benchmark for structure-level molecular toxicity repair, and the resource release is solid, but the headline results lean on a single learned oracle whose own sensitivity is visible in the paper's ablations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three components carry the argument. ToxiMol is the benchmark: 660 toxic molecules sampled from the Therapeutics Data Commons toxicity tasks, balanced at 60 per endpoint across 11 primary tasks (with 12 Tox21 and 10 ToxCast subtasks evaluated separately), selected via ECFP4 fingerprints, Tanimoto similarity, and Butina clustering so each task covers chemically diverse scaffolds. ToxiEval is the evaluation chain: an RDKit validity check first, then a strict conjunctive gate that requires the TxGemma-Predict safety oracle to call the molecule non-toxic, QED ≥ 0.5, SAS ≤ 6, Lipinski violations ≤ 1, and Tanimoto similarity ≥ 0.4 to the original; a sample is repaired only if any one of the three candidates clears all five hurdles. The mechanism-aware prompt annotation pipeline is the third piece: a base template that casts the model as a medicinal chemistry expert is layered with task-level mechanism annotations and, for Tox21 and ToxCast, subtask-specific instructions such as 'reduce AhR agonism by altering molecular coplanarity,' yielding a unique multimodal prompt per molecule that ties the repair goal to the specific toxic mechanism.
What would settle it
Collect the molecules that the best model, Claude Opus 4.5, 'successfully repaired' — roughly 43% of the 660 — and re-judge them independently: run them through actual Ames or hERG assays, or at minimum re-score them with a structurally different toxicity predictor that the paper itself lists (ProTox, ADMETlab, or pkCSM). If the pass rate under the second oracle or the wet-lab assay collapses, or if the per-task ordering (LD50 hardest, hERG_Central easiest) flips, then the benchmark's success signal is an artifact of the single oracle rather than a measure of detoxification capability.
Extended reading notes
Core claim
The paper's claim is that molecular toxicity repair is a real, distinct capability — different from predicting toxicity or optimizing a single property like solubility — and that it can be measured for the first time with a fixed benchmark. The task is defined concretely: given a toxic molecule's SMILES string, its 2D structure image, and a mechanism-aware prompt, the model must return three candidate SMILES, and the sample is repaired if at least one candidate passes every stage of ToxiEval — RDKit structural validity, a non-toxic verdict from the TxGemma-Predict oracle, QED ≥ 0.5, SAS ≤ 6, at most one Lipinski violation, and at least 0.4 Tanimoto similarity to the original. On this yardstick, 43 evaluated MLLMs cluster between roughly 2% and 43% overall success, with the strongest result held by Claude Opus 4.5; closed-source models average 32.9%, open-source models 22.7%, and chemistry-specific models 25.4%. The task-level pattern is stable across families: hERG-channel, Tox21, and ToxCast repairs are comparatively tractable while LD50 and DILI are nearly always failures, and failure attribution shows that some endpoints bottleneck on toxicity while others bottleneck on drug-likeness. The authors' conclusion is measured: current models exhibit preliminary capability in toxicity understanding, semantic constraint adherence, and structure-aware editing, but molecular toxicity repair as a dependable service is not yet within reach.
Load-bearing premise
A repair is declared successful only when a single AI model, TxGemma-Predict, judges the new molecule non-toxic; if that learned oracle's verdict is wrong or biased for these particular molecules, then every reported success rate, per-task difficulty ranking, and model comparison in the paper measures the oracle instead of real detoxification.
Editorial extensions
If this is right
- Toxicity repair becomes a standard reproducible test: any future multimodal model can be scored on the same 660-molecule benchmark, so claims about molecular understanding can be checked against a fixed yardstick instead of anecdotes.
- No current model is trustworthy enough for autonomous drug redesign — the best succeeds on under half of molecules, with LD50 and DILI at near zero — so the realistic deployment is model proposes, chemist disposes.
- Repair difficulty is endpoint-specific, not generic: LD50, DILI, and hERG failures are toxicity bottlenecks while Tox21 and ToxCast failures are drug-likeness bottlenecks, implying that a single 'make it less toxic' instruction will not work and prompts must be mechanism-aware.
- Sampling more candidates from the same model saturates quickly — success rises from roughly 25% at one candidate to 35% at three or four, then plateaus — so the binding constraint is per-attempt edit quality, not candidate diversity.
- Adding the 2D structure image yields a modest but consistent gain (for GPT-4.1, 30.5% with image versus 28.9% with SMILES only), so visual perception contributes to repair skill but is still under-exploited by current models.
Reading between the lines
- The entire leaderboard rests on a single learned oracle: if TxGemma-Predict is systematically lenient on some chemistries and strict on others, the per-model and per-task ordering — including the sharp LD50-versus-hERG contrast — could be an artifact. A direct robustness test is to re-score the same candidate pools with an independent predictor (the paper itself lists ProTox, ADMETlab, and pkCSM a
- Because a sample counts as repaired when any one of three candidates passes, the saturation of success near three or four candidates implies each model's output distribution has a thin 'good tail' that sampling quickly exhausts; this points toward training models to produce several distinct plausible edits per molecule rather than boosting test-time sampling.
- The near-total LD50 failure suggests the gap is continuous dose-response reasoning rather than chemical knowledge. A testable extension is to give the model the target log(LD50) range explicitly and ask for a molecule predicted inside that window, then measure whether success moves off zero.
- The evaluation machinery generalizes beyond the current scope: the paper names macromolecular therapeutics as a future direction, and its Type-T/Type-O failure attribution already provides a template for diagnosing where any structure-editing system — for toxicity or for other properties — fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper defines molecular toxicity repair as a benchmark task for general-purpose multimodal large language models (MLLMs). It introduces ToxiMol, a dataset of 660 toxic molecules sampled in a structure-aware, task-balanced way from 11 TDC toxicity tasks, together with a mechanism-aware prompt annotation pipeline. The paper also proposes ToxiEval, an automated evaluation chain in which a repaired molecule is deemed successful only if it passes TxGemma-Predict-based safety scoring, QED, SAS, RO5, and Tanimoto similarity thresholds. The authors evaluate 43 MLLMs, reporting low absolute success rates (best overall 43.3% for Claude Opus 4.5) and presenting ablations on structural validity, candidate count, metric combinations, multimodal input, oracle scale, and failure attribution. The central claim is that current MLLMs, while far from reliable on this task, show measurable capability in toxicity understanding, semantic constraint adherence, and structure-aware editing.
Significance. The paper is a serious and potentially useful benchmark contribution. Its strengths include cluster-based sampling with UMAP coverage checks, balanced task construction, a transparent multi-criteria evaluation protocol, evaluation across 43 models, public code and data releases, and a useful failure-mode analysis. The multi-round robustness check in Appendix C.3 and the oracle-scale study in Appendix C.6 are honest and informative, and the paper explicitly acknowledges the surrogate nature of the toxicity oracle and the risk of data leakage. If the safety oracle can be validated on the repaired-molecule distribution and the headline results are accompanied by uncertainty estimates, ToxiMol would be a valuable community resource for studying MLLM-driven molecular editing. As it stands, however, the central quantitative claims rest on an unvalidated single-oracle success definition and on single-run model comparisons, so the reported rates cannot yet be interpreted as established detoxification capability.
major comments (3)
- [Section 3.3, Appendix C.6 (Table 10)] The success label in ToxiEval is determined entirely by TxGemma-Predict's safety score, and Appendix C.6 shows that this choice is load-bearing: re-scoring the same Claude 3.7 Sonnet candidates with TxGemma-2B instead of TxGemma-27B changes the overall success rate from 43.3% to 35.5%, with the largest per-task shifts on hERG (10.0 to 35.0) and SkinRxn (6.7 to 33.3). Because the paper's central claim of 'promising capability' is expressed through these absolute success rates, the benchmark currently measures a conjunction of MLLM repair ability and oracle behavior. The paper acknowledges the surrogate nature of the predictor in Section 6 and Appendix A, but it does not validate the oracle on the repaired-molecule distribution or against independent toxicity predictors. I ask for a concrete validation study: sample successful and failed candidates, obtain labels from at least one independent predictor (for example, ADMETlab 2.0, ProTox-II, or pkCSM) or from expert toxicological review, and report per-endpoint agreement with TxGemma. If agreement is low, report success rates under an ensemble or majority-vote oracle. Without this, the headline numbers are not yet interpretable as detoxification capability.
- [Section 4.1, Table 1, Appendix C.3] Table 1 reports single-run success rates for all 43 models even though the MLLMs are stochastic and are sampled at temperature 0.7 (Section 4.1). Appendix C.3 provides five-run means and standard deviations for only four models; those runs show task-level standard deviations up to about 4.7 percentage points (for example, GPT-4.1 on hERG_Central), even though overall rates are stable. Several model-level claims in Section 4.2 involve differences of only a few points (for example, GPT-5.2 at 34.7% versus GPT-o3 at 39.7%, or ChemVLM-8B at 20.8% versus ChemVLM-26B at 15.3%). Without multi-run estimates or confidence intervals for the headline table, these differences and the associated conclusions about reasoning variants, scale effects, and model-family comparisons are not statistically supported. Please either report mean and standard deviation for all models (or for a representative subset covering each claim), or explicitly label these comparisons as exploratory.
- [Section 4.3, Table 2, Table 15] The failure-mode attribution table is internally inconsistent for DILI. Table 2 reports Type-T = 41.7% for Claude-3.7 Sonnet on DILI, but Table 15 reports a validity rate of only 18.3% for the same model and task. Since Type-T is defined as an RDKit-valid candidate that fails the safety threshold, the Type-T percentage cannot exceed the validity rate. This suggests either that the validity rates and Type-T rates use different denominators, or that the DILI row is miscalculated. The conclusion that DILI is a 'strong toxicity bottleneck' rather than a validity bottleneck is load-bearing for the task-level analysis in Section 4.2, so please clarify the denominator and recompute the attribution.
minor comments (6)
- [Appendix C.6, Table 10] The comparison across TxGemma-Predict scales is confounded by numerical precision: the 27B variant is evaluated in float32 while the 2B and 9B variants use float16. Please report whether this affects the conclusions, or run all scales in the same precision.
- [Table 15] The header of Table 15 says 'Average Success Rate' but the table reports validity rates. Please rename this to 'Average Valid Rate' or similar.
- [Figure 1] The dual-scale radar plot with the light blue inset is visually complex. A separate panel with a shared axis or a small-multiples plot would make the task-level comparisons easier to read.
- [Section 4.1] The text states that temperature is set to 0.7 uniformly, but also says 'default settings are used when this is not configurable.' Please list which models did not honor the temperature setting, since this affects reproducibility.
- [Appendix I] The data-leakage caveat is an important interpretive boundary for the benchmark, but it appears only in an appendix and the limitations section. Please state this caveat prominently in the main text or abstract, because the paper otherwise reads as a generalization assessment.
- [Appendix J] The TxGemma citation in Appendix J is dated 2024, while reference [126] is dated 2025. Please correct this inconsistency.
Circularity Check
No significant circularity: ToxiEval uses an external, fixed oracle (TxGemma-Predict) and performs no parameter fitting, so the reported success rates do not reduce to the benchmark's own inputs.
full rationale
The central derivation chain is: define repair success via ToxiEval (Section 3.3 and Algorithm 2), which requires RDKit-valid structure, a TxGemma-Predict safety score, and thresholds on QED, SAS, RO5, and Tanimoto similarity. TxGemma-Predict is an external, frozen model selected for reproducibility and coverage in Appendix A; no parameter of ToxiEval is fitted to the outputs of the 43 MLLMs, and no equation in the paper defines an output in terms of an input it is supposed to predict. The headline claim that MLLMs show promising capability in toxicity repair is therefore a measurement result under a stated, transparent evaluation coordinate system, not a prediction derived from its own inputs. The dataset and the oracle share the TDC source, and the paper explicitly acknowledges potential data leakage in Appendix I and disclaims equivalence with real-world toxicological correctness in the Limitations section; this is an acknowledged validity limitation, not a circular reduction. The oracle-scale ablation in Appendix C.6 shows sensitivity of success rates to the choice of TxGemma variant, which is a robustness concern about the metric, not evidence that the success definition is circular. No load-bearing self-citation was found: the authors' own prior publications appear only as general background in the related-work discussion, and no uniqueness theorem or prior result by the same authors is invoked to force the benchmark design. Consequently, the paper is self-contained as an evaluation benchmark, and its central empirical claims rest on external evidence rather than on its own assumptions.
Assumptions & free parameters
free parameters (8)
- QED threshold =
0.5
- SAS threshold =
6
- RO5 violations threshold =
1
- Tanimoto similarity threshold =
0.4
- Safety score threshold =
label A or normalized >0.5
- Number of candidates =
3
- Sampling size per task =
60
- Sampling temperature =
0.7
assumptions (4)
- domain assumption TxGemma-Predict provides a valid and standardized toxicity oracle for all 11 endpoints.
- domain assumption TDC toxicity labels are reliable ground truth for selecting toxic molecules.
- domain assumption ECFP4 fingerprints and Tanimoto similarity capture meaningful structural similarity for scaffold preservation.
- standard math RDKit parsing is a sufficient structural validity check.
Cite this review
Pith. "Pith review of Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?." pith.science (2026). https://pith.science/paper/SKMASNE3
@misc{pith2026250610912,
author = {Pith},
title = {Pith review of: Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?},
year = {2026},
howpublished = {\url{https://pith.science/paper/SKMASNE3}},
note = {Machine review of arXiv:2506.10912}
}
read the original abstract
Toxicity remains a leading cause of early-stage drug development failure. Despite advances in molecular design and property prediction, the task of molecular toxicity repair, generating structurally valid molecular alternatives with reduced toxicity, has not yet been systematically defined or benchmarked. To fill this gap, we introduce ToxiMol, the first benchmark task for general-purpose Multimodal Large Language Models (MLLMs) focused on molecular toxicity repair. We construct a standardized dataset covering 11 primary tasks and 660 representative toxic molecules spanning diverse mechanisms and granularities. We design a prompt annotation pipeline with mechanism-aware and task-adaptive capabilities, informed by expert toxicological knowledge. In parallel, we propose an automated evaluation framework, ToxiEval, which integrates toxicity endpoint prediction, synthetic accessibility, drug-likeness, and structural similarity into a high-throughput evaluation chain for repair success. We systematically assess 43 mainstream general-purpose MLLMs and conduct multiple ablation studies to analyze key issues, including evaluation metrics, candidate diversity, and failure attribution. Experimental results show that although current MLLMs still face significant challenges on this task, they begin to demonstrate promising capabilities in toxicity understanding, semantic constraint adherence, and structure-aware editing.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
01.AI, Alex Young, Bei Chen, et al. 2024. Yi: Open Foundation Models by 01.AI. arXiv:2403.04652
arXiv 2024
-
[2]
Raghad J AbuNasser, Mostafa Z Ali, Yaser Jararweh, Mustafa Daraghmeh, and Talal Z Ali. 2024. Large Language Models in Drug Discovery: A Comprehensive Analysis of Drug-Target Interaction Prediction. In2024 2nd International Con- ference on Foundation and Large Language Models (FLLM). IEEE, Dubai, United Arab Emirates, 417–431
2024
-
[3]
Nicholas Aksamit, Alain Tchagang, Yifeng Li, and Beatrice Ombuki-Berman
-
[4]
Vinicius M Alves, Eugene Muratov, Denis Fourches, Judy Strickland, Nicole Kleinstreuer, Carolina H Andrade, and Alexander Tropsha. 2015. Predicting chemically-induced skin reactions. Part I: QSAR models of skin sensitization and their application to identify potentially hazardous compounds.Toxicology and applied pharmacology284, 2 (2015), 262–272
2015
-
[5]
Anthropic. 2024. Claude 3.7 Sonnet and Claude Code. https://www.anthropic. com/news/claude-3-7-sonnet. Accessed: 2025-05-05
2024
-
[6]
Anthropic. 2025. Introducing Claude Opus 4.5. https://www.anthropic.com/cl aude/opus. Accessed: 2026-01-06
2025
-
[7]
Anthropic. 2025. Introducing Claude Sonnet 4.5. https://www.anthropic.com/ news/claude-sonnet-4-5. Accessed: 2026-01-06
2025
-
[8]
Reza Averly, Frazier N. Baker, Ian A. Watson, and Xia Ning. 2025. LIDDIA: Language-based Intelligent Drug Discovery Agent. arXiv:2502.13959 [cs.CL] https://arxiv.org/abs/2502.13959
arXiv 2025
Show all 165 references
-
[9]
Changsen Bai, Lianlian Wu, Ruijiang Li, Yang Cao, Song He, and Xiaochen Bo
-
[10]
Jinze Bai and Shuai Bai et al. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
2023
-
[11]
Lei Bai, Zhongrui Cai, Yuhang Cao, Maosong Cao, Weihan Cao, Chiyu Chen, Haojiong Chen, Kai Chen, Pengcheng Chen, Ying Chen, et al. 2025. Intern-S1: A Scientific Multimodal Foundation Model. arXiv:2508.15763 [cs.LG] https: //arxiv.org/abs/2508.15763
2025
-
[12]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. 2025. Qwen3-VL Technical Report. arXiv:2511.21631 [cs.CV] https://arxiv.org/abs/2511.21631
2025 arXiv
-
[13]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923 [cs.CV] https://arxiv.org/abs/2502.13923
2025 arXiv
-
[14]
Priyanka Banerjee, Andreas O Eckert, Anna K Schrey, and Robert Preissner
-
[15]
Anna O Basile, Alexandre Yahi, and Nicholas P Tatonetti. 2019. Artificial intelligence for drug toxicity and safety.Trends in pharmacological sciences40, 9 (2019), 624–635
2019
-
[16]
Romualdo Benigni and Cecilia Bossa. 2008. Predictivity of QSAR.Journal of chemical information and modeling48, 5 (2008), 971–980
2008
-
[17]
Volker Bergen, Konstantia Kodella, Sreenath Srikrishnan, Ornella Barrandon, Sara Anderson, Max Rogers-Grazado, Casey Fowler, Hirit Beyene, Nicole Ro- bichaud, Timothy Fulton, et al . 2025. A large-scale human toxicogenomics resource for drug-induced liver injury prediction.Nat...
2025
-
[18]
G Richard Bickerton, Gaia V Paolini, Jérémy Besnard, Sorel Muresan, and Andrew L Hopkins. 2012. Quantifying the chemical beauty of drugs.Nature chemistry4, 2 (2012), 90–98
2012
-
[19]
Pietro Bongini, Monica Bianchini, and Franco Scarselli. 2021. Molecular gener- ative graph neural networks for drug discovery.Neurocomputing450 (2021), 242–252
2021
-
[20]
Rodolpho C Braga, Vinicius M Alves, Meryck FB Silva, Eugene Muratov, Denis Fourches, Luciano M Lião, Alexander Tropsha, and Carolina H Andrade. 2015. Pred-hERG: A novel web-accessible computational tool for predicting cardiac toxicity.Molecular informatics34, 10 (2015), 698–701
2015
-
[21]
Nathan Brown, Marco Fiscato, Marwin HS Segler, and Alain C Vaucher. 2019. GuacaMol: benchmarking models for de novo molecular design.Journal of chemical information and modeling59, 3 (2019), 1096–1108
2019
-
[22]
Krishna C Bulusu, Rajarshi Guha, Daniel J Mason, Richard PI Lewis, Eugene Muratov, Yasaman Kalantar Motamedi, Murat Cokol, and Andreas Bender. 2016. Modelling of compound combination effects and applications to efficacy and toxicity: state-of-the-art, challenges and perspectiv...
2016
-
[23]
Darko Butina. 1999. Unsupervised data base clustering based on daylight’s fingerprint and Tanimoto similarity: A fast and automated way to cluster small and large data sets.Journal of Chemical Information and Computer Sciences39, 4 (1999), 747–750
1999
-
[24]
ByteDance Seed Team. 2024. Doubao Seed 1.6 Vision Model Introduction. https://www.volcengine.com/docs/82379/1320672. Official Volcano Engine documentation for doubao-seed-1.6-vision. Accessed: 2026-01-06
2024
-
[25]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. 2024. InternLM2 Technical Report. arXiv:2403.17297 [cs.CL] https://arxiv.org/abs/2403.17297
2024 arXiv
-
[26]
Daniel Campos and Heng Ji. 2021. IMG2SMI: Translating Molecular Structure Images to Simplified Molecular-input Line-entry System. arXiv:2109.04202
2021 arXiv
-
[27]
He Cao, Zijing Liu, Xingyu Lu, Yuan Yao, and Yu Li. 2025. InstructMol: Multi- Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery. InProceedings of the 31st International Conference on Computa- tional Linguistics. Association for Compu...
2025
-
[28]
Juan Manuel Zambrano Chaves, Eric Wang, Tao Tu, Eeshit Dhaval Vaishnav, Byron Lee, S Sara Mahdavi, Christopher Semturs, David Fleet, Vivek Natarajan, and Shekoofeh Azizi. 2024. Tx-LLM: A Large Language Model for Therapeutics. arXiv:2406.06316
2024 arXiv
-
[29]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. InProceedings of the IEEE/CVF Conference on Compute...
2024
-
[30]
Feixiong Cheng, Weihua Li, Yadi Zhou, Jie Shen, Zengrui Wu, Guixia Liu, Philip W Lee, and Yun Tang. 2012. admetSAR: a comprehensive source and free tool for assessment of chemical ADMET properties
2012
-
[31]
Core Team, Bangjun Xiao, Bingquan Xia, Bo Yang, Bofei Gao, Bowen Shen, Chen Zhang, Chenhong He, Chiheng Lou, Fuli Luo, Gang Wang, et al . 2026. MiMo-V2-Flash Technical Report. arXiv:2601.02780 [cs.CL] https://arxiv.org/ abs/2601.02780
2026 arXiv
-
[32]
Yiming Cui, Xin Yao, Yuxuan Qin, Xin Li, Shijin Wang, and Guoping Hu. 2025. Evaluating large language models on multimodal chemistry olympiad exams. Communications Chemistry8 (2025), 402. doi:10.1038/s42004-025-01782-x
2025 doi
-
[33]
Alex GC de Sá, Yangyang Long, Stephanie Portelli, Douglas EV Pires, and David B Ascher. 2022. toxCSM: comprehensive prediction of small molecule toxicity profiles.Briefings in Bioinformatics23, 5 (2022), bbac337
2022
-
[34]
Fabian Dey and Amedeo Caflisch. 2008. Fragment-based de novo ligand design by multiobjective evolutionary optimization.Journal of chemical information and modeling48, 3 (2008), 679–690
2008
-
[35]
Bing-Xue Du, Yi Xu, Siu-Ming Yiu, Hui Yu, and Jian-Yu Shi. 2023. ADMET property prediction via multi-task graph learning under adaptive auxiliary task selection.iScience26, 11 (2023), 108285. doi:10.1016/j.isci.2023.108285
2023
-
[36]
Fang Du, Haibo Yu, Beiyan Zou, Joseph Babcock, Shunyou Long, and Min Li
-
[37]
Emanuel SR Ehmki, Robert Schmidt, Farina Ohm, and Matthias Rarey. 2019. Comparing molecular patterns using the example of SMARTS: applications and filter collection analysis.Journal of Chemical Information and Modeling59, 6 (2019), 2572–2586. Lin et al
2019
-
[38]
Peter Ertl and Ansgar Schuffenhauer. 2009. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions.Journal of cheminformatics1 (2009), 1–11
2009
-
[39]
Dominga Evangelista, Elliot Nelson, Rachael Skyner, Ben Tehan, Mattia Ber- netti, Marinella Roberti, Maria Laura Bolognesi, and Giovanni Bottegoni. 2025. Application of Deep Learning to Predict the Persistence, Bioaccumulation, and Toxicity of Pharmaceuticals.Journal of Chemic...
2025
-
[40]
Alessio Fallani, Ramil Nugmanov, Jose Arjona-Medina, Jörg Kurt Wegner, Alexandre Tkatchenko, and Kostiantyn Chernichenko. 2025. Pretraining graph transformers with atom-in-a-molecule quantum properties for improved AD- MET modeling.Journal of Cheminformatics17, 1 (2025), 25
2025
-
[41]
Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. 2024. Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.Advances in Neural Information Processing Systems37 (2024), 89098–89124
2024
-
[42]
Mojca Fuart Gatnik and Andrew P. Worth. 2010.Review of Software Tools for Tox- icity Prediction. Technical Report EUR 24489 EN. European Commission, Joint Research Centre, Institute for Health and Consumer Protection, Luxembourg. 13 pages. doi:10.2788/60101
2010 doi
-
[43]
Wenhao Gao, Tianfan Fu, Jimeng Sun, and Connor Coley. 2022. Sample efficiency matters: a benchmark for practical molecular optimization.Advances in neural information processing systems35 (2022), 21342–21357
2022
-
[44]
Kaitlyn M Gayvert, Neel S Madhukar, and Olivier Elemento. 2016. A data-driven approach to predicting successes and failures of clinical trials.Cell chemical biology23, 10 (2016), 1294–1301
2016
-
[45]
Goh, Charles Siegel, Abhinav Vishnu, and Nathan O
Garrett B. Goh, Charles Siegel, Abhinav Vishnu, and Nathan O. Hodas. 2018. Using Rule-Based Labels for Weak Supervised Learning: A ChemNet for Trans- ferable Chemical Property Prediction. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Da...
2018
-
[46]
Google. 2025. Gemini 2.5 Flash Model Description. https://cloud.google.c om/vertex-ai/generative-ai/docs/models/gemini/2-5-flash. Official model documentation page for Gemini 2.5 Flash. Accessed: 2026-01-06
2025
-
[47]
Google DeepMind. 2025. Gemini 2.5 Pro Experimental. https://blog.google/ technology/google-deepmind/gemini-model-thinking-updates-march-2025/ Accessed: May 5, 2025
2025
-
[48]
Google DeepMind. 2025. Gemini 3 Pro — Official Model Page. https://deepmi nd.google/models/gemini/pro/. Accessed: 2026-01-06
2025
-
[49]
Nigel Greene. 2002. Computer systems for the prediction of toxicity: an update. Advanced Drug Delivery Reviews54, 3 (2002), 417–431
2002
-
[50]
Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, et al. 2025. Seed1.5-VL Technical Report. arXiv:2505.07062 [cs.CV] https://arxiv.org/abs/2505.07062
2025 arXiv
-
[51]
Chawla, Olaf Wiest, and Xiangliang Zhang
Kehan Guo, Bozhao Nan, Yujun Zhou, Taicheng Guo, Zhichun Guo, Mihir Surve, Zhenwen Liang, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Can LLMs Solve Molecule Puzzles? A Multimodal Benchmark for Molecular Structure Elucidation. InAdvances in Neural Information Pro...
2024
-
[52]
Sumin Ha, Jun Hyeong Kim, Yinhua Piao, and Sun Kim. 2025. MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model. arXiv:2503.04780 [cs.CL] https://arxiv.org/abs/2503.04780
2025 arXiv
-
[53]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi
-
[54]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Con- verge to a Local Nash Equilibrium. InAdvances in Neural Information Process- ing Systems, Vol. 30. Curran Associates, Inc., Red Ho...
2017
-
[55]
Smith, and Jiebo Luo
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A. Smith, and Jiebo Luo. 2023. PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE/CVF, Paris, France, 2963–2975
2023
-
[56]
David Huang, Sauhaarda Raunak Chowdhuri, Andrew Li, Alex Li, Ayush Agrawal, Kameron Gano, and Andy Zhu. 2022. A Unified System for Molecular Property Predictions: Oloren ChemEngine and Its Applications.ChemRxiv (2022). doi:10.26434/chemrxiv-2022-zz776
2022 doi
-
[57]
Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. 2022. Artificial intelligence foundation for therapeutic science.Nature chemical biology18, 10 (2022), 1033–1036
2022
-
[58]
Roohani, Jure Leskovec, Connor W
Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf H. Roohani, Jure Leskovec, Connor W. Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. 2021. Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development. InNeurIPS 2021 Track on Datasets ...
2021
-
[59]
Ruili Huang, Menghang Xia, Dac-Trung Nguyen, Tongan Zhao, Srilatha Saka- muru, Jinghua Zhao, Sampada A Shahane, Anna Rossoshek, and Anton Sime- onov. 2016. Tox21Challenge to build predictive models of nuclear receptor and stress response pathways as mediated by exposure to env...
2016
-
[60]
Yuqing Huang, Rongyang Zhang, Xuesong He, Xuyang Zhi, Hao Wang, Xin Li, Feiyang Xu, Deguang Liu, Huadong Liang, Yi Li, et al. 2024. ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models. arXiv:2409.13989
2024 arXiv
-
[61]
Yunhui Jang, Jaehyung Kim, and Sungsoo Ahn. 2024. Chain-of-Thoughts for Molecular Understanding. arXiv:2410.05610
2024 arXiv
-
[62]
Woojeong Jin, Yu Cheng, Yelong Shen, Weizhu Chen, and Xiang Ren. 2022. A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volum...
2022 doi
-
[63]
Ives, Yoseph Barash, Cesar de la Fuente-Nunez, Jacob R
Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan, Maggie Ziyu Huan, Jiaming Liang, Marcelo Der Torossian Torres, Jiatao Liang, Zachary G. Ives, Yoseph Barash, Cesar de la Fuente-Nunez, Jacob R. Gardner, Mark Yatskar, et al
-
[64]
Abdul Karim, Matthew Lee, Thomas Balle, and Abdul Sattar. 2021. CardioTox net: a robust predictor for hERG channel blockade based on deep learning meta-feature ensembles.Journal of cheminformatics13 (2021), 1–13
2021
-
[65]
Jihyung Kil, Zheda Mai, Justin Lee, Arpita Chowdhury, Zihe Wang, Kerrie Cheng, Lemeng Wang, Ye Liu, and Wei-Lun Harry Chao. 2024. Mllm-compbench: A comparative reasoning benchmark for multimodal llms.Advances in Neural Information Processing Systems37 (2024), 28798–28827
2024
-
[66]
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al . 2021. PubChem in 2021: new data content and improved web interfaces.Nucleic acids research49, D1 (2021), D1388–D1395
2021
-
[67]
Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, et al
-
[68]
Kerstin Kläser, Blazej Banaszewski, Samuel Maddrell-Mander, Callum McLean, Luis Müller, Ali Parviz, Shenyang Huang, and Andrew Fitzgibbon. 2024. MiniMol: A Parameter-Efficient Foundation Model for Molecular Learning. arXiv:2404.14986
2024 arXiv
-
[69]
arXiv:2508.10899
A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design. arXiv:2508.10899
-
[70]
Siddhartha Laghuvarapu, Namkyeong Lee, Chufan Gao, and Jimeng Sun. 2024. MolTextQA: A Curated Question-Answering Dataset and Benchmark for Molec- ular Structure-Text Relationship Learning. https://openreview.net/forum?id= gwGHBD9ZKU
2024
-
[71]
Alexey Lagunin, Dmitrii Filimonov, Alexey Zakharov, Wei Xie, Ying Huang, Fucheng Zhu, Tianxiang Shen, Jianhua Yao, and Vladimir Poroikov. 2009. Computer-aided prediction of rodent carcinogenicity by PASS and CISOC-PSCT. QSAR & Combinatorial Science28, 8 (2009), 806–810
2009
-
[72]
Yunshi Lan, Xiang Li, Xin Liu, Yang Li, Wei Qin, and Weining Qian. 2023. Improving Zero-shot Visual Question Answering via Large Language Models with Reasoning Question Prompts. InProceedings of the 31st ACM International Conference on Multimedia. Association for Computing Mac...
2023
-
[73]
Greg Landrum. 2013. Rdkit documentation.Release1, 1-79 (2013), 4
2013
-
[74]
arXiv:2504.07491 [cs.CV] https://arxiv.org/ab s/2504.07491
Kimi-VL Technical Report. arXiv:2504.07491 [cs.CV] https://arxiv.org/ab s/2504.07491
-
[75]
Chanhui Lee, Hanbum Ko, Yuheon Song, YongJun Jeong, Rodrigo Hormazabal, Sehui Han, Kyunghoon Bae, Sungbin Lim, and Sungwoong Kim. 2025. Mol- LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization. arXiv:2502.02810 [cs.LG] https://arxiv.org/abs/2502.02810
2025 arXiv
-
[76]
Jiayi Kuang, Ying Shen, Jingyou Xie, Haohao Luo, Zhe Xu, Ronghao Li, Yinghui Li, Xianfeng Cheng, Xika Lin, and Yu Han. 2025. Natural language understand- ing and inference with MLLM in visual question answering: A survey.Comput. Surveys57, 8 (2025), 1–36
2025
-
[77]
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024. LLaVA- OneVision: Easy Visual Task Transfer. arXiv:2408.03326 [cs.CV] https://arxiv. org/abs/2408.03326
2024 arXiv
-
[78]
Haorui Li, Shengchao Liu, Hongyu Guo, and Anima Anandkumar. 2024. Geometry-text Multi-modal Foundation Model for Reactivity-oriented Molecule Editing. NeurIPS 2024 Workshop on AI for New Drug Modalities (OpenReview). https://openreview.net/forum?id=A9FlQMKxJ4 Are MLLMs Ready f...
2024
-
[79]
Hao Li, Liuzhenghao Lv, He Cao, Zijing Liu, Zhiyuan Yan, Yu Wang, Yonghong Tian, Yu Li, and Li Yuan. 2025. How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Compre- hension. arXiv:2504.12314 [cs.CL] https://arxiv.org/...
2025 arXiv
-
[80]
Jiatong Li, Junxian Li, Weida Wang, Yunqing Liu, Changmeng Zheng, Dongzhan Zhou, Xiao-yong Wei, and Qing Li. 2025. Speak-to-Structure: Evaluat- ing LLMs in Open-domain Natural Language-Driven Molecule Generation. arXiv:2412.14642 arXiv preprint
2025 arXiv
-
[81]
Greg Landrum et al. 2024. RDKit: Open-source cheminformatics. https://www. rdkit.org. Accessed: 2025-05-11
2024
-
[82]
Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu. 2021. Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models. InProceedings of the IEEE/CVF International Conference on Computer Vision. IEEE/CVF, Montreal, QC, Canada, 2042–2051
2021
-
[83]
Albert P Li. 2004. Accurate prediction of human drug toxicity: a major challenge in drug development.Chemico-biological interactions150, 1 (2004), 3–7
2004
-
[84]
Fei Lin, Jing Yang, Dali Sun, Levente Kovács, and Fei-Yue Wang. 2025. Au- tonomous drug discovery with parallel intelligence.IEEE/CAA Journal of Automatica Sinica12, 8 (2025), 1742–1744
2025
-
[85]
Christopher A Lipinski, Franco Lombardo, Beryl W Dominy, and Paul J Feeney
-
[86]
Pengfei Liu, Yiming Ren, Jun Tao, and Zhixiang Ren. 2024. Git-mol: A multi- modal large language model for molecular science with graph, image, and text. Computers in biology and medicine171 (2024), 108073
2024
-
[87]
Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, and Animashree Anandkumar. 2023. Multi-modal molecule structure–text model for text-based retrieval and editing.Nature Machine Intelligence5, 12 (2023), 1447–1457
2023
-
[88]
Junxian Li, Di Zhang, Xunzhi Wang, Zeying Hao, Jingdi Lei, Qian Tan, Cai Zhou, Wei Liu, Yaotian Yang, Xinrui Xiong, et al . 2025. ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area. InProceedings of the AAAI Conference on Artificial Intelligence...
2025
-
[89]
Gen Luo, Xue Yang, Wenhan Dou, Zhaokai Wang, Jiawen Liu, Jifeng Dai, Yu Qiao, and Xizhou Zhu. 2024. Mono-InternVL: Pushing the Boundaries of Mono- lithic Multimodal Large Language Models with Endogenous Visual Pre-training. arXiv:2410.08202 arXiv preprint
2024 arXiv
-
[90]
Sihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang, Xiangnan He, Kenji Kawaguchi, Tat-Seng Chua, and Qi Tian. 2024. Towards 3D Molecule-Text Interpretation in Language Models. arXiv:2401.13923 arXiv preprint
2024 arXiv
-
[91]
Oscar Mañas, Benno Krojer, and Aishwarya Agrawal. 2024. Improving Auto- matic VQA Evaluation Using Large Language Models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. AAAI Press, Vancouver, Canada, 4171–4179. doi:10.1609/aaai.v38i5.28212
2024 doi
-
[92]
Andreas Mayr, Günter Klambauer, Thomas Unterthiner, and Sepp Hochreiter
-
[93]
Leland McInnes, John Healy, and James Melville. 2018. UMAP: Uniform Mani- fold Approximation and Projection for Dimension Reduction. arXiv:1802.03426 arXiv preprint
2018 arXiv
-
[94]
Oscar Méndez-Lucio, Christos A Nicolaou, and Berton Earnshaw. 2024. MolE: a foundation model for molecular graphs using disentangled attention.Nature Communications15, 1 (2024), 9431
2024
-
[95]
Samar Monem, Alaa H Abdel-Hamid, and Aboul Ella Hassanien. 2025. Drug toxicity prediction model based on enhanced graph neural network.Computers in Biology and Medicine185 (2025), 109614
2025
-
[96]
Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. 2024. Conversational Drug Editing Using Retrieval and Domain Feedback. InThe Twelfth International Conference on Learning Representations. OpenReview.net, Vienna, Austria, 1–1
2024
-
[97]
Oscar Méndez-Lucio, Christos Nicolaou, and Berton Earnshaw. 2022. MolE: a molecular foundation model for drug discovery. arXiv:2211.02657 [q-bio.QM] https://arxiv.org/abs/2211.02657
2022 arXiv
-
[98]
Qitan Lv, Tianyu Liu, and Hong Wang. 2025. Exploiting Edited Large Language Models as General Scientific Optimizers. arXiv:2503.09620 [math.OC] https: //arxiv.org/abs/2503.09620
2025 arXiv
-
[99]
Zhangming Niu, Xianglu Xiao, Wenfan Wu, Qiwei Cai, Yinghui Jiang, Wangzhen Jin, Minhao Wang, Guojian Yang, Lingkang Kong, Xurui Jin, et al . 2024. PharmaBench: Enhancing ADMET benchmarks with large language models. Scientific Data11, 1 (2024), 985
2024
-
[100]
Notwell and Michael W
James H. Notwell and Michael W. Wood. 2023. ADMET property prediction through combinations of molecular fingerprints. arXiv:2310.00174 [q-bio.BM] https://arxiv.org/abs/2310.00174
2023 arXiv
-
[101]
OpenAI. 2024. GPT-4V System Card. https://openai.com/index/gpt-4v-system- card/ Accessed: 2024-11-16
2024
-
[102]
OpenAI. 2025. GPT-5.1: A smarter, more conversational ChatGPT. https: //openai.com/index/gpt-5-1/. Accessed: 2025-12-06
2025
-
[103]
OpenAI. 2025. Introducing GPT-4.1 in the API. https://openai.com/index/gpt- 4-1/. Accessed: 2025-05-05
2025
-
[104]
OpenAI. 2025. Introducing GPT-5.2. https://openai.com/index/introducing-gpt- 5-2/
2025
-
[105]
Elwood A Mullins, Rongxin Shi, and Brandt F Eichman. 2017. Toxicity and repair of DNA adducts produced by the natural product yatakemycin.Nature chemical biology13, 9 (2017), 1002–1008
2017
-
[106]
Goucher, Adam Perelman, Aditya Ramesh, et al
OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, et al. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https: //arxiv.org/abs/2410.21276
2024 arXiv
-
[107]
Alaa-Eldin F Nassar, Amin M Kamel, and Caroline Clarimont. 2004. Improv- ingthe decision-making process in structural modification of drug candidates: reducing toxicity.Drug Discovery Today9, 24 (2004), 1055–1064
2004
-
[108]
Douglas EV Pires, Tom L Blundell, and David B Ascher. 2015. pkCSM: predicting small-molecule pharmacokinetic and toxicity properties using graph-based signatures.Journal of medicinal chemistry58, 9 (2015), 4066–4072
2015
-
[109]
Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Arta- monov, Vladimir Aladinskiy, Mark Veselov, et al. 2020. Molecular sets (MOSES): a benchmarking platform for molecular generation m...
2020
-
[110]
Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Gunter Klambauer. 2018. Fréchet ChemNet distance: a metric for generative models for molecules in drug discovery.Journal of chemical information and modeling 58, 9 (2018), 1736–1741
2018
-
[111]
Mayk Caldas Ramos, Christopher J Collison, and Andrew D White. 2025. A review of large language models and autonomous agents in chemistry.Chemical science16, 6 (2025), 2514–2572
2025
-
[112]
Ann M Richard, Richard S Judson, Keith A Houck, Christopher M Grulke, Patra Volarath, Inthirany Thillainadarajah, Chihae Yang, James Rathman, Matthew T Martin, John F Wambaugh, et al. 2016. ToxCast chemical landscape: paving the road to 21st century toxicology.Chemical researc...
2016
-
[113]
David Rogers and Mathew Hahn. 2010. Extended-connectivity fingerprints. Journal of chemical information and modeling50, 5 (2010), 742–754
2010
-
[114]
OpenAI. 2025. Introducing OpenAI o3 and o4-mini. https://openai.com/index/i ntroducing-o3-and-o4-mini/. Accessed: 2026-01-06
2025
-
[115]
Robert Schmidt, Emanuel SR Ehmki, Farina Ohm, Hans-Christian Ehrlich, An- driy Mashychev, and Matthias Rarey. 2019. Comparing molecular patterns using the example of SMARTS: theory and algorithms.Journal of Chemical Information and Modeling59, 6 (2019), 2560–2571
2019
-
[116]
OpenGVLab. 2024. InternVL Family: Closing the Gap to Commercial Multimodal Models with Open-Source Suites —— A Pioneering Open-Source Alternative to GPT-5. https://github.com/OpenGVLab/InternVL. Accessed: 2025-05-05
2024
-
[117]
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 7881–7892
2020
-
[118]
Duxin Sun, Wei Gao, Hongxiang Hu, and Simon Zhou. 2022. Why 90% of clinical drug development fails and how to improve it?Acta Pharmaceutica Sinica B12, 7 (2022), 3049–3062
2022
-
[119]
Iurii Sushko, Elena Salmina, Vladimir A Potemkin, Gennadiy Poda, and Igor V Tetko. 2012. ToxAlerts: a web server of structural alerts for toxic chemicals and compounds with potential adverse reactions
2012
-
[120]
Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu, Wanli Ouyang, Lei Bai, Yuqiang Li, and Tianfan Fu. 2025. ChemMLLM: Chemical Multimodal Large Language Model. arXiv:2505.16326 [cs.LG] https://arxiv.org/abs/2505.16326
2025 arXiv
-
[121]
T. T. Tanimoto. 1958.An Elementary Mathematical Theory of Classification and Prediction. International Business Machines Corporation, New York, NY, USA. https://books.google.com.hk/books?id=yp34HAAACAAJ
1958
-
[122]
Tencent Hunyuan. 2024. Hunyuan Vision. https://cloud.tencent.com/docume nt/product/1729/105701 Accessed: 2025-05-07
2024
-
[123]
Zohar Rosenwasser, Erez Levanon, Michael Levitt, and Gal Oren. 2025. Leverag- ing GPT Continual Fine-Tuning for Improved RNA Editing Site Prediction. ICLR 2025 Workshop on Machine Learning for Genomics Explorations. Workshop paper
2025
-
[124]
Gail A Van Norman. 2019. Limitations of animal studies for predicting toxicity in clinical trials: is it time to rethink our current approach?JACC: Basic to Translational Science4, 7 (2019), 845–854
2019
-
[125]
Arne Schneuing, Charles Harris, Yuanqi Du, Kieran Didi, Arian Jamasb, Ilia Igashov, Weitao Du, Carla Gomes, Tom L Blundell, Pietro Lio, et al . 2024. Structure-based drug design with equivariant diffusion models.Nature Compu- tational Science4, 12 (2024), 899–909
2024
-
[127]
Jianmin Wang, Peng Zhou, Zixu Wang, Wei Long, Yangyang Chen, Kyoung Tai No, Dongsheng Ouyang, Jiashun Mao, and Xiangxiang Zeng. 2024. Diffusion- based Generative Drug-like Molecular Editing with Chemical Natural Language. Journal of Pharmaceutical Analysis14, 2 (2024), 101137
2024
-
[128]
Shuangquan Wang, Huiyong Sun, Hui Liu, Dan Li, Youyong Li, and Tingjun Hou. 2016. ADMET evaluation in drug discovery. 16. Predicting hERG blockers by combining multiple pharmacophores and machine learning approaches. Molecular pharmaceutics13, 8 (2016), 2855–2866
2016
-
[129]
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency. arXiv:2508.18265 [cs.CV] https://arxiv.org...
2025 arXiv
-
[130]
Zixu Wang, Yangyang Chen, Pengsen Ma, Zhou Yu, Jianmin Wang, Yuansheng Liu, Xiucai Ye, Tetsuya Sakurai, and Xiangxiang Zeng. 2025. Image-based generation for molecule design with SketchMol.Nature Machine Intelligence7, 2 (2025), 244–255
2025
-
[131]
Chengyue Wu, Yixiao Ge, Qiushan Guo, Jiahao Wang, Zhixuan Liang, Zeyu Lu, Ying Shan, and Ping Luo. 2024. Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots. arXiv:2405.07990 [cs.CL] https://arxiv.org/a...
2024 arXiv
-
[132]
Gemma Turon, Jason Hlozek, John G Woodland, Ankur Kumar, Kelly Chibale, and Miquel Duran-Frigola. 2023. First fully-automated AI/ML virtual screening cascade implemented at a drug discovery centre in Africa.Nature Communica- tions14, 1 (2023), 5736
2023
-
[133]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang,...
2024 arXiv
-
[134]
Li, Xiang Lin, Wenhao Gao, Tianfan Fu, Manolis Kellis, Bradley L
Alejandro Velez-Arce, Kexin Huang, Michelle M. Li, Xiang Lin, Wenhao Gao, Tianfan Fu, Manolis Kellis, Bradley L. Pentelute, and Marinka Zitnik. 2024. TDC-2: Multimodal Foundation for Therapeutic Science. doi:10.1101/2024.06. Lin et al. 12.598655
2024 doi
-
[135]
xAI. 2025. Grok 4. https://x.ai/news/grok-4. Accessed: 2026-01-06
2025
-
[136]
Guoli Xiong, Zhenxing Wu, Jiacai Yi, Li Fu, Zhijiang Yang, Changyu Hsieh, Mingzhu Yin, Xiangxiang Zeng, Chengkun Wu, Aiping Lu, et al . 2021. AD- METlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties.Nucleic acids research49, ...
2021
-
[137]
Congying Xu, Feixiong Cheng, Lei Chen, Zheng Du, Weihua Li, Guixia Liu, Philip W Lee, and Yun Tang. 2012. In silico prediction of chemical Ames mutagenicity.Journal of chemical information and modeling52, 11 (2012), 2840–2847
2012
-
[138]
Youjun Xu, Ziwei Dai, Fangjin Chen, Shuaishi Gao, Jianfeng Pei, and Luhua Lai. 2015. Deep learning for drug-induced liver injury.Journal of chemical information and modeling55, 10 (2015), 2085–2093
2015
-
[139]
Hongbin Yang, Chaofeng Lou, Lixia Sun, Jie Li, Yingchun Cai, Zhuang Wang, Weihua Li, Guixia Liu, and Yun Tang. 2019. admetSAR 2.0: web-service for prediction and optimization of chemical ADMET properties.Bioinformatics35, 6 (2019), 1067–1069
2019
-
[140]
Hengzheng Yang, Jian Xiu, Weiqi Yan, Kaifeng Liu, Huizi Cui, Zhibang Wang, Qizheng He, Yilin Gao, and Weiwei Han. 2025. Large Language Models as Tools for Molecular Toxicity Prediction: AI Insights into Cardiotoxicity.Journal of Chemical Information and Modeling65, 5 (2025), 2268–2282
2025
-
[141]
Jiaxin Wu, Ting Zhang, Rubing Chen, Wengyu Zhang, Chen Jason Zhang, Xiaoyong Wei, and Li Qing. 2025. MolGround: A Benchmark for Molecular Grounding. arXiv:2503.23668 arXiv preprint
2025 arXiv
-
[142]
Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang, Junhong Huang, Longyue Wang, Wei Liu, and Xiangxiang Zeng. 2025. Drugassist: A large language model for molecule optimization.Briefings in Bioinformatics26, 1 (2025), bbae693
2025
-
[143]
xAI. 2024. Grok 2 Vision Model Description. https://docs.x.ai/developers/relea se-notes. Accessed: 2025-12-06
2024
-
[144]
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2024. MM-Vet: Evaluating Large Multi- modal Models for Integrated Capabilities. InProceedings of the 41st International Conference on Machine Learning. PMLR, Vienna, Aus...
2024
-
[145]
Xiangxiang Zeng, Hongxin Xiang, Linhui Yu, Jianmin Wang, Kenli Li, Ruth Nussinov, and Feixiong Cheng. 2022. Accurate prediction of molecular prop- erties and drug targets using a self-supervised image representation learning framework.Nature Machine Intelligence4, 11 (2022), 1004–1016
2022
-
[146]
Di Zhang, Wei Liu, Qian Tan, Jingdan Chen, Hang Yan, Yuliang Yan, Jiatong Li, Weiran Huang, Xiangyu Yue, Wanli Ouyang, Dongzhan Zhou, Shufei Zhang, Mao Su, Han-Sen Zhong, and Yuqiang Li. 2024. ChemLLM: A Chemical Large Language Model. arXiv:2402.06852 [cs.AI] https://arxiv.org...
2024 arXiv
-
[147]
Odin Zhang, Haitao Lin, Hui Zhang, Huifeng Zhao, Yufei Huang, Chang-Yu Hsieh, Peichen Pan, and Tingjun Hou. 2024. Deep Lead Optimization: Leverag- ing Generative AI for Structural Modification.Journal of the American Chemical Society146, 46 (2024), 31357–31370
2024
-
[148]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[149]
Xuanle Zhao, Shuxin Zeng, Xinyuan Cai, Xiang Cheng, Duzhen Zhang, Xiuyi Chen, and Bo Xu. 2025. TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks. arXiv:2511.06283 [cs.CV] https://arxiv.org/abs/2511.06283
2025
-
[150]
Lijuan Yang, Chao Jin, Guanghui Yang, Zhitong Bing, Liang Huang, Yuzhen Niu, and Lei Yang. 2023. Transformer-based deep learning method for optimizing ADMET properties of lead compounds.Physical Chemistry Chemical Physics 25, 3 (2023), 2377–2385
2023
-
[151]
Zehua Zhao, Zhixian Huang, Junren Li, Siyu Lin, Junting Zhou, Fengqi Cao, Kun Zhou, Rui Ge, Tingting Long, Yuexiang Zhu, et al . 2025. SUPERChem: A Multimodal Reasoning Benchmark in Chemistry. arXiv:2512.01274 arXiv preprint
2025
-
[152]
Baker, Ziqi Chen, Xia Ning, and Huan Sun
Botao Yu, Frazier N. Baker, Ziqi Chen, Xia Ning, and Huan Sun. 2024. LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Compre- hensive, High-Quality Instruction Tuning Dataset. arXiv:2402.09391 [cs.AI] https://arxiv.org/abs/2402.09391
2024 arXiv
-
[153]
Zhipu AI. 2024. GLM-4V-Plus-0111. https://docs.bigmodel.cn/cn/guide/models /vlm/glm-4v-plus-0111 Accessed: 2025-05-05
2024
-
[154]
Hao Zhu, Todd M Martin, Lin Ye, Alexander Sedykh, Douglas M Young, and Alexander Tropsha. 2009. Quantitative structure- activity relationship modeling of rat acute toxicity by oral exposure.Chemical research in toxicology22, 12 (2009), 1913–1921
2009
-
[155]
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. 2025. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv:2504.10479 [cs.CV] https://arxiv.org/abs/2504.10479
2025 arXiv
-
[156]
coordinate system
Marinka Zitnik, Monica Agrawal, and Jure Leskovec. 2018. Modeling polyphar- macy side effects with graph convolutional networks.Bioinformatics34, 13 (2018), i457–i466. Are MLLMs Ready for Structure-Level Molecular Detoxification? A Rationale for Selecting the TxGemma-Predict M...
2018 arXiv
-
[160]
Zihan Zhao, Bo Chen, Jingpiao Li, Lu Chen, Liyang Wen, Pengyu Wang, Zichen Zhu, Danyang Zhang, Yansi Li, Zhongyang Dai, et al . 2024. ChemDFM-X: towards large multimodal model for chemistry.Science China Information Sciences67, 12 (2024), 220109
2024
-
[162]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623
2023
-
[2011]
hERGCentral: a large database to store, retrieve, and analyze compound- human Ether-a-go-go related gene channel interactions to facilitate cardiotoxi- city assessment in drug development.Assay and drug development technologies 9, 6 (2011), 580–588
2011
-
[2012]
Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings.Advanced drug delivery reviews64 (2012), 4–17
2012
-
[2016]
DeepTox: toxicity prediction using deep learning.Frontiers in Environ- mental Science3 (2016), 80
2016
-
[2018]
ProTox-II: a webserver for the prediction of toxicity of chemicals.Nucleic acids research46, W1 (2018), W257–W263
2018
-
[2019]
arXiv:1904.09675 arXiv preprint
BERTScore: Evaluating Text Generation with BERT. arXiv:1904.09675 arXiv preprint
1904 arXiv
-
[2021]
In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Punta Cana, Dominican Republic, 7514–7528. doi:10.18653/v1/2021.emnlp-main.595
2021 doi
-
[2024]
Hybrid fragment-SMILES tokenization for ADMET prediction in drug discovery.BMC bioinformatics25, 1 (2024), 255
2024
-
[2025]
doi:10.1002/advs.202413405
Machine Learning-Enabled Drug-Induced Toxicity Prediction.Advanced Science12, 16 (2025), e2413405. doi:10.1002/advs.202413405
2025 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.