Pith. sign in

REVIEW 3 major objections 6 minor 69 references

CovDocker: Benchmarking Covalent Drug Design with Tasks, Datasets, and Solutions

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read CovDocker decomposes covalent docking into three learnable tasks and releases 2,754 curated complexes with baselines.

desk verdict The dataset is the contribution; the headline RMSD(IB) metric is largely measuring the postprocessor, not the learned pose. read the letter →

arxiv 2506.21085 v1 pith:GOEFFSYE submitted 2025-06-26 q-bio.BM cs.AIcs.LG

classification q-bio.BMcs.AIcs.LG
keywords moleculardockingcovalentdrugdesigndeeplearningbenchmarksprotein-ligandinteractionreactivesitepredictionreactionpose
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most docking tools and deep-learning models treat protein–ligand binding as non-covalent, so they miss the covalent bond that makes many drugs bind strongly and persistently. This paper argues that covalent docking should be decomposed into three separate learnable steps: locating the reactive pocket and residue, predicting the product of the covalent reaction, and predicting the docked pose of the post-reactive ligand. To support that decomposition, it assembles 2,754 covalent complexes from two public structural databases, reconstructs the post-reactive ligands, and releases preprocessed data, code, and trained weights. It also adapts existing models—Uni-Mol for site and pose prediction, Chemformer for reaction prediction—as baselines, including an auxiliary loss that keeps the covalent bond at the shortest ligand–pocket distance. If the benchmark is sound, deep learning can be trained and compared on covalent docking in a reproducible way, which would speed up the design of selective covalent inhibitors.

What carries the argument

The machinery is the task decomposition plus the data pipeline that feeds it. CovDocker converts PDB LINK records and ligand metadata from CovPDB and CovBinderInPDB into a consistent set of pre-reactive SMILES, post-reactive SMILES, and docked structures; the post-reactive ligand is reconstructed by aligning the stable Chemical Component Dictionary form to the PDB coordinates, deleting extra atoms, patching bond orders, and adding hydrogens. On top of this data, each task has a defined model: a residue-level Uni-Mol encoder with cross-attention predicts the pocket center and reactive residue; Chemformer, a Transformer sequence model, predicts reaction products; and a finetuned Uni-Mol docking model with the covalent-distance auxiliary loss $L_{cov}$ and optional bond-length postprocessing predicts poses. The time-based split is the guard against data leakage.

What would settle it

Reconstruct the post-reactive ligand for a random sample of CovDocker entries directly from the raw PDB LINK records and the published preprocessing rules, without consulting the released labels; if the resulting SMILES and bond orders fail to match the release on even a few entries, the benchmark targets inherit those errors. A second check is to retrain the Task 3 model without the $L_{cov}$ loss and compare RMSD (IB) on the same test split: the paper reports a large gap at tight thresholds, so an independent run should reproduce that gap if the loss is doing the claimed work.

Watch

Extended reading notes

Core claim

On its own terms, the paper’s central claim is that covalent docking is not one monolithic prediction problem but three coupled tasks, and that a benchmark built from PDB-derived covalent complexes can make each task tractable for deep learning. Task 1 predicts the pocket center and the reactive residue; Task 2 predicts the post-reactive product SMILES from the pre-reactive ligand and reactive residue; Task 3 predicts the docked pose of the product within a pocket, with a loss term $L_{cov} = \mathrm{ReLU}(d_{ij} - D_{inter})$ that biases the predicted covalent bond toward the shortest inter-molecular distance. The paper reports that this decomposition, together with time-based splits of 2,308 training, 223 validation, and 223 test entries, yields baselines that beat traditional non-covalent and covalent docking tools on pose accuracy, and that the covalent constraint plus a bond-length postprocessing step nearly saturates the new covalent-bond RMSD metric.

Load-bearing premise

The entire benchmark depends on the hand-reconstructed post-reactive ligand labels: for each complex, the authors align the stable ligand form from the Chemical Component Dictionary to PDB coordinates, delete or add atoms, patch bond orders, and add hydrogens by hand, so if those reconstructed products are wrong or ambiguous, the reaction-prediction targets and the pose labels built from them are wrong.

Editorial extensions

If this is right

  • A machine-learning model can now be trained on covalent docking from a single downloadable dataset, with held-out splits chosen by discovery date rather than by curation quality.
  • Covalent bond quality becomes directly measurable: the new RMSD (IB) metric checks whether the bonded ligand atom lands on the bonded protein atom, not just whether the whole pose is close.
  • Reaction prediction on this dataset is harder than on standard small-molecule reaction sets, so gains scored on CovDocker may transfer to real covalent inhibitor chemistry.
  • The blind-docking pipeline result (0.4% under 3 Å RMSD versus 41.3% for site-specific docking) shows that the three stages must be improved jointly rather than in isolation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the benchmark labels hold up, the same three-task decomposition could be applied to other irreversible binding modalities, such as covalent protein–DNA or protein–carbohydrate cross-links, by swapping the residue vocabulary and reaction templates.
  • A direct next experiment is to feed oracle reactive-site and oracle reaction labels into the pose model and measure how much each stage’s error contributes to the final blind-docking score, which would locate where the pipeline loses the most accuracy.
  • The generality of the $L_{cov}$ loss could be tested on a non-covalent docking benchmark by imposing an artificial close-contact anchor between ligand and pocket, to see whether the gain comes from a chemistry-specific signal or from a generic short-distance prior.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes CovDocker, a benchmark for covalent drug design that decomposes covalent docking into three tasks: reactive location prediction, covalent reaction prediction, and covalent docking pose prediction. The authors construct a dataset of 2,754 complexes from CovPDB and CovBinderInPDB, provide time-based splits, adapt Uni-Mol and Chemformer as baselines, introduce an auxiliary covalent loss and a postprocessing step, and release the preprocessed data, code, and trained weights. The paper claims that this provides a comprehensive, rigorous, and reproducible framework for advancing covalent drug design.

Significance. The dataset release and task decomposition are genuinely valuable: 2,754 entries, 22 reaction mechanisms, 10 target amino acids, time-based splits, and publicly available code and weights are a clear step beyond the existing evaluation-only covalent docking benchmarks. If the reconstructed post-reactive labels are reliable, Tasks 1 and 2 provide useful ML-ready testbeds for reactive-site and reaction-product prediction. The open availability of the resource is a concrete strength that should be credited. However, the Task 3 evaluation is weakened by the RMSD(IB) metric and postprocessor issue described below, so the benchmark's claim to rigorously evaluate covalent docking accuracy is not yet fully established.

major comments (3)
  1. [Sections 3.4, 4.3, and Tables 4 and 7] The RMSD(IB) metric is not a measure of pose accuracy and is almost entirely determined by the postprocessor. In Section 3.4, the postprocessor samples a bond length l' from the dataset-wide normal distribution N(mu, sigma) and moves the bonded ligand atom so that the predicted covalent bond length equals l'. RMSD(IB) as defined in Section 4.3 is the distance between the bonded ligand atom and the bonded protein atom, which is exactly this bond length. Therefore a model that places the warhead atom at any plausible bond length from the reactive residue will score near-perfectly on RMSD(IB) regardless of whether the rest of the pose is correct. The paper's own ablation in Table 7 demonstrates this: rows (b) to (a), postprocessing alone raises RMSD(IB)<0.5 Å from 59.2% to 77.1% while whole-ligand RMSD<2 Å stays at 37.7%; Table 4 shows the same pattern for Ours-p (60.1% to 79.1% with RMSD<2 Å unchanged at 37.2%). Since Table 4 presents RMSD(IB) as the 'covalent bond precision' result, these numbers are mechanical and do not support the claim that the benchmark rigorously evaluates covalent docking. Please make whole-ligand RMSD the primary pose-quality criterion and introduce a warhead-position RMSD that compares the predicted bonded ligand atom with the ground-truth bonded ligand atom after alignment; the bond length can be reported only as a sanity check, not as a pose-accuracy score.
  2. [Appendix A.2] The post-reactive ligand labels underpin both the Task 2 reaction targets and the Task 3 pose labels, but they are reconstructed heuristically: stable ligand forms from the Chemical Component Dictionary are aligned to PDB coordinates, atoms are deleted or added, bond orders are patched, and hydrogen atoms are added manually. The paper does not provide quantitative validation of these reconstructions. Errors or ambiguities in this step would propagate to the reaction-prediction SMILES and to the docking pose labels, invalidating the reported baseline numbers. Please validate a sample against hand-curated covalent complexes such as the Keseru benchmark, report alignment failure rates and the frequency of each patch type, and assess how much the benchmark results change under alternative reconstruction choices.
  3. [Section 3.4 and Table 4 reproducibility] The postprocessing step samples l' randomly from a normal distribution at inference, which makes the Ours-p results stochastic and not exactly reproducible without a fixed seed or multiple sampling. Please either fix the random seed, average results over multiple samples, or replace the random sample with a fixed quantile, and report the resulting variance.
minor comments (6)
  1. [Section 3.2] The text says the pocket is defined as a circle with radius 20 Å around the pocket center; this should be a sphere in three-dimensional space.
  2. [Section 4.4] There is a typo: 'entry numer' should be 'entry number'.
  3. [Table 4 caption] The caption contains the typo 'ligand aotm' and should be 'ligand atom'; also, the asterisk for AutoDock4(cov) should be defined in the caption before it appears in the table.
  4. [Section 5.2] The dataset name 'USTPO' appears to be a typo for 'USPTO'.
  5. [Section 4.1 and Appendix A.2] The main text states that chains exceeding 1,024 amino acids are excluded for Task 1, while Appendix A.2 says the cutoff is 1,022 residues; please clarify which value was actually used.
  6. [Equation (5)] The notation is dimensionally unclear: d_ij is described as a scalar bond distance while D_inter is a distance map matrix; please specify whether the loss is computed on the single matrix entry corresponding to the covalent pair or aggregated over all ligand-pocket pairs.

Circularity Check

1 steps flagged · score 6.0 of 10

RMSD(IB) improvement is enforced by the §3.4 bond-length postprocessor, making the cov-docking 'covalent bond precision' result partially circular.

  1. fitted input called prediction [Section 3.4 (post-processing), Section 4.3 (RMSD(IB) metric), Tables 4 and 7]
    "We have analyzed all the covalent bond lengths, with their mean value μ and variance σ. We then randomly sample one value l′ from the normal distribution of N(μ,σ), and if the predicted covalent bond length exceeds l′, we will move the ligand atom involved in the covalent bond toward the reactive aa to achieve the target bond length l′. ... To better evaluate the model’s performance in covalent bond formation, we introduce a new metric, RMSD (IB), which measures the RMSD between the covalently bonded ligand atom and the bonded protein atom."

    RMSD(IB) is computed from exactly one inter-bond pair: the bonded ligand atom and the bonded protein atom, i.e., the covalent bond length. The postprocessor explicitly moves the bonded ligand atom until that inter-bond distance equals a freshly sampled l′ from N(μ,σ), the global distribution of covalent bond lengths fitted from the same benchmark data. Thus the quantity being scored after post-processing is not a model prediction; it is the postprocessor’s own assignment. Table 7 makes this reduction visible: changing only the postprocessor (b)→(a) raises RMSD(IB)<0.5 Å from 59.2% to 77.1% while RMSD<2 Å stays at 37.7%. Table 4 shows the same pattern: Ours→Ours-p jumps from 60.1% to 79.1% on RMSD(IB)<0.5 Å while RMSD<2 Å remains 37.2%.

full rationale

The circular step is confined to the new RMSD(IB) metric used as headline evidence of covalent-bond precision. Post-processing directly sets the inter-bond distance, and RMSD(IB) measures exactly that distance, so the Ours→Ours-p improvement in Tables 4 and 7 is produced by the postprocessor rather than by the learned pose. Tasks 1 and 2 (reactive location and covalent reaction prediction) are not circular: their baselines are compared against external methods on fixed time-based splits, and the whole-ligand RMSD numbers in Task 3 retain independent content because post-processing does not move the non-bonded ligand atoms. The manual reconstruction of post-reactive ligands is a data-quality risk, not a circularity. Overall, one central prediction reduces by construction, so the benchmark still has substantial independent content but the specific covalent-bond-precision claim is partially circular.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central benchmark claim stands on reconstructed chemical structures, curated covalent-bond annotations, and an empirically motivated distance-minimum constraint. The free parameters are hyperparameters and a data-fitted bond-length distribution used in postprocessing.

free parameters (4)
  • reactive site loss weight alpha = 0.05
    Hand-set weight in Eq. (2) for the reactive-site classification subtask; the paper does not report a sweep.
  • covalent auxiliary loss weight = 1
    Weight for L_cov in Eq. (6), chosen after 'initial experiments' (Appendix B.3); changes the balance of docking accuracy and bond constraint.
  • postprocessing bond length distribution (mu, sigma) = not reported numerically
    Mean and variance of all covalent bond lengths in the dataset, fitted to the data and used to sample the target bond length l' in Section 3.4; the paper does not state whether test samples contribute to these statistics.
  • inference pocket radius = 20 Angstrom
    Hand-chosen radius around the predicted pocket center used to define the pocket for Task 3; the training pocket radius is 10 Angstrom, so this is an additional design choice.
assumptions (4)
  • domain assumption Post-reactive ligands reconstructed from Chemical Component Dictionary alignment and manual patches are chemically correct.
    Appendix A.2 generates Task 2/3 labels this way; if reconstruction is wrong, the benchmark labels are wrong.
  • domain assumption PDB LINK records and CovPDB/CovBinderInPDB annotations correctly identify the biologically relevant covalent bond.
    Dataset curation (Sections 4.1 and A.1) relies on these annotations to define reactive sites, bond atoms, and products.
  • domain assumption The covalent inter-bond distance is the minimum entry of the inter-distance map.
    Section 3.4 and Appendix E justify this with a KS test and 99.6% frequency, then hard-code it into L_cov; this is an empirical regularity, not a physical law.
  • domain assumption A time-based split by PDB deposition date prevents data leakage.
    Section 4.2 uses the EquiBind split; no sequence identity or ligand similarity filter is described, so similar proteins may appear in both train and test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CovDocker: Benchmarking Covalent Drug Design with Tasks, Datasets, and Solutions." pith.science (2026). https://pith.science/paper/GOEFFSYE

@misc{pith2026250621085,
  author       = {Pith},
  title        = {Pith review of: CovDocker: Benchmarking Covalent Drug Design with Tasks, Datasets, and Solutions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GOEFFSYE}},
  note         = {Machine review of arXiv:2506.21085}
}
read the original abstract

Molecular docking plays a crucial role in predicting the binding mode of ligands to target proteins, and covalent interactions, which involve the formation of a covalent bond between the ligand and the target, are particularly valuable due to their strong, enduring binding nature. However, most existing docking methods and deep learning approaches hardly account for the formation of covalent bonds and the associated structural changes. To address this gap, we introduce a comprehensive benchmark for covalent docking, CovDocker, which is designed to better capture the complexities of covalent binding. We decompose the covalent docking process into three main tasks: reactive location prediction, covalent reaction prediction, and covalent docking. By adapting state-of-the-art models, such as Uni-Mol and Chemformer, we establish baseline performances and demonstrate the effectiveness of the benchmark in accurately predicting interaction sites and modeling the molecular transformations involved in covalent binding. These results confirm the role of the benchmark as a rigorous framework for advancing research in covalent drug design. It underscores the potential of data-driven approaches to accelerate the discovery of selective covalent inhibitors and addresses critical challenges in therapeutic development.

Figures

Figures reproduced from arXiv: 2506.21085 by the authors.

Figure 1
Figure 1. Graphical overview of different methodological [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the three proposed tasks in our covalent docking framework. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of our proposed models. All green blocks represent the Uni-Mol [ [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Overview on the data construction of CovDocker with two main steps: Data Collection from Website and Data [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Visualization result of two case study. Structures predicted by Uni-Mol (orange) and ours (green) are placed together [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 9
Figure 9. Figure 9: Reactive Amino Acid Type Statistics (n=2717) [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 7
Figure 7. Figure 7: Reactive Amino Acid Type Statistics (n=2754) [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Reactive Mechanism Type Statistics (n=2717) [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 60 canonical work pages

  1. [1]

    Mohammad Hossein Haghir Ebrahim Abadi, Abdulrahman Ghasemlou, Fate- meh Bayani, Yahya Sefidbakht, Massoud Vosough, Sina Mozaffari-Jovin, and Vladimir N. Uversky. 2024. AI-driven covalent drug design strategies targeting main protease (mpro) against SARS-CoV-2: structural insights and molecular mechanisms.Journal of biomolecular structure & dynamics(2024), 1–29

  2. [2]

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. 2024. Accurate structure prediction of biomolecular interactions with AlphaFold 3.Nature(2024), 1–3

  3. [3]

    Thomas A Baillie. 2016. Targeted covalent inhibitors for drug design.Angewandte Chemie International Edition55, 43 (2016), 13408–13421

  4. [4]

    Berman, John Westbrook, Zukang Feng, Gary Gilliland, T

    Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, T. N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne. 2000. The Protein Data Bank. Nucleic Acids Research28, 1 (2000), 235–242. https://doi.org/10.1093/nar/28.1.235 arXiv:https://academic.oup.com/nar/article-pdf/28/1/235/9895144/280235.pdf

  5. [5]

    Goodsell, and Arthur J

    Giulia Bianco, Stefano Forli, David S. Goodsell, and Arthur J. Olson. 2016. Co- valent Docking Using Autodock: Two-point Attractor and Flexible Side Chain Methods.Protein Science25, 1 (2016), 295–301. https://doi.org/10.1002/pro.2733

  6. [6]

    Chemical Computing Group ULC 910-1010 Sherbrooke St. W. Montreal QC H3A 2R7 Canada. 2024. Molecular Operating Environment (MOE) v2022.02

  7. [7]

    Gabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay, and Tommi Jaakkola

  8. [8]

    Sandro Cosconati, Stefano Forli, Alex L Perryman, Rodney Harris, David S Good- sell, and Arthur J Olson. 2010. Virtual screening with AutoDock: Theory and practice.Expert opinion on drug discovery5, 6 (2010), 597–607

Show all 69 references
  1. [9]

    Hongyan Du, Junbo Gao, Gaoqi Weng, Junjie Ding, Xin Chai, Jinping Pang, Yu Kang, Dan Li, Dongsheng Cao, and Tingjun Hou. 2021. CovalentInDB: A Com- prehensive Database Facilitating the Discovery of Covalent Inhibitors.Nucleic Acids Research49, D1 (2021), D1122–D1129. https://d...

  2. [10]

    Khaled M Elokely and Robert J Doerksen. 2013. Docking challenge: Protein sampling and molecular docking performance.Journal of chemical information and modeling53, 8 (2013), 1934–1945

  3. [11]

    Todd JA Ewing and Irwin D Kuntz. 1997. Critical evaluation of search algo- rithms for automated molecular docking and database screening.Journal of computational chemistry18, 9 (1997), 1175–1189

  4. [12]

    Todd JA Ewing, Shingo Makino, A Geoffrey Skillman, and Irwin D Kuntz. 2001. DOCK 4.0: Search strategies for automated molecular docking of flexible molecule databases.Journal of computer-aided molecular design15 (2001), 411–428

  5. [13]

    Jiyu Fan, Ailing Fu, and Le Zhang. 2019. Progress in molecular docking.Quanti- tative Biology7 (2019), 83–89

  6. [14]

    Berman, and John D

    Zukang Feng, Li Chen, Himabindu Maddula, Ozgur Akcan, Rose Oughtred, He- len M. Berman, and John D. Westbrook. 2004. Ligand Depot: A Data Warehouse for Ligands Bound to Macromolecules.Bioinformatics20 13 (2004), 2153–5

  7. [15]

    Kaiyuan Gao, Qizhi Pei, Jinhua Zhu, Tao Qin, Kun He, Tie-Yan Liu, and Lijun Wu. 2024. FABind+: Enhancing Molecular Docking through Improved Pocket Prediction and Pose Generation.ArXiv preprintabs/2403.20261 (2024)

  8. [16]

    Mingjie Gao and Stefan Günther. 2023. HyperCys: A Structure- and Sequence- Based Predictor of Hyper-Reactive Druggable Cysteines.International Journal of Molecular Sciences24, 6 (2023), 5960. https://doi.org/10.3390/ijms24065960

  9. [17]

    Mingjie Gao, Aurélien F A Moumbock, Ammar Qaseem, Qianqing Xu, and Stefan Günther. 2022. CovPDB: A High-Resolution Coverage of the Covalent Protein– Ligand Interactome.Nucleic Acids Research50, D1 (2022), D445–D450. https: //doi.org/10.1093/nar/gkab868

  10. [18]

    Mathilde Goullieux, Vincent Zoete, and Ute F. Röhrig. 2023. Two-Step Covalent Docking with Attracting Cavities.Journal of Chemical Information and Modeling (2023). https://doi.org/10.1021/acs.jcim.3c01055

  11. [19]

    Courville, and Yoshua Bengio

    Anirudh Goyal, Alex Lamb, Ying Zhang, Saizheng Zhang, Aaron C. Courville, and Yoshua Bengio. 2016. Professor Forcing: A New Algorithm for Training Recurrent Networks. InAdvances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Sys...

  12. [20]

    Vincent Le Guilloux, Peter Schmidtke, and Pierre Tufféry. 2009. Fpocket: An open source platform for ligand pocket detection.BMC Bioinformatics10 (2009), 168 – 168

  13. [21]

    Xiao-Kang Guo and Yingkai Zhang. 2022. CovBinderInPDB: A Structure-Based Covalent Binder Database.Journal of Chemical Information and Modeling62, 23 (2022), 6057–6068. https://doi.org/10.1021/acs.jcim.2c01216

  14. [22]

    Joseph L. Hodges. 1958. The significance probability of the smirnov two-sample test.Arkiv för Matematik3 (1958), 469–486

  15. [23]

    Peter J. Huber. 1964. Robust Estimation of a Location Parameter.Annals of Mathematical Statistics35 (1964), 492–518

  16. [24]

    Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. 2022. Chemformer: A Pre-Trained Transformer for Computational Chemistry.Machine Learning: Science and Technology3, 1 (2022), 015022. https://doi.org/10.1088/2632- 2153/ac3ffb

  17. [25]

    Yize Jiang, Xinze Li, Yuanyuan Zhang, Jin Han, Youjun Xu, Ayush Pandit, Zaixi Zhang, Mengdi Wang, Mengyang Wang, Chong Liu, et al . 2025. PoseX: AI Defeats Physics Approaches on Protein-Ligand Cross Docking.ArXiv preprint abs/2505.01700 (2025)

  18. [26]

    Byrd, Vladimir Tseitin, Dongcheng Dai, Eugene Raush, Maxim Totrov, Ruben Abagyan, Robert Jordan, and Dennis E

    Vsevolod Katritch, Chelsea M. Byrd, Vladimir Tseitin, Dongcheng Dai, Eugene Raush, Maxim Totrov, Ruben Abagyan, Robert Jordan, and Dennis E. Hruby. 2007. Discovery of Small Molecule Inhibitors of Ubiquitin-like Poxvirus Proteinase I7L Using Homology Modeling and Covalent Docki...

  19. [27]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.)

  20. [28]

    Baumgartner, and Carlos J

    David Ryan Koes, Matthew P. Baumgartner, and Carlos J. Camacho. 2013. Lessons Learned in Empirical Scoring with smina from the CSAR 2011 Benchmarking Exercise.Journal of Chemical Information and Modeling53, 8 (2013), 1893–1904

  21. [29]

    2023.Generalized Biomolecular Modeling and Design with RoseTTAFold All-Atom

    Rohith Krishna, Jue Wang, Woody Ahern, Pascal Sturmfels, Preetham Venkatesh, Indrek Kalvet, Gyu Rie Lee, Felix S Morey-Burrows, Ivan Anishchenko, Ian R Humphreys, Ryan McHugh, Dionne Vafeados, Xinting Li, George A Sutherland, Andrew Hitchcock, C Neil Hunter, Minkyung Baek, Fra...

  22. [30]

    Radoslav Krivák and David Hoksza. 2018. P2Rank: Machine Learning Based Tool for Rapid and Accurate Prediction of Ligand Binding Sites from Protein Structure. Journal of Cheminformatics10, 1 (2018), 39. https://doi.org/10.1186/s13321-018- 0285-8

  23. [31]

    Harold W Kuhn. 1955. The Hungarian method for the assignment problem.Naval research logistics quarterly2, 1-2 (1955), 83–97

  24. [32]

    Mark Payne, Colin A

    Mengjie Liu, Alon Grinberg Dana, Matthew Stanley Johnson, Mark Jacob Gold- man, Agnes Jocher, A. Mark Payne, Colin A. Grambow, Kehang Han, Nathan W. Yee, Emily J. Mazeau, Katrin Blondal, Richard H. West, C. Franklin Goldsmith, and William H. Green. 2020. Reaction Mechanism Gen...

  25. [33]

    Ruibin Liu, Joseph Clayton, Mingzhe Shen, Shubham Bhatnagar, and Jana Shen. 2024. Machine Learning Models to Interrogate Proteome-Wide Cova- lent Ligandabilities Directed at Cysteines.JACS Au4, 4 (2024), 1374–1384. https://doi.org/10.1021/jacsau.3c00749

  26. [34]

    Nir London, Rand M Miller, John J Irwin, Oliv Eidam, Lucie Gibold, Richard Bonnet, Brian K Shoichet, and Jack Taunton. 2014. Covalent docking of large libraries for the discovery of chemical probes.Biophysical Journal106, 2 (2014), 264a

  27. [35]

    Jieyu Lu and Yingkai Zhang. 2022. Unified Deep Learning Model for Multitask Reaction Predictions with Explanation.Journal of Chemical Information and Modeling62, 6 (2022), 1376–1387. https://doi.org/10.1021/acs.jcim.1c01467

  28. [36]

    Wei Lu, Qifeng Wu, Jixian Zhang, Jiahua Rao, Chengtao Li, and Shuangjia Zheng

  29. [37]

    Nicolas Moitessier, Joshua Pottel, Eric Therrien, Pablo Englebienne, Zhaomin Liu, Anna Tomberg, and Christopher R. Corbeil. 2016. Medicinal Chemistry Projects Requiring Imaginative Structure-Based Drug Design Methods.Accounts of Chemi- cal Research49, 9 (2016), 1646–1657. http...

  30. [38]

    Tankbind: Trigonometry-aware neural networks for drug-protein binding structure prediction.Advances in neural information processing systems35 (2022), 7236–7249

  31. [39]

    Garrett M Morris, David S Goodsell, Ruth Huey, and Arthur J Olson. 1996. Dis- tributed automated docking of flexible ligands to proteins: parallel applications of AutoDock 2.4.Journal of computer-aided molecular design10 (1996), 293–304

  32. [40]

    Garrett M Morris, David S Goodsell, Robert S Halliday, Ruth Huey, William E Hart, Richard K Belew, and Arthur J Olson. 1998. Automated docking using a Lamarckian genetic algorithm and an empirical binding free energy function. Journal of computational chemistry19, 14 (1998), 1639–1662

  33. [41]

    Garrett M Morris and Marguerita Lim-Wilby. 2008. Molecular docking.Molecular modeling of proteins(2008), 365–382

  34. [42]

    Garrett M Morris, Ruth Huey, and Arthur J Olson. 2008. Using autodock for ligand-receptor docking.Current protocols in bioinformatics24, 1 (2008), 8–14

  35. [43]

    Xuchang Ouyang, Shuo Zhou, Chinh Tran To Su, Zemei Ge, Runtao Li, and Chee Keong Kwoh. 2013. CovalentDock: Automated Covalent Docking with Parameterized Covalent Linkage Energy Estimation and Molecular Geometry Constraints.Journal of Computational Chemistry34, 4 (2013), 326–33...

  36. [44]

    O’Boyle, Michaela S

    Noel M. O’Boyle, Michaela S. Banck, Craig A. James, Chris Morley, Tim Van- dermeersch, and Geoffrey R. Hutchison. 2011. Open Babel: An open chemical toolbox.Journal of Cheminformatics3 (2011), 33 – 33

  37. [45]

    Qizhi Pei, Kaiyuan Gao, Lijun Wu, Jinhua Zhu, Yingce Xia, Shufang Xie, Tao Qin, Kun He, Tie-Yan Liu, and Rui Yan. 2024. FABind: Fast and accurate protein-ligand binding.Advances in Neural Information Processing Systems36 (2024)

  38. [46]

    Abdul-Quddus Kehinde Oyedele, Abdeen Tunde Ogunlana, Ibrahim Damilare Boyenle, Ayodeji Oluwadamilare Adeyemi, Temionu Oluwakemi Rita, Temi- tope Isaac Adelusi, Misbaudeen Abdul-Hammed, Oluwabamise Emmanuel Eleg- beleye, and Tope Tunji Odunitan. 2023. Docking Covalent Targets f...

  39. [47]

    Sereina Riniker and Gregory A. Landrum. 2015. Better Informed Distance Geom- etry: Using What We Know To Improve Conformation Generation.Journal of chemical information and modeling55 12 (2015), 2562–74

  40. [48]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res.21 (2020), 140:1–140:67

  41. [49]

    Tatsuya Sagawa and Ryosuke Kojima. 2023. ReactionT5: A Large-Scale Pre- Trained Model towards Application of Limited Reaction Data

  42. [50]

    JPGLM Rodrigues, JMC Teixeira, M Trellet, and AMJJ Bonvin. 2018. pdb-tools: a swiss army knife for molecular structures [version 1; peer review: 2 approved]. F1000Research7, 1961 (2018). https://doi.org/10.12688/f1000research.17456.1

  43. [51]

    Ferenczy, Stanislav Gobec, and György M

    Andrea Scarpino, László Petri, Damijan Knez, Tímea Imre, Péter Ábrányi-Balogh, György G. Ferenczy, Stanislav Gobec, and György M. Keserű. 2021. WIDOCK: A Reactive Docking Protocol for Virtual Screening of Covalent Inhibitors.Journal of Computer-Aided Molecular Design35, 2 (202...

  44. [52]

    Ferenczy, and György M

    Andrea Scarpino, György G. Ferenczy, and György M. Keserű. 2018. Comparative Evaluation of Covalent Docking Tools.Journal of Chemical Information and Modeling58, 7 (2018), 1441–1458. https://doi.org/10.1021/acs.jcim.8b00228

  45. [53]

    Juswinder Singh, Russell C Petter, Thomas A Baillie, and Adrian Whitty. 2011. The resurgence of covalent drugs.Nature reviews Drug discovery10, 4 (2011), 307–317

  46. [54]

    Christoph Scholz, Sabine Knorr, Kay Hamacher, and Boris Schmidt. 2015. DOCKTITE—A Highly Versatile Step-by-Step Workflow for Covalent Docking and Virtual Screening in the Molecular Operating Environment.Journal of Chem- ical Information and Modeling55, 2 (2015), 398–406. https...

  47. [55]

    Oleg Trott and Arthur J. Olson. 2010. AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading.Journal of Computational Chemistry31, 2 (2010), 455–461

  48. [56]

    Jaakkola

    Hannes Stärk, Octavian Ganea, Lagnajit Pattanaik, Regina Barzilay, and Tommi S. Jaakkola. 2022. EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction. InInternational Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA (Procee...

  49. [57]

    Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. 2022. Diffusers: State-of-the- art diffusion models. https://github.com/huggingface/diffusers

  50. [58]

    Verdonk, Jason C

    Marcel L. Verdonk, Jason C. Cole, Michael J. Hartshorn, Christopher W. Murray, and Richard D. Taylor. 2003. Improved Protein–Ligand Docking Using GOLD. Proteins: Structure, Function, and Bioinformatics52, 4 (2003), 609–623. https: //doi.org/10.1002/prot.10465

  51. [59]

    David Weininger. 1988. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules.Journal of chemical information and computer sciences28, 1 (1988), 31–36

  52. [60]

    Lin Wei, Yaru Chen, Jiaqi Liu, Li Rao, Yanliang Ren, Xin Xu, and Jian Wan

  53. [61]

    https://doi

    Cov_DOX: A Method for Structure Prediction of Covalent Protein–Ligand Bindings.Journal of Medicinal Chemistry65, 7 (2022), 5528–5538. https://doi. org/10.1021/acs.jmedchem.1c02007

  54. [62]

    Brooks III

    Yujin Wu and Charles L. Brooks III. 2022. Covalent Docking in CDOCKER. Journal of Computer-Aided Molecular Design36, 8 (2022), 563–574. https://doi. org/10.1007/s10822-022-00472-3

  55. [63]

    Chang Wen, Xin Yan, Qiong Gu, Jiewen Du, Di Wu, Yutong Lu, Huihao Zhou, and Jun Xu. 2019. Systematic Studies on the Protocol and Criteria for Selecting a Covalent Docking Tool.Molecules24, 11 (2019), 2183. https://doi.org/10.3390/ molecules24112183

  56. [64]

    Qilong Wu and Sheng-You Huang. 2023. HCovDock: An Efficient Docking Method for Modeling Covalent Protein–Ligand Interactions.Briefings in Bioinformatics 24, 1 (2023), bbac559. https://doi.org/10.1093/bib/bbac559

  57. [65]

    Kai Zhu, Kenneth W Borrelli, Jeremy R Greenwood, Tyler Day, Robert Abel, Ramy S Farid, and Edward Harder. 2014. Docking covalent inhibitors: A parameter free approach to pose prediction and scoring.Journal of chemical information and modeling54, 7 (2014), 1932–1940

  58. [66]

    Xujun Zhang, Odin Zhang, Chao Shen, Wanglin Qu, Shicheng Chen, Hanqun Cao, Yu Kang, Zhe Wang, Ercheng Wang, Jintu Zhang, et al. 2023. Efficient and accurate large library ligand docking with KarmaDock.Nature Computational Science3, 9 (2023), 789–804

  59. [67]

    Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. 2023. Uni-Mol: A Universal 3D Molecular Representation Learning Framework. (2023)

  60. [69]

    Publicly Accessible

    Kai Zhu, Kenneth W. Borrelli, Jeremy R. Greenwood, Tyler Day, Robert Abel, Ramy S. Farid, and Edward Harder. 2014. Docking Covalent Inhibitors: A Param- eter Free Approach To Pose Prediction and Scoring.Journal of Chemical Infor- mation and Modeling54, 7 (2014), 1932–1940. htt...

  61. [2022]

    DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.