{"id":"0c13b37a-57bd-4d92-9007-32f065e08376","arxiv_id":"2412.02889","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A fair benchmark shows conventional docking programs beat DiffDock, and DiffDock's remaining accuracy mostly comes from cases with near-duplicate training structures.","lead":"This paper reruns DiffDock's own benchmarks with standard docking software and finds the older tools win. It also shows DiffDock's successes cluster on test cases with near-identical training examples, suggesting memorization rather than general learning.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'table-lookup' conclusion lacks a control: no conventional docking method is evaluated on the same near-neighbor/hard split, so the 40-point gap may reflect task difficulty rather than memorization.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but the weakest link is not only the ad hoc thresholds. The paper's most striking claim - that DiffDock's success is 'essentially a form of table-lookup' - depends on showing a performance dichotomy that is specific to DiffDock. The presented evidence is purely correlational: DiffDock does well on cases with near neighbors and poorly on cases without. Without a control group of conventional docking methods on the same split, this dichotomy could simply reflect the well-known fact that some docking targets are harder than others. The missing control is concrete and checkable with the data the authors already provide. If conventional methods also drop sharply on the hard subset, the 'table-lookup' claim would need to be withdrawn, although the main comparative result (conventional docking outperforms DiffDock on the full Clean Test Set) may still stand. Thus the final verdict remains CONDITIONAL pending this additional analysis, and I recommend no change to the reader's verdict.","tokens_in":18525,"tokens_out":5369,"duration_ms":52155,"concrete_test":"Using the authors' published near-neighbor lists (191 easy, 99 hard) and the provided data archive, recompute Surflex-Dock, Glide, and Vina Top-1/Top-5 success rates at 2.0 Å RMSD separately on the easy and hard subsets. If conventional methods also exhibit a large performance drop on the hard subset (e.g., Top-1 near 21/28%), the memorization claim fails; if conventional methods maintain substantially higher success on the hard subset while DiffDock drops to 21/28%, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference in Section 2.3 is that DiffDock's success is 'inextricably linked' to near-neighbor training cases (191/290 easy, 99/290 hard), with Top-1/Top-5 success rates of 57/65% versus 21/28% at 2.0 Å RMSD (Figure 7). From this, the paper concludes that DiffDock has 'apparently encoded a type of table-lookup.' This causal claim is not supported because the analysis never reports the same easy/hard split for Surflex-Dock, Glide, Vina, or Gnina. The overall conventional success rates (e.g., Surflex-Dock 68/81% on the complete Clean Test Set) are not an adequate control: the hard subset is a different, more difficult population, and conventional docking would also be expected to perform worse on it. Without showing that conventional methods maintain high success on the same 99 hard cases while DiffDock collapses to 21/28%, the 40-point gap is equally consistent with the simple explanation that near-neighbor cases are easier docking problems for any method. The reader's concern about ad hoc similarity thresholds is secondary; even with a robust and externally validated partition, the missing control would still undermine the 'table-lookup' interpretation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks four conventional docking methods (Surflex-Dock, Glide, AutoDock Vina, Gnina) against the deep-learning method DiffDock on the PDBBind 2020 temporal split, using a cleaned 290-case subset of DiffDock's 363-case test set. The authors report that, with a known binding site, Surflex-Dock and Glide succeed at 2.0 Å RMSD at rates of 68/81% and 67/73% (Top-1/Top-5), compared with DiffDock's 45/51%, and that Vina and Gnina also outperform DiffDock. In the blind-docking condition, Surflex-Dock shows a smaller advantage. The paper's second main claim is that DiffDock's success is concentrated in the 191/290 test cases that have near-neighbor protein-ligand complexes in the training set, with success rates of 57/65% on those versus 21/28% on the 99 remaining 'hard' cases, leading the authors to conclude that DiffDock has 'apparently encoded a type of table-lookup' and that its reported results are artifactual.","tokens_in":18776,"tokens_out":9831,"duration_ms":129573,"significance":"This is a timely and potentially useful benchmark contribution. The authors provide fully automatic, scripted workflows for four conventional docking methods on a cleaned version of the DiffDock test set, together with a public data archive, which is a valuable reproducibility resource for the docking community. The known-site results for conventional methods, if properly compared, would be an important reference point. The near-neighbor analysis raises a legitimate concern about temporal-split benchmarking at an extreme 98/2 ratio. However, the central memorization conclusion is not yet supported because the analysis lacks a critical control, and the known-site comparison is asymmetric. If the missing analyses are provided, the paper could become a significant corrective to optimistic assessments of deep-learning docking; in its current form, the evidence is incomplete.","major_comments":[{"comment":"The 'known binding site' comparison is asymmetric. The DiffDock results shown are those from the original DiffDock report, which used a blind docking protocol, as the paper itself notes in §2.2 for the corresponding Glide baseline. Surflex-Dock, Glide, Vina, and Gnina, by contrast, are given the binding site defined by the cognate ligand. The paper never reports DiffDock's performance when it also receives the pocket as input. Consequently, the headline gap of 68/81% versus 45/51% conflates method quality with the amount of information provided. To support the claim that conventional methods 'far exceed' DiffDock in the known-site condition, the authors should either run DiffDock in its pocket-conditioned mode or explicitly frame the comparison as 'conventional docking with known site versus DiffDock blind' and temper the wording accordingly.","section":"§2.2.1, Figure 2 (left)"},{"comment":"The conclusion that DiffDock's success is 'inextricably linked' to near-neighbor training cases and that it 'has apparently encoded a type of table-lookup' is not supported because no conventional docking method is evaluated on the same near-neighbor/hard split. The 57/65% versus 21/28% success rates for DiffDock are equally consistent with the simple explanation that near-neighbor cases are easier docking problems for any method. The overall Clean Test Set success rates reported for Surflex-Dock and Glide do not serve as the needed control, since they mix the two populations. The paper should report the same split for Surflex-Dock, Glide, Vina, and Gnina, all of which have already been run on the Clean Test Set, and test whether the performance gap between the near-neighbor and hard subsets is significantly larger for DiffDock than for conventional methods.","section":"§2.3, Figure 7"},{"comment":"The classification into 191 near-neighbor and 99 hard cases depends on ad hoc similarity thresholds: Gsim in the top 1% or ≥0.3, followed by PSIM ≥0.65, with 'extreme' cases defined by Gsim ≥0.80. These thresholds are not validated against an external benchmark, and no sensitivity analysis is provided. Since the roughly 40-percentage-point performance gap is the central evidence for the memorization claim, the authors should vary the thresholds (for example, PSIM from 0.5 to 0.8 and alternative Gsim cutoffs) and show that the easy/hard partition and the performance gap are stable, or report how the conclusions change under reasonable alternative definitions.","section":"§2.3 and Appendix 'Finding Near-Neighbor Training Cases'"},{"comment":"The claim that Surflex-Dock is statistically superior to DiffDock in blind docking is based on a paired t-test computed on a post-hoc subset of 160/290 cases where both methods achieved Top-5 RMSD ≤4.0 Å. Selecting the subset based on outcomes biases the test and invalidates the p-values as a statement about the full test set. The authors should provide a pre-specified analysis on all 290 cases, for example by treating non-converged dockings as 20 Å or using a non-parametric test on ranked RMSDs, or should explicitly limit the statistical claim to the 160-case subset and present the full-set comparison descriptively.","section":"§2.2.1, unknown binding site"}],"minor_comments":[{"comment":"The paper should state explicitly whether the DiffDock RMSD values shown for the Clean Test Set were recomputed by the authors or provided by the DiffDock team, and confirm that the same RMSD definition was used for DiffDock and for the conventional methods.","section":"§2.1, Figure 1 caption"},{"comment":"The paired t-test is applied to binary success indicators at a single RMSD threshold; McNemar's test or a bootstrap of the paired difference would be more appropriate for these binary outcomes, even though the large gaps likely remain significant.","section":"§2.2.1"},{"comment":"The statement that the Surflex-Dock approach 'made no use of prior information from pre-2019 structures' is overstated, since the Surflex-Dock scoring function is empirically parameterized on protein-ligand complexes. The more defensible claim is that no specific pre-2019 co-crystal structures near to the test cases were used for binding-site identification or pose selection.","section":"§2.2.1"},{"comment":"The exclusion of identical training ligands from the near-neighbor definition is not explained. If any test complex has a fully identical protein-ligand training complex, that would be the clearest possible memorization case, and its handling should be stated explicitly.","section":"§2.3, Appendix"},{"comment":"The recommendation that a 25/75% temporal split is 'more reasonable' is presented as a general benchmarking guideline without supporting evidence beyond a citation to a prior paper; it would benefit from being framed as an opinion or from being supported by additional analysis.","section":"§4, Conclusions"},{"comment":"The Addendum notes the later release of DockGen and DiffDock-L, but the main text does not discuss whether those developments affect the near-neighbor criticism. A brief discussion of whether the same issue persists in DiffDock-L would strengthen the paper's relevance.","section":"§5, Addendum"}],"recommendation":"major_revision","confidential_remarks":"The authors have a commercial affiliation with Optibrium, the developer of Surflex-Dock, and this potential competing interest is not declared in the manuscript; I would ask the authors to add a conflict-of-interest statement. The manuscript is more of a brief perspective/benchmark than a full research article, and it may be better suited to a venue that publishes such contributions. The revisions requested in the major comments are essential before the paper's central claims about 'table-lookup' and the known-site comparison can be considered supported; however, they are feasible within the scope of the current manuscript because the authors already have the computational results needed for the missing control and sensitivity analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: you should know two things. First, this paper delivers a genuinely useful, reproducible re-benchmark: cleaned 290-case test set from DiffDock's own set, with automated protocols for Surflex-Dock, Glide, Vina, and Gnina, and a data archive. The known-site results are clear and large: Surflex-Dock 68/81%, Glide 67/73%, Vina 47/73%, Gnina also above DiffDock, with p-values <1e-7. That resets the baseline claim from the original DiffDock paper.\n\nSecond, the headline interpretation, that DiffDock 'encoded a type of table-lookup,' is not established. The near-neighbor analysis splits the test set into 191 easy and 99 hard cases, and shows DiffDock collapses from 57/65% to 21/28%. But the same split is never computed for any conventional method. Near-neighbor cases are probably easier for every dock-er; without showing that Surflex-Dock or Glide maintain high success on the same 99 hard cases, the 40-point gap is equally consistent with task difficulty as with memorization. The thresholds for Gsim and PSIM are also ad hoc, but even a robust partition wouldn't fix the missing control.\n\nLesser soft spots: the blind-docking statistical test is done on a post-hoc 160-case subset where both methods were within 4Å, which is selection; and RMSD comparability between DiffDock's reported values and the re-processed proteins is not fully discussed. Neither changes the main known-site result.\n\nWhat's genuinely new: the cleaned test set and the systematic near-neighbor mapping. The paper is honest about its own protocols and ships code and data, which makes the central benchmark independently checkable. The citation pattern is appropriate; no self-serving citation beyond the methods' own provenance.\n\nFor docking practitioners and benchmark developers, this is a useful resource: the cleaned set and scripts lower the cost of checking any future docking claim. If this were submitted to me, I'd send it to peer review. The core empirical comparison is valuable and likely correct; the memorization conclusion is overreached and should be either qualified or tested with the missing control. A solid paper with one weak interpretive section, not a fatal flaw.","headline":"A useful, reproducible re-benchmark that puts DiffDock's reported advantage in doubt, but the 'table-lookup' claim overreaches because the near-neighbor split is never run on conventional baselines.","tokens_in":19329,"tokens_out":3158,"would_cite":true,"duration_ms":28328,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiffDock's reported docking superiority is largely an artifact of near-neighbor memorization from its training set, and conventional docking workflows outperform it when run properly.","keywords":["DiffDock","molecular docking","deep learning","memorization","near-neighbor analysis","Surflex-Dock","PDBBind","benchmark contamination"],"falsifier":"Delete from DiffDock's training set all complexes that meet the near-neighbor criteria for the 290 clean test cases, retrain, and re-dock; if success on the previously 'hard' cases rises from about 21% toward the 57% seen on near-neighbor cases, the table-lookup claim is refuted, whereas if it stays near 21%, the claim is confirmed.","tokens_in":18310,"feed_emoji":"💊","tokens_out":6130,"duration_ms":56729,"temperature":0.7,"pith_summary":"This paper argues that the much-publicized superiority of the deep-learning docking method DiffDock over conventional docking programs is an artifact. The authors re-ran the DiffDock benchmark with fully automatic conventional docking tools (Surflex-Dock, Glide, Vina, Gnina) and found that the conventional tools outperform DiffDock on the same test set, both when the binding site is given and when it must be found automatically. The central explanatory claim is that DiffDock memorized close to identical protein-ligand complexes from its 17,000-structure training set: on the two-thirds of test cases with a near-neighbor training analog, its success rate was 57/65%, but on the one-third without such an analog it fell to 21/28%. The authors conclude that DiffDock's performance is closer to table-lookup than to generalizable docking.","feed_headline":"Deep-learning docking wins vanish without near-neighbor training data","feed_subtitle":"Retested on cases with no near-identical training analog, DiffDock's success drops by roughly 40 points and conventional docking wins.","key_machinery":"The argument is carried by a contamination diagnostic that splits the test set by near-neighbor status. A training case counts as a near neighbor when the 2D ligand topological similarity Gsim is in the top 1% or ≥0.3 and the binding-pocket similarity PSIM is ≥0.65, with Gsim ≥0.80 defining an extreme subset. The complementary machinery is a fully automatic conventional docking pipeline, including automated protein and ligand preparation, protomol-based site definition, docking, pose-family clustering, and automorph-corrected RMSD calculation, applied to a Clean Test Set of 290 complexes curated from DiffDock's 363-case test set to ensure the conventional baselines use the methods as intended.","core_discovery":"The paper's central claim is that DiffDock's reported performance cannot be taken as evidence that deep learning solves molecular docking. Recomputing DiffDock's own benchmark with a fully automatic Surflex-Dock workflow shows conventional docking is far better at redocking cognate ligands when the binding site is known (Top-1/Top-5 success 68/81% vs. 45/51% at 2.0 Å RMSD) and no worse when the site is unknown. The decisive new result is a diagnostic: 191 of the 290 clean test complexes have a near-neighbor training case, defined by high topological ligand similarity (Gsim) and high binding-pocket similarity (PSIM), and DiffDock's success is dichotomized by that label, 57/65% on near-neighbor cases versus 21/28% on the 99 hard cases, with extreme near-neighbors above 90%. The paper therefore states that DiffDock has apparently encoded a type of table-lookup and that its comparisons to other methods used those methods in a nonstandard, blinded way that handicapped them.","pith_inferences":["The Gsim/PSIM contamination audit could be adopted as a standard reporting metric for any machine-learned docking or structure-prediction model, making memorization effects visible without additional experiments.","Because the test set is drawn from PDBBind complexes released after 2019, other deep-learning methods trained on the same pre-2019 PDBBind data may show the same split; re-running this audit on them is a direct next test.","The paper's addendum notes that DiffDock-L and a new DockGen benchmark appeared after its analyses; applying this same near-neighbor audit to DiffDock-L would reveal whether its claimed generalization to new pockets is real or a similar artifact.","If this finding is correct, reported state-of-the-art docking accuracy in the deep-learning literature may be substantially overstated, and progress should be judged on truly novel-scaffold and novel-pocket test sets."],"forward_implications":["DiffDock's apparent gains over conventional docking do not survive a properly run comparison: conventional tools win by 20–30 points at the 2.0 Å threshold in the known-site condition.","Benchmark claims for learning-based docking methods should be audited for near-neighbor contamination before being read as evidence of generalization.","Temporal splits of 98% training / 2% testing are too skewed to prevent memorization from dominating results; more balanced splits (the paper suggests 25/75) would better reveal true capabilities.","The roughly 40-point gap between near-neighbor and hard cases means predictions for targets and ligands without near-identical precedents are the real challenge, and on those DiffDock is weak.","When binding sites are unknown, a conventional automated workflow with pocket detection remains competitive or better, so blind-docking comparisons should be run with the conventional method's own site-finding tools."],"supporting_citations":[{"why":"Defines DiffDock, its training set, and the reported results that this paper re-benchmarks.","marker":"[1]"},{"why":"Supplies the expected 60–80% cognate redocking success for mature docking methods that frames the comparison.","marker":"[4]"},{"why":"Describes the knowledge-guided Surflex-Dock protocol used to build the Clean Test Set and the automatic pipeline.","marker":"[3]"},{"why":"Independent evidence that AI docking methods produce strained poses and fail to generalize, supporting the contamination interpretation.","marker":"[2]"},{"why":"Original Surflex-Dock method paper, the engine for the primary conventional baseline.","marker":"[5]"},{"why":"Supplies the binding-site surface similarity and pocket-finding method used in the unknown-site blind docking condition.","marker":"[13]"}],"fun_headline_variants":["DiffDock's edge vanishes without near-neighbor training data","Conventional docking wins when deep learning gets fair test","Deep-learning docking exposed as table lookup on benchmark","DiffDock's success drops 40 points on unfamiliar protein-ligand pairs","Surflex-Dock outdoes DiffDock in fair docking benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The memorization conclusion rests on the chosen similarity thresholds (Gsim top 1% or ≥0.3 and PSIM ≥0.65) that split the test set into near-neighbor and hard cases; if those cutoffs do not validly capture 'near-identical', the 40-point gap and the table-lookup reading could be overstated.","fun_headline_variants_meta":{"raw":{"variants":["DiffDock's edge vanishes without near-neighbor training data","Conventional docking wins when deep learning gets fair test","Deep-learning docking exposed as table lookup on benchmark","DiffDock's success drops 40 points on unfamiliar protein-ligand pairs","Surflex-Dock outdoes DiffDock in fair docking benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2774,"prompt_tokens":1155,"completion_tokens":1619,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":771,"completion_tokens_details":{"reasoning_tokens":1533}},"tokens_in":771,"tokens_out":1619,"duration_ms":12195,"temperature":1.0,"reasoning_tokens":1533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:58:52.422764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Delete from DiffDock's training set all complexes that meet the near-neighbor criteria for the 290 clean test cases, retrain, and re-dock; if success on the previously 'hard' cases rises from about 21% toward the 57% seen on near-neighbor cases, the table-lookup claim is refuted, whereas if it stays near 21%, the claim is confirmed.","supporting_citations":[{"cited_title":"Surflex-Dock: Docking benchmarks and real-world application","cited_arxiv_id":null,"evidence_quote":"Supplies the expected 60–80% cognate redocking success for mature docking methods that frames the comparison."},{"cited_title":"Knowledge-guided docking: Accurate prospective prediction of bound config- urations of novel ligands using Surflex-Dock","cited_arxiv_id":null,"evidence_quote":"Describes the knowledge-guided Surflex-Dock protocol used to build the Clean Test Set and the automatic pipeline."},{"cited_title":"PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences","cited_arxiv_id":"2308.05777","evidence_quote":"Independent evidence that AI docking methods produce strained poses and fail to generalize, supporting the contamination interpretation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Original Surflex-Dock method paper, the engine for the primary conventional baseline."},{"cited_title":"Cleves, Rocco Varela, and Ajay N","cited_arxiv_id":null,"evidence_quote":"Supplies the binding-site surface similarity and pocket-finding method used in the unknown-site blind docking condition."}],"review_version":1}