{"id":"03d1efc4-d63d-4058-a07e-63ad9fa10d85","arxiv_id":"2506.01302","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that categorizes recent GNN-based methods for drug discovery tasks and lists benchmark datasets, without presenting new experimental results.","lead":"This paper surveys how graph neural networks (GNNs) are used in drug discovery, covering molecule generation, property prediction, and drug-drug interaction prediction. It organizes recent methods (mostly since 2021) into taxonomies and lists common benchmark datasets and code repositories.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central 'dominant paradigm' claim is not established: the review is GNN-only and its molecule-generation section misclassifies conformer-generation papers as de novo generation.","rationale":"Reader's verdict is CONDITIONAL with high confidence, on grounds that taxonomy and reference list have internal inconsistencies. I agree those inconsistencies exist, and the DeepDDS/SQUID examples are valid. My stress-test adds a more load-bearing issue: the concluding dominance claim is an empirical statement about the field, but the study design—a GNN-only survey with no comparative or quantitative baseline—cannot support it. A survey that assembles only GNN success stories cannot demonstrate that GNNs dominate the broader literature. The misclassification of ConfVAE and V onMisesNet as unconstrained molecule generation is a concrete instance of the evidence being overcounted: both generate conformers, not drug candidates, so the 'molecule generation' pillar of the dominance claim is partly built on non-examples. This is internal to the paper (Section 1.1's definition vs Section 3.1's contents), so it does not depend on outside consensus. The recommended verdict remains CONDITIONAL because the paper is still a usable catalog of GNN methods; the condition should include weakening or carefully qualifying the dominance claim and correcting the Table 1/Appendix misclassifications. I do not advocate REJECT: the taxonomy errors are fixable and the survey has standalone value as a reference. The concrete test proposed would settle whether the dominance claim and the §3.1 categorization hold.","tokens_in":26161,"tokens_out":6653,"duration_ms":67429,"concrete_test":"Manually or via a filtered search of DBLP/OpenReview, collect papers from ICML, NeurIPS, ICLR, KDD, and AAAI (2021–2025) matching the three task families; classify each as GNN-based, non-GNN, or hybrid, and for molecule generation classify whether the method generates new molecules or only conformations of known molecules. If the share of GNN-based methods among top-performing systems is below 50%, or if a substantial fraction of Section 3.1's examples are conformation-generation papers, the dominance claim and the taxonomy are both unsupported. As a minimal check, read the abstracts of references [6], [7], and [15] and compare their stated tasks with their Table 1 entries.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest concluding claim (Section 7) is that 'molecular graph representation learning has become a dominant paradigm' across molecule generation, property prediction, and DDI prediction. The only evidence offered is the set of papers the authors chose to review, all of which are GNN-based; no comparison is made with non-GNN approaches (e.g., sequence-based transformers, diffusion or flow models, or LLM-based methods) that have also advanced these tasks. This is selection bias: a review that only includes GNN papers cannot, by itself, demonstrate dominance. The problem becomes visible inside the selection. Section 3.1 places ConfVAE [6] and V onMisesNet [7] under 'unconstrained generation' of molecules, but both papers generate 3D conformations of existing molecules, not novel chemical entities; Section 1.1 defined molecule generation as producing 'valid but different molecules' and evaluation in terms of validity and variety. Including conformation-generation work inflates the molecule-generation evidence for GNN dominance. The taxonomy reliability is also questionable elsewhere: Table 1 lists SQUID as 'Ligand-Protein Based Generation' while Section 3.3 describes it as shape-conditioned and target-agnostic, and Appendix Table 3 attaches 'DeepDDS [109]' where [109] is SSF-DDI. These are not merely cosmetic: the survey's only novel contribution is its organization, and the central claim draws on that organization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of graph neural network methods applied to three drug-discovery tasks: molecule generation, molecular property prediction, and drug-drug interaction prediction. It organizes recent GNN-based methods by input type and task, catalogs commonly used benchmark datasets, provides a table of code URLs, and concludes that molecular graph representation learning has become a dominant paradigm across these applications.","tokens_in":26454,"tokens_out":4233,"duration_ms":47090,"significance":"If its organizational and bibliographic claims were accurate, the survey would serve as a useful entry point to GNN-based drug discovery by assembling a broad set of methods, datasets, and code links in one place. The paper makes no empirical claims and contains no derivations, and it offers no machine-checked artifacts; its contribution is purely taxonomic and bibliographic. That makes the correctness of its categorization and reference list load-bearing. The strengths are the breadth of coverage, the inclusion of code URLs, and the concise descriptions of benchmark datasets.","major_comments":[{"comment":"The central concluding claim that molecular graph representation learning has 'become a dominant paradigm' is not established by the evidence presented. The review includes only GNN-based methods, by design, and contains no comparison with or systematic discussion of non-GNN approaches (e.g., sequence-based transformers, diffusion or flow models, or LLM-based methods) that have also advanced molecule generation, property prediction, and interaction prediction. To support the word 'dominant,' the authors would need to broaden the survey to include representative non-GNN methods or explicitly qualify the claim as a statement about trends within GNN-based work.","section":"Section 7"},{"comment":"ConfVAE [6] and von MisesNet [7] are presented as unconstrained molecule generation methods, but both papers address molecular conformation generation (generating 3D geometries of a given molecular graph), not generation of novel chemical entities. This conflicts with the definition in Section 1.1, where molecule generation produces 'valid but different molecules' evaluated by validity and variety. Placing these methods under unconstrained generation inflates the molecule-generation evidence for GNNs, and they should be moved to a separate conformation-generation category or explicitly reframed.","section":"Section 3.1"},{"comment":"SQUID is listed under 'Ligand-Protein Based Generation,' but Section 3.3 describes it as shape-conditioned generation from an input 3D shape treated as a point cloud, with no protein binding site as input. This is internally inconsistent and misleading for readers who use the table as a guide. The row should be reclassified or the description in Section 3.3 should be corrected.","section":"Table 1, SQUID row"},{"comment":"The reference number for DeepDDS is inconsistent: Section 5.3 and Table 1 cite it as [107], while Appendix Table 3 lists 'DeepDDS [109]' even though reference [109] is SSF-DDI and the SSF-DDI URL appears on the same line. This is a concrete citation error in a survey whose main contribution is bibliographic organization, and the numbering in the appendix should be checked and corrected.","section":"Appendix Table 3 vs. Section 5.3"}],"minor_comments":[{"comment":"References [6] and [61] are the same paper (Xu et al., ICML 2021) but are listed as two distinct entries; one should be removed or cross-referenced.","section":"References"},{"comment":"The sentence 'by appling NLP algorithms' contains a typo and should read 'by applying NLP algorithms.'","section":"Section 6.4.2"},{"comment":"The sentence 'the two database' should read 'the two databases.'","section":"Section 6.1"},{"comment":"The count '1.440 protein-nucleic acid complex' should be written as '1,440 protein-nucleic acid complexes.'","section":"Section 6.4.1"},{"comment":"The name 'V onMisesNet' appears with a stray space in several places; it should be 'von MisesNet' to match the cited work.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is suitable for a journal or venue that publishes surveys, and the main revisions concern the taxonomy, the overreach of the 'dominant paradigm' claim, and reference consistency. The self-citations are relevant to the topics covered and do not appear to distort the survey's organization."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a competent survey with real practical value, but it overclaims in its conclusion and has a handful of taxonomy errors that need fixing before I'd trust it as a map of the field.\n\nWhat's actually good: The paper collects recent GNN-based work across three tasks—molecule generation, property prediction, DDI prediction—into one place. The tables (Table 1 for models, Table 2 for datasets, Table 3 for code URLs) are genuinely useful, especially for a newcomer. The appendix of code URLs is a nice touch and makes the survey actionable. The prose is clear, and the descriptions of individual methods are mostly accurate.\n\nWhere it's soft: The central claim in Section 7 that 'molecular graph representation learning has become a dominant paradigm' is not established by the evidence. The review only includes GNN-based papers; it never benchmarks against sequence-based Transformers, diffusion models, or LLM-based methods, so it cannot justify a claim about dominance. That's a selection-bias problem, not just a wording quibble. The taxonomy has some real mistakes. ConfVAE and vonMisesNet are placed under 'unconstrained generation' of novel molecules, but both are 3D conformation generation models—they generate conformations for a given molecule, not new molecules. That inflates the molecule-generation evidence. Similarly, SQUID is listed as 'Ligand-Protein Based Generation' in Table 1, but Section 3.3 correctly describes it as shape-conditioned and target-agnostic. And the DeepDDS citation is [107] in the text and Table 1, but [109] in the Appendix Table 3, where [109] is actually SSF-DDI. These are not fatal to the whole review, but they undermine the organizational value that is the paper's only novel contribution.\n\nAlso, the paper uses 'state-of-the-art' without quantified comparisons in several places—minor for a survey, but it should be rephrased.\n\nIf these issues are addressed, this could be a serviceable reference. As is, I'd want a major revision before putting it in front of readers. I'd still send it to peer review rather than desk-reject—surveys this current and this organized deserve referee time, and the errors are fixable.\n\nOverall: worth engaging with, but the authors need to either narrow the claims or broaden the scope.\n\nBest,\n[You]","headline":"Useful but uneven survey of GNNs for drug discovery; the 'dominant paradigm' conclusion overreaches, and a few taxonomy errors need fixing.","tokens_in":26924,"tokens_out":3546,"would_cite":false,"duration_ms":35720,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that molecular graph representation learning has become the dominant paradigm across molecule generation, property prediction, and drug-drug interaction prediction.","keywords":["drug discovery","graph neural networks","molecular graph representation learning","molecule generation","molecular property prediction","drug-drug interaction prediction","self-supervised learning","benchmark datasets"],"falsifier":"A reader could reproduce the paper's tables from the original sources: if a substantial share of recent state-of-the-art methods in the three tasks are not GNN-based, or if the table's classifications do not match the cited papers (for instance, SQUID is listed as protein-ligand generation although the text describes it as shape-conditioned and target-agnostic, and DeepDDS is cited as [107] in the text but [109] in the appendix), the taxonomy and the dominance claim would need revision.","tokens_in":26002,"feed_emoji":"🧪","tokens_out":6406,"duration_ms":54675,"temperature":0.7,"pith_summary":"This review tries to establish that graph neural networks have become the dominant computational approach in drug discovery, uniting molecule generation, molecular property prediction, and drug-drug interaction prediction under molecular graph representation learning. It surveys work published since 2021, organizing methods by input type (2D graphs, 3D graphs, multi-molecule complexes) and by task, and collects the benchmark datasets that support these applications. The authors argue that GNNs naturally capture intricate molecular patterns, and that recent advances such as self-supervised pre-training, few-shot learning, and multi-modal fusion have made these models more accurate and generalizable. If the claim holds, new drug discovery systems will be built around GNNs operating on molecular graphs, with data quality, interpretability, and multi-omics integration as the main remaining bottlenecks.","feed_headline":"Graph neural networks now anchor drug discovery modeling","feed_subtitle":"A survey shows GNNs unify molecule generation, property prediction, and drug-drug interaction work.","key_machinery":"The central object is the molecular graph, with atoms as nodes and bonds as edges, optionally augmented with three-dimensional coordinates and auxiliary information. GNN message passing aggregates information from neighboring nodes to produce latent node and molecular representations, which are then used for generation, property prediction, or interaction prediction. The review's organizing machinery is a taxonomy that crosses input types (2D graphs, 3D graphs, multi-molecule complexes) with tasks and learning strategies such as self-supervised pre-training and few-shot learning; this taxonomy carries the argument by grouping all surveyed methods and revealing the field's common trends.","core_discovery":"The paper's central claim is that molecular graph representation learning has become a dominant paradigm in computational drug discovery, with evidence drawn from three application areas: molecule generation, molecular property prediction, and drug-drug interaction prediction. It argues that GNNs, by operating directly on molecular graphs rather than one-dimensional string representations, capture chemical structure and complex molecular patterns more effectively, and that recent progress has shifted toward 3D molecular graphs, pre-trained and few-shot learning strategies, and multi-modal input fusion. The review presents its taxonomy of methods and benchmark datasets as a map of this emerging mainstream, and identifies data scarcity, model interpretability, and integration of multi-omics data as the central open challenges.","pith_inferences":["The 'dominant paradigm' claim is an assessment of the reviewed literature rather than a quantitative comparison; a systematic benchmark covering both GNN and non-GNN methods on common datasets would test it directly.","The taxonomy's internal inconsistencies (SQUID listed as protein-ligand based despite being described as shape-conditioned and target-agnostic, and the DeepDDS citation appearing as [107] in the text but [109] in the appendix) suggest readers should verify details against the original papers before using the tables as a guide.","The emphasis on pre-training and multi-modal fusion hints that future GNN drug discovery models may integrate sequence, graph, and 3D geometry into a single pretraining framework, an extension the authors mention but do not develop.","If interpretability remains unresolved, adoption of GNNs in regulated pharmaceutical settings may depend on explainable model variants as much as on predictive accuracy."],"forward_implications":["If molecular graph representation learning is dominant, new drug discovery systems will likely treat molecular graphs as the default input representation for generation and prediction tasks.","The shift toward 3D molecular graphs implies that property and interaction prediction will continue to adopt geometry-aware architectures that exploit distances, angles, and conformations.","Self-supervised pre-training and few-shot learning will remain standard tools for coping with scarce labeled data, especially for toxicity and clinical endpoints.","Drug-drug interaction prediction will increasingly incorporate patient medical histories to produce personalized and context-aware medication recommendations.","Data heterogeneity, interpretability, and multi-omics integration are the next frontiers that GNN-based drug discovery must address to move from benchmarks to real-world use."],"supporting_citations":[{"why":"GraphINVENT supplies the early demonstration of iterative GNN-based molecular graph generation.","marker":"[56]"},{"why":"MoLeR provides the motif-based constrained generation approach that anchors the constrained generation category.","marker":"[8]"},{"why":"Pocket2Mol is a representative protein-ligand based generation model operating on 3D protein pockets.","marker":"[13]"},{"why":"MGSSL exemplifies motif-based self-supervised pre-training for molecular property prediction.","marker":"[63]"},{"why":"GraphMVP supports the 2D-3D consistency pretraining trend for geometry-aware property prediction.","marker":"[75]"},{"why":"MGraphDTA is a GNN-based drug-target binding affinity predictor combining compound graphs and protein sequences.","marker":"[27]"},{"why":"DeepDDS is the attention-based GNN model cited for synergistic drug-drug interaction prediction.","marker":"[107]"},{"why":"SafeDrug is the personalized drug combination recommendation model that incorporates patient histories with molecular graphs.","marker":"[40]"},{"why":"MoleculeNet supplies the standard benchmark datasets used to evaluate molecular property prediction models.","marker":"[21]"},{"why":"PDBbind provides experimentally determined binding affinities and structures that ground protein-ligand interaction benchmarks.","marker":"[126]"}],"fun_headline_variants":["GNNs turn drug discovery from strings to graphs","Graph neural nets become drug discovery's new backbone","Survey: GNNs dominate molecular modeling in drug discovery","How GNNs reshape drug design and prediction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's central claim depends on the assumption that its selected papers and their categorization accurately represent the primary literature on GNNs in drug discovery; missed or misclassified works could make the dominance claim overstate the field's actual state.","fun_headline_variants_meta":{"raw":{"variants":["GNNs turn drug discovery from strings to graphs","Graph neural nets become drug discovery's new backbone","Survey: GNNs dominate molecular modeling in drug discovery","How GNNs reshape drug design and prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1108,"prompt_tokens":776,"completion_tokens":332,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":392,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":392,"tokens_out":332,"duration_ms":4167,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:43:54.891381+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could reproduce the paper's tables from the original sources: if a substantial share of recent state-of-the-art methods in the three tasks are not GNN-based, or if the table's classifications do not match the cited papers (for instance, SQUID is listed as protein-ligand generation although the text describes it as shape-conditioned and target-agnostic, and DeepDDS is cited as [107] in the text but [109] in the appendix), the taxonomy and the dominance claim would need revision.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GraphINVENT supplies the early demonstration of iterative GNN-based molecular graph generation."},{"cited_title":"& Lee, C.-K","cited_arxiv_id":null,"evidence_quote":"MGSSL exemplifies motif-based self-supervised pre-training for molecular property prediction."},{"cited_title":"InInternational Confer- ence on Learning Representations (ICLR)(OpenReview","cited_arxiv_id":null,"evidence_quote":"GraphMVP supports the 2D-3D consistency pretraining trend for geometry-aware property prediction."},{"cited_title":"& Liu, H","cited_arxiv_id":null,"evidence_quote":"DeepDDS is the attention-based GNN model cited for synergistic drug-drug interaction prediction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PDBbind provides experimentally determined binding affinities and structures that ground protein-ligand interaction benchmarks."}],"review_version":1}