{"id":"4cc688dc-5dfb-4152-bd39-349c0b5d308a","arxiv_id":"2504.16152","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of graph-based and heterogeneous-network methods for drug-target interaction prediction published mostly between 2020 and 2024, with datasets, metrics, future challenges, and claimed source code links.","lead":"This paper surveys graph machine learning methods for predicting drug-target interactions, focusing on heterogeneous networks and covering random walk, GNN, metapath, and hybrid approaches. It is a literature review, not a new method, and its value is as a map of recent work and datasets.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's 'comprehensive' and 'wider range' claims are unsupported because selection criteria are absent and known network-based methods are explicitly excluded; a reproducible literature count is needed.","rationale":"The reader's conditional verdict identifies the core weakness: the survey's central utility claim depends on an undocumented and potentially biased paper selection. My reading confirms this and sharpens it. The paper itself says it 'includes a wider range of methods and approaches' than prior reviews [43-46], yet it excludes known network-based methods [31,47-51] without applying a documented scope rule. Since the excluded methods include knowledge-graph and multi-modal frameworks that are close to the surveyed territory, the comparative claim is at risk. This is not an internal inconsistency in any derived result; it is an unsupported completeness claim, which is the load-bearing element for a survey. The proposed concrete test would settle the concern by making the selection reproducible and quantifying omissions. I therefore agree with the reader's conditional verdict and would not change it: the manuscript should be accepted only if the authors document their search and selection protocol and either include or explicitly justify the exclusion of the named network-based methods. No independent evidence such as machine-checked proofs or released code is present, and the unresolved formatting placeholders further support a conditional rather than unconditional acceptance.","tokens_in":26561,"tokens_out":3079,"duration_ms":29828,"concrete_test":"Run a reproducible literature scan: query PubMed/Web of Science for ('drug-target interaction' OR 'drug-target binding affinity') AND ('heterogeneous network' OR 'graph neural network' OR 'metapath' OR 'random walk'), restricted to 2020-2024, screen titles/abstracts with the survey's own stated scope, and compare the resulting eligible set to the surveyed set in Tables 6-10. If the scan identifies eligible network-based DTI methods absent from the survey (e.g., knowledge-graph DTI frameworks such as [49] or [50]) beyond the explicitly excluded list, the 'comprehensive/wider range' claim fails; if all eligible papers are present or accounted for, the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it 'provides comprehensive details' of graph-ML DTI methods and 'includes a wider range of methods and approaches' than prior reviews [43-46]. Both claims rest on the set of included papers being representative of the 2020-2024 literature. The manuscript never states search databases, query terms, inclusion/exclusion criteria, or a screening process, and the methods section explicitly sets aside several recent network-based DTI methods ([31,47-51]) because they are not random-walk, GNN, metapath, or hybrid. Among those exclusions are knowledge-graph-based ([49]), multi-modal representation ([50]), arbitrary-order proximity deep forest ([48]), and multi-label community detection ([31]) approaches. If these fall within the survey's stated scope of network-based graph methods, then the 'wider range' claim is not merely under-documented but likely false relative to reviews that cover such methods. The taxonomy and future-challenges sections could overstate the dominance of the four included families.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of graph machine learning methods for drug-target interaction (DTI) prediction, organized around heterogeneous biological networks. It provides definitions of heterogeneous networks, metapaths, and GNN components; describes benchmark datasets (Yamanishi, Luo, Davis, KIBA) and evaluation metrics; and then groups selected methods into random walk-based, GNN-based, metapath-based, and hybrid categories, with tables summarizing each method's graph mining technique, dataset, contribution, and source-code link. The survey closes with data-related and method-related future challenges. The abstract claims comprehensive coverage of graph-ML DTI methods and a wider range of methods than prior reviews.","tokens_in":26709,"tokens_out":5056,"duration_ms":49508,"significance":"If the survey's coverage were representative, the manuscript would be a useful entry point for practitioners seeking a 2020-2024 overview of graph-based DTI prediction, particularly because the tables consistently list source-code links and datasets. The taxonomy (random-walk, GNN, metapath, hybrid) is clear, and the emphasis on negative-sample selection, over-smoothing, and evaluation metrics is valuable. However, the central 'comprehensive/wider range' claim currently rests on an undeclared literature selection: the paper gives no search protocol and explicitly excludes several recent network-based methods. Because the survey's usefulness as a reference depends on this claim, the significance is presently conditional on the authors documenting or softening their scope.","major_comments":[{"comment":"The central claim that the survey 'provides comprehensive details' and 'includes a wider range of methods and approaches' than prior reviews [43-46] is not backed by any documented selection protocol. The paper states no search databases, query terms, inclusion/exclusion criteria, or screening process, and it explicitly excludes the network-based methods in [31,47-51], several of which (e.g., the knowledge-graph method [49] and arbitrary-order proximity deep forest [48]) fall within the survey's own network-based scope. As written, a reader cannot distinguish deliberate scope from omission, so the comprehensiveness claim is unverifiable; please add a literature search and screening description, or revise the abstract and the 'wider range' claim to a clearly scoped selection.","section":"Abstract and Section 'Network-based methods in DTI prediction'"},{"comment":"The statement that 'MHGNN and AMGDTI have AUC and AUPR of over 97% on Lou's dataset' is given without a citation, table, or specification of the evaluation split. Because the survey itself reports no performance numbers elsewhere, this claim cannot be checked; please remove it or replace it with a reference to the original papers and their reported metrics for the specific dataset version and split.","section":"Discussion and future challenges, method-related challenges"},{"comment":"The manuscript's title and most of its framing concern heterogeneous networks, but Table 7 includes methods applied only to drug molecular graphs and protein sequences without any heterogeneous network (e.g., GraphDTA, DGraphDTA, GEFA). The introduction should explicitly explain how these single-graph methods fit within the 'heterogeneous networks' scope, or the title and framing should be broadened to 'graph-based methods'.","section":"Section 'GNN-based methods', Table 7"}],"minor_comments":[{"comment":"There are unresolved placeholder references ('Error! Reference source not found.') for Figure 3, Figure 4, Table 2, Equation (6), and other locations; these must be fixed before publication.","section":"Throughout"},{"comment":"The dataset named after Luo is consistently spelled 'Lou's dataset' in the Discussion section; please standardize to 'Luo's dataset' to match the rest of the paper and the cited reference [32].","section":"Discussion and future challenges"},{"comment":"The definition of r_m^2 is incomplete: r_0^2 is not defined, and the expression under the square root requires a stated condition (or absolute value) to remain real; please add the definitions and conditions.","section":"Equation (10)"},{"comment":"The caption ends mid-sentence ('...the light blue node.'); it should be completed to describe what the alpha values represent after the walk transitions from the green node to the purple node.","section":"Figure 3 caption"},{"comment":"The text introducing Luo's dataset says the details 'are presented in  and Table 4,' omitting the table number for the node details; please insert the correct cross-reference.","section":"Table 3 and Table 4"}],"recommendation":"major_revision","confidential_remarks":"I see no evidence of misconduct or citation problems. The main issue is that the abstract and the 'wider range' claim substantially overpromise relative to what the body documents; this is fixable in revision by adding a selection protocol or softening the scope claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a narrative survey of graph-ML methods for drug-target interaction prediction, organized into random-walk, GNN, metapath, and hybrid families. It is not a new-results paper; its value is as a moderately useful map of the 2020-2024 literature, with a few gaps that matter.\n\nWhat the paper does well: the taxonomy is clean and mostly accurate. The preliminaries on metapaths, metagraphs, and GNN variants are compact and correct. The tables list methods, datasets, and claimed contributions, and the paper is honest enough to flag the network-based methods it deliberately excludes [31,47-51]. The future-challenges section makes reasonable points: negative-sample selection, class imbalance, inductive/cross-dataset evaluation, uncertainty quantification, and 3D structure via AlphaFold.\n\nThe soft spots are real but not fatal. The 'comprehensive' claim is not backed by a search protocol. There is no stated database, query, inclusion/exclusion criteria, or screening process, so a reader cannot judge whether the coverage is representative. The 'wider range' claim is also fragile: several recent graph-based DTI methods (knowledge-graph-based framework, multi-modal representation, arbitrary-order proximity deep forest) are excluded because they don't fit the four categories, which is a defensible scoping decision but undermines the word 'wider' relative to prior reviews that cover such approaches. The discussion asserts MHGNN and AMGDTI exceed 97% AUROC/AUPR on Luo's dataset without citation or a results table; that is an unsupported number and should be removed or sourced. The manuscript also has unresolved 'Error! Reference source not found.' placeholders and the source-code links are literally the word 'link', which would need to be fixed.\n\nWho the paper is for: a newcomer wanting a first map of the field, or a DTI researcher checking they haven't missed a method family. It does not change practice. It deserves peer review because it is a citable survey and the taxonomy has organizational value, but it needs revision before publication: add a documented selection procedure, soften or justify 'comprehensive', fix the placeholders, and either support or drop the 97% claim.\n\nIf I were editor, I would send it to a knowledgeable referee rather than desk reject; the core can be fixed with moderate effort.","headline":"A serviceable but under-documented narrative survey of graph-based DTI prediction; the taxonomy is useful, the 'comprehensive' claim is not backed by a search protocol, and the 97% performance assertion needs a citation.","tokens_in":27183,"tokens_out":2879,"would_cite":false,"duration_ms":26145,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that graph-based machine learning on heterogeneous biological networks has become the core tool for drug-target interaction prediction, and maps the 2020–2024 literature into four method families with datasets, metrics…","keywords":["drug-target interaction prediction","heterogeneous networks","graph neural networks","metapath","random walk","binding affinity","drug repurposing","survey"],"falsifier":"A reproducible literature search for network-based DTI prediction from 2020 to 2024, with explicit queries and inclusion rules, would settle the coverage claim; if it surfaces a substantial method family missing from the four categories—for instance knowledge-graph embedding approaches built on heterogeneous links, several of which the survey itself cites as excluded—then the survey's 'wider range' claim would need revision.","tokens_in":26370,"feed_emoji":"🧬","tokens_out":4158,"duration_ms":40009,"temperature":0.7,"pith_summary":"This survey tries to establish that graph-based machine learning, especially methods built on heterogeneous networks that join drugs, proteins, diseases, and side effects, is now central to computational drug-target interaction (DTI) prediction. It maps the 2020–2024 literature into four method families—random-walk embeddings, graph neural networks (GNNs), metapath-based methods, and hybrids—and gives each method's framework, main contribution, benchmark dataset, and source code. The authors argue this coverage is wider than earlier network-based reviews, and they identify future priorities: richer multimodal benchmark data, cross-dataset generalization, interpretability, uncertainty estimation, and the use of predicted protein 3D structure. A sympathetic reader comes away with an organized orientation to the field and its open problems.","feed_headline":"Survey maps four method families for drug-target prediction","feed_subtitle":"From random walks to GNNs, metapaths, and hybrids, with datasets, metrics, and code links for each family.","key_machinery":"The survey's organizing device is a taxonomy of network-based methods built on heterogeneous information networks, defined as graphs $G=(V,E)$ with node-type and edge-type mappings. The conceptual machinery includes metapaths (typed paths such as protein–disease–drug), metagraphs formed by metapath-based neighbors, random-walk embedding methods such as node2vec and DeepWalk, message-passing GNNs (GCN, GraphSAGE, GAT), and hybrid combinations. This taxonomy carries the argument: the survey's claim of wider coverage rests on showing that each reviewed method is an instance of one of these four families, and that the families are distinct in how they extract knowledge from heterogeneous biological graphs.","core_discovery":"The paper's central claim is that DTI prediction has converged on heterogeneous-network graph learning, and that the resulting methods divide into four recognizable families. It further claims that the regression formulation—predicting binding affinity rather than a binary interaction label—is the more meaningful task, because DTI datasets are incomplete and binary classifiers cannot distinguish true negatives from missing labels. The survey presents the common benchmark datasets (Yamanishi, Luo's, KIBA, Davis), the evaluation metrics appropriate for imbalanced data, and, for each reviewed method, its overall framework, contribution, dataset, and source-code link. Its contribution is organizational: a reader can use it to identify which graph-based approach fits a given prediction problem and what the field currently treats as unsolved.","pith_inferences":["If the reviewed selection is representative, the trend toward metapath and hybrid methods suggests that explicit biological semantics remain valuable even as GNNs automate representation learning, so the next generation of models may combine both rather than replace metapaths with pure message passing.","The paper itself notes that several network-based methods are excluded from its taxonomy, which implies the four-family map may be incomplete; a fuller map would need to place knowledge-graph embedding and other heterogeneous-network deep learning approaches relative to these families.","Because the survey reports that the large BETA benchmark is overlooked, a testable extension is to re-run the reviewed methods on BETA's seven validation tasks to see whether current performance rankings change under a broader evaluation protocol.","The source-code links collected in the tables would allow a direct reproducibility comparison across families, which the survey does not perform; such a benchmark would be a natural follow-up."],"forward_implications":["A reader can choose a method family by requirement: random walks for cheap topology-based embeddings, GNNs for structure-aware representations, metapaths for explicit biological semantics, and hybrids for combining multiple signal types.","Because affinity regression is presented as the more meaningful task, the KIBA and Davis benchmarks should be extended to include heterogeneous associations such as drug–disease, drug–drug, and protein–disease links.","Evaluation practice should shift from accuracy toward AUPR, F1-score, and MCC, and papers reporting AUPR should also report precision and recall separately.","Generalization claims should be tested inductively and across datasets, for instance by training on Luo's dataset and testing on an extended version, to prevent data leakage.","Future models should provide uncertainty estimates, since wet-lab validation is expensive and point predictions alone are hard to act on.","Adoption of predicted protein 3D structure, via tools such as AlphaFold, is expected to improve both accuracy and generalization once structure coverage ceases to be a barrier."],"supporting_citations":[{"why":"Supplies the prior survey of machine-learning databases and methods that frames the DTI prediction problem.","marker":"[7]"},{"why":"Introduces the Yamanishi benchmark dataset used by many reviewed classification methods.","marker":"[25]"},{"why":"Introduces Luo's heterogeneous dataset and the DTINet random-walk-with-restart approach that anchors the random-walk family.","marker":"[32]"},{"why":"Presents NeoDTI, a foundational GNN on heterogeneous networks that many later GNN-based methods build on.","marker":"[65]"},{"why":"Introduces GraphDTA, the reference method for applying GNNs to drug molecular graphs for affinity prediction.","marker":"[95]"},{"why":"Introduces SkipGNN, a representative hybrid method that combines skip-graph construction with representation learning.","marker":"[134]"},{"why":"Introduces the BETA benchmark, which the survey highlights as an overlooked resource for evaluating DTI predictors.","marker":"[144]"},{"why":"Provides an earlier GNN-focused survey that the paper contrasts with to support its claim of wider coverage.","marker":"[46]"}],"fun_headline_variants":["Four graph families for drug-target interaction prediction","Survey splits DTI prediction into four method classes","Graph learning survey maps drug-target prediction methods","Heterogeneous graph survey: four DTI prediction approaches","DTI survey: from random walks to GNNs, metapaths, hybrids"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's map of the field is only as reliable as the undeclared selection of papers it reviews, since it gives no search databases, query terms, inclusion criteria, or exclusion rules and explicitly sets aside several network-based methods.","fun_headline_variants_meta":{"raw":{"variants":["Four graph families for drug-target interaction prediction","Survey splits DTI prediction into four method classes","Graph learning survey maps drug-target prediction methods","Heterogeneous graph survey: four DTI prediction approaches","DTI survey: from random walks to GNNs, metapaths, hybrids"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1237,"prompt_tokens":800,"completion_tokens":437,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":416,"completion_tokens_details":{"reasoning_tokens":358}},"tokens_in":416,"tokens_out":437,"duration_ms":4819,"temperature":1.0,"reasoning_tokens":358,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:12:13.262861+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reproducible literature search for network-based DTI prediction from 2020 to 2024, with explicit queries and inclusion rules, would settle the coverage claim; if it surfaces a substantial method family missing from the four categories—for instance knowledge-graph embedding approaches built on heterogeneous links, several of which the survey itself cites as excluded—then the survey's 'wider range' claim would need revision.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents NeoDTI, a foundational GNN on heterogeneous networks that many later GNN-based methods build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the BETA benchmark, which the survey highlights as an overlooked resource for evaluating DTI predictors."}],"review_version":1}