{"id":"1139a7bf-7c82-4abf-881a-7961dafbbab6","arxiv_id":"2411.15418","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SPRINT co-embeds drugs and proteins with a structure-aware language model and attention pooling, achieving leading virtual screening enrichment and billion-scale retrieval speed.","lead":"SPRINT is a machine learning method that embeds drugs and proteins into a shared space, then searches that space to predict which molecules will bind to which proteins. It reports top virtual screening accuracy on several benchmarks and can scan a 6.7 billion molecule database in minutes, which could make whole-proteome drug screens practical.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LIT-PCBA zero-shot evaluation may be contaminated by ligand overlap with the PubChem-derived MERGED training set; reported enrichment factors could reflect memorized ligand labels rather than target-specific generalization.","rationale":"The reader's weakest assumption is exactly the ligand-side leakage in the LIT-PCBA zero-shot evaluation, and the paper's own text supports the plausibility of that leakage: MERGED is dominated by PubChem data, and LIT-PCBA labels originate from PubChem bioassays. The reported protocol filters only by protein homology, not by ligand identity or similarity, so the model could in principle memorize active-ligand fingerprints rather than learn target-specific interactions. This is load-bearing because the headline virtual-screening SOTA rests on Table 2. The proposed retraining test would settle whether the concern actually lands. The CACHE2 and TDC affinity results provide some independent support but do not by themselves establish the LIT-PCBA SOTA claim. Since the reader already issued a CONDITIONAL verdict based substantially on this same concern, my read does not change the verdict; if the leakage test fails, the verdict should move to REJECT.","tokens_in":14814,"tokens_out":5105,"duration_ms":46642,"concrete_test":"Compute, for every LIT-PCBA active, the maximum Morgan fingerprint (radius 2, 2048-bit) Tanimoto similarity to any molecule in the MERGED training split (and the exact InChIKey overlap). Then retrain SPRINT on MERGED after removing all training molecules with Tanimoto ≥0.6 to any LIT-PCBA active, while keeping the existing protein-homology filter, and re-evaluate Table 2 with the same 3:1 negative sampling. If EF@0.5% or BEDROC falls below DrugCLIP/Gnina, the headline LIT-PCBA SOTA is not leakage-free; if the metrics are stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central SOTA claim for virtual screening rests on the LIT-PCBA results in Table 2. The protocol described in Section 2.1 removes only protein sequences with ≥90% homology to LIT-PCBA from MERGED before pre-training; there is no ligand-side filter. This matters because Section 3.1 states LIT-PCBA activity labels are derived from dose-response bioassays in PubChem, and the MERGED training set is 98.31% PubChem data. The same active molecules, or close analogs, can therefore appear in training with positive labels, possibly against different targets. Since SPRINT's drug encoder is a Morgan-fingerprint MLP (Section 3), the model can learn a ligand-only shortcut: fingerprint patterns of known actives are mapped to high similarity with any protein embedding, inflating enrichment factors and BEDROC independent of target-specific binding. Table C2 intensifies the concern by showing the 3:1 negative-sampling ratio was chosen from LIT-PCBA ablation, so the benchmark is not a fully held-out zero-shot evaluation. If ligand overlap is substantial, the reported EF 15.90 at 0.5% and 73.4 AUROC are not evidence of a new SOTA for target-specific structure-aware screening; they may reflect memorization of PubChem activity labels.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces SPRINT, a deep co-embedding model for drug–target interaction prediction in which Morgan fingerprints and structure-aware protein language model (SaProt) embeddings are projected into a shared space, with interaction probability given by a sigmoid of a scaled cosine similarity and protein residues aggregated by multi-head attention pooling. The authors report state-of-the-art results on the LIT-PCBA virtual screening benchmark, on several DTI classification benchmarks (particularly with the smaller SPRINT-sm variant), and match the top leaderboard method on the TDC BindingDB Patent affinity prediction task. They also demonstrate a large-scale screen of the Enamine REAL database against the human proteome in minutes, analyze attention maps for interpretability, and release code and data.","tokens_in":15091,"tokens_out":6213,"duration_ms":51130,"significance":"If the LIT-PCBA results were leakage-free and the CACHE2 proxy were a valid measure of hit finding, SPRINT would be a significant advance: it combines an interpretable, structure-aware protein representation with a highly scalable vector-retrieval screening pipeline, and the authors provide a reproducible codebase. Strengths include the transparent architecture, the use of shared splits with reported variance in Table 1, the application of a structure-aware PLM in this co-embedding setting, and the concrete large-scale screening demonstration. However, the central virtual-screening claim depends on the LIT-PCBA evaluation being uncontaminated and on a model-selection protocol that does not use the test benchmark; these conditions are not currently established.","major_comments":[{"comment":"The zero-shot LIT-PCBA evaluation decontaminates only the protein side. Section 2.1 states that all protein sequences with at least 90% homology to LIT-PCBA were removed from MERGED using MMSeqs2, but no ligand-side filter is described. Since LIT-PCBA activity labels derive from PubChem bioassays and MERGED is 98.31% PubChem (Section 3.1), the same active molecules or close analogs can occur in the training set. Because SPRINT's drug encoder is a Morgan-fingerprint MLP, the model can memorize ligand-label associations and inflate EF and BEDROC regardless of target-specific binding. The authors should report the overlap between LIT-PCBA ligands and MERGED training molecules (at identity and at typical analog thresholds) and re-run the evaluation with these molecules removed, or otherwise demonstrate that the reported metrics are unchanged.","section":"Section 2.1 and Section 3.1"},{"comment":"The choice of the 3:1 negative sampling ratio is made using LIT-PCBA performance. Table C2 presents LIT-PCBA results for 1:1 and 3:1 sampling, and the 3:1 configuration is then reported as the final model in Table 2. This is a form of model selection on the test benchmark, which invalidates the 'zero-shot' characterization and may inflate the reported numbers if the benchmark also drives hyperparameter selection. The authors should either hold out LIT-PCBA entirely for final evaluation, or show that the chosen ratio was selected on a separate validation set without reference to LIT-PCBA.","section":"Section 2.1 and Table C2"},{"comment":"The CACHE2 comparison uses Gnina CNN VS docking scores as a proxy for hit-finding performance. The claim that SPRINT 'finds almost three times the number of high-scoring molecular scaffolds' is based on a threshold of CNN VS > 6, not on experimentally confirmed hits. The distributional comparison in Figure 1 shows that SPRINT-selected molecules have higher predicted docking scores than DeepDocking's Batch 0, but this does not establish that SPRINT would produce more hits in the prospective CACHE2 setting. The authors should re-frame this section as a docking-score distribution analysis and remove or temper the hit-finding language.","section":"Section 2.2 and Figure 1"},{"comment":"The text states that attention pooling achieves 'SOTA predictive scores for DTIs on most benchmarks,' but the full SPRINT (16M) model is not the best model on BIOSNAP, Unseen Drugs, Unseen Targets, DAVIS, or BindingDB; the smaller SPRINT-sm (10M) model is. The SOTA claim is only consistently supported for SPRINT-sm on those benchmarks and for SPRINT on MERGED. The text should be revised to attribute each result to the appropriate model variant.","section":"Table 1"},{"comment":"Table 2 reports no standard deviations or replicate runs for the LIT-PCBA AUROC, BEDROC, and EF values. Given the reported sensitivity to random seeds in Section 2.3 and Appendix E, the absence of uncertainty estimates makes it difficult to assess whether the improvements over DrugCLIP and other baselines are statistically meaningful. Please provide confidence intervals or standard errors for the main virtual screening results.","section":"Table 2"}],"minor_comments":[{"comment":"The reference for the CACHE2 method [33] is incomplete and should be updated to a proper citation.","section":"References"},{"comment":"Please check the 'T able' spacing artifacts in table captions and ensure the final PDF renders table titles correctly.","section":"Table captions"},{"comment":"The sentence 'Crystal structures with bound fragments in the RNA-binding site exist, with PDB ID 5RLZ used for virtual screening' could be clarified to state which structures were used for the CACHE2 screen.","section":"Section 2.2"},{"comment":"The statement that 'all but one' ProtBert heads attend less to binding residues should be checked against Figure 2(a); if all four heads show this pattern, the text should be corrected.","section":"Section 2.3"},{"comment":"The term 'zero-shot' is used in two different senses (the unseen drugs/targets splits in Table 1 and the LIT-PCBA evaluation); consider defining the term explicitly to avoid confusion.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central issue for the editor: the LIT-PCBA evaluation, which is the paper's headline result, is at risk of ligand-level data leakage because the training set is dominated by PubChem while LIT-PCBA labels also originate from PubChem. This is not a matter of outside-consensus disagreement but an unaddressed potential confound that can be fixed with additional computational experiments. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: SPRINT is a solid engineering advance in vector-based virtual screening, and the billion-scale retrieval numbers are real. But the headline LIT-PCBA result rests on a zero-shot claim that the paper doesn't actually secure: training data is ~98% PubChem, LIT-PCBA labels come from PubChem, and only protein homology was filtered. That needs to be addressed before I'd buy the SOTA claim.\n\nWhat's genuinely new and good: the combination of SaProt structure tokens with attention pooling over the ConPLex-style co-embedding is a reasonable, well-tested extension. Table 1 shows consistent gains over ConPLex on shared splits with variance reported. The TDC BindingDB Patent result is credible—matching an ensemble without knowledge graphs is a nice result. The retrieval stack (ChromaDB, 7 ms per target, 16-minute human-proteome screen of Enamine REAL) is an impressive engineering contribution, and they've put code and data up.\n\nSoft spots, in order:\n1. Ligand leakage in LIT-PCBA. Section 2.1 states only ≥90% protein homology was removed from MERGED. No ligand-side filter is described. Since MERGED is 98.31% PubChem and LIT-PCBA labels are PubChem bioassay data, the drug encoder (Morgan MLP) could memorize active scaffolds and produce inflated enrichment. The paper doesn't discuss this. This is the load-bearing concern.\n2. Benchmark tuning. Table C2 shows the 3:1 negative-sampling ratio was chosen after looking at LIT-PCBA performance. That makes the Table 2 numbers a selected result, not a held-out zero-shot evaluation.\n3. CACHE2 is a docking-proxy comparison, not an experimental hit-finding result. Screening with SPRINT and rescoring with Gnina shows higher CNN VS scores, which is suggestive, but it's not a prospective validation. The text overstates it by calling it 'enhancing real-world virtual screening.'\n4. Interpretability. The abstract touts residue-level attention maps, but Section 2.3 says 'there is little biological relevance for the attention patterns at its current scale.' That's a mismatch worth fixing.\n\nMinor: Table 2 has no error bars for the flagship numbers. Alpha scaling and homology thresholds are reasonable choices; those aren't concerns.\n\nBottom line: this deserves serious peer review. The scaling result alone is publishable. But the LIT-PCBA SOTA claim needs ligand-side deduplication or at least a similarity analysis, plus honest reframing of the CACHE2 and interpretability claims. I'd send it to a good venue with a request for revision, not desk reject.","headline":"SPRINT's billion-scale retrieval is real and worth publishing, but the headline LIT-PCBA zero-shot claim doesn't rule out ligand-side leakage from a PubChem-heavy training set.","tokens_in":15605,"tokens_out":2971,"would_cite":true,"duration_ms":26027,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SPRINT's central claim is that structure-aware protein-language co-embeddings with learned attention pooling make cosine similarity a reliable, ultra-fast predictor of drug-target binding, enabling billion-molecule virtual screens.","keywords":["virtual screening","drug-target interaction prediction","co-embedding","protein language models","structure-aware embeddings","multi-head attention pooling","vector search","LIT-PCBA"],"falsifier":"Compute exact and Tanimoto-similarity overlap between the MERGED training molecules and the LIT-PCBA ligands, retrain SPRINT after excluding any training pair whose drug resembles a benchmark ligand, and re-measure AUROC and enrichment factor at 0.5%; a substantial drop would indicate that the reported enrichment came from memorization rather than generalization.","tokens_in":14621,"feed_emoji":"⚡","tokens_out":8331,"duration_ms":74530,"temperature":0.7,"pith_summary":"The paper sets out to show that virtual screening no longer needs to be bottlenecked by molecular docking. It presents SPRINT, a model that learns a shared vector space for drugs and protein targets, where binding likelihood is read off as a sigmoid-scaled cosine similarity. The authors claim that by pooling per-residue embeddings from a structure-aware protein language model with learned multi-head attention, SPRINT beats the leading contrastive and docking baselines on the LIT-PCBA benchmark, matches or exceeds prior methods on DTI classification and binding-affinity prediction, and can query a 6.7-billion-molecule library against a whole proteome in minutes. If these results hold, the practical payoff is that screening for off-target effects, repurposing, and new mechanisms of action becomes a fast retrieval task rather than a weeks-long docking campaign.","feed_headline":"SPRINT screens 6.7B molecules against human proteome in 16 minutes","feed_subtitle":"A learned vector space turns drug-target search into milliseconds-per-target retrieval across whole proteomes.","key_machinery":"The load-bearing object is the drug–target co-embedding space defined by $P(Y=1|Z_d,Z_t)=\\sigma(\\alpha\\, \\text{cosine}(Z_d,Z_t))$ with $\\alpha=5$. A frozen Morgan fingerprint encoder and a frozen structure-aware protein language model (SaProt) feed modality-specific MLPs whose outputs $Z_d$ and $Z_t$ are the co-embeddings; the protein side uses multi-head attention pooling over per-residue embeddings instead of the average pooling used by ConPLex. Structure is injected through Foldseek tokens computed from AlphaFold2 structures, so the model sees sequence and geometry without explicit docking. The same space serves three tasks: binary interaction classification under cross-entropy, affinity regression when the sigmoid and cosine are replaced by a dot product, and ultra-fast retrieval when the embeddings are indexed in a vector store.","core_discovery":"On the paper's own terms, SPRINT's discovery is that structure-aware protein language embeddings, when aggregated by a learnable attention mechanism, form a co-embedding space in which cosine distance is a sufficient statistic for drug-target interaction. Trained with binary cross-entropy on the large MERGED dataset containing PubChem, BindingDB, and ChEMBL interactions, the 16-million-parameter model reaches 73.4% AUROC and a 15.90 enrichment factor at a 0.5% false-positive rate on LIT-PCBA in a zero-shot setting, outperforming DrugCLIP, Gnina, and docking baselines. The same co-embedding space doubles as a binding-affinity predictor, matching the top TDC BindingDB Patent leaderboard ensemble, and as a vector-search index that screens 6.7 billion Enamine REAL molecules against the human proteome in 16 minutes. The paper also reports that SPRINT's selected compounds score higher in a docking-based CACHE2 screen than the first random batch of DeepDocking, and that its molecule embeddings improve antibacterial and toxicity property prediction when concatenated with Morgan fingerprints.","pith_inferences":["Editorial inference: the 16-minute screen is reported as query time; a full deployment would also need to account for index construction, memory, and update costs as a 6.7-billion-embedding store grows, which the paper does not quantify.","Editorial inference: the zero-shot claim would be much stronger with a ligand-side leakage check; measuring Tanimoto similarity between MERGED training drugs and LIT-PCBA ligands would show whether enrichment comes from generalization or from memorization of related molecules.","Editorial inference: because the best-screening attention maps are the least interpretable, the architecture likely separates two useful properties; a future design could train an interpretable pooler for explanation while keeping a high-recall pooler for screening.","Editorial inference: the co-embedding recipe may transfer to non-protein targets such as RNA or modified peptides, but SaProt's amino-acid vocabulary and structure tokens would need to be replaced or extended."],"forward_implications":["If the reported LIT-PCBA numbers are right, ligand enrichment no longer requires docking: a frozen protein-language representation plus learned pooling outperforms structure-based screeners on a benchmark designed to avoid DUD-E-style bias.","Whole-proteome and pan-species screens become cheap enough to run routinely, making off-target prediction and drug repurposing practical at the scale of billions of molecules.","The same co-embedding space can be used directly for binding-affinity prediction, so classification, regression, and retrieval need not be separate models.","DTI pre-training produces molecule embeddings that add value on top of Morgan fingerprints for antibacterial and toxicity prediction, suggesting the co-embedding space captures target-neighborhood information.","SPRINT can replace the initial random sample in iterative docking pipelines, yielding higher-scoring candidates with roughly one-sixth of the docking effort."],"supporting_citations":[{"why":"Supplies the contrastive co-embedding framework and the dataset splits that SPRINT extends and compares against.","marker":"[2]"},{"why":"Provides the structure-aware protein language model and per-residue embeddings that SPRINT pools with attention.","marker":"[8]"},{"why":"Is the main contrastive virtual-screening baseline whose LIT-PCBA results SPRINT is designed to beat.","marker":"[9]"},{"why":"Is the vector store that enables the reported 7 ms per-target queries and 16-minute proteome-scale screens.","marker":"[13]"},{"why":"Provides the MERGED training corpus of PubChem, BindingDB, and ChEMBL interactions used for large-scale pretraining.","marker":"[16]"},{"why":"Defines the LIT-PCBA benchmark that supplies the paper's zero-shot virtual screening evaluation.","marker":"[26]"},{"why":"Is used to split the MERGED data by protein homology and to remove proteins with high homology to LIT-PCBA.","marker":"[28]"},{"why":"Provides the docking software and CNN scoring used in the CACHE2 comparison.","marker":"[23]"},{"why":"Supplies the DeepDocking screen that SPRINT is measured against and used to seed with higher-scoring molecules.","marker":"[33]"}],"fun_headline_variants":["SPRINT: 6.7B drug–protein screens in 16 min","SPRINT maps whole proteome to 6.7B drugs in 16 min","SPRINT: 16-min screen of 6.7B compounds vs proteome","SPRINT: interpretable billion-scale drug-target screening","SPRINT: whole-proteome screen of 6.7B molecules in 16 min"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the LIT-PCBA zero-shot evaluation is free of ligand leakage: only proteins with at least 90% sequence homology were removed from pretraining, and no check is described for whether the same or closely similar drug molecules appear in the PubChem, BindingDB, or ChEMBL training data.","fun_headline_variants_meta":{"raw":{"variants":["SPRINT: 6.7B drug–protein screens in 16 min","SPRINT maps whole proteome to 6.7B drugs in 16 min","SPRINT: 16-min screen of 6.7B compounds vs proteome","SPRINT: interpretable billion-scale drug-target screening","SPRINT: whole-proteome screen of 6.7B molecules in 16 min"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001393,"raw_usage":{"total_tokens":5685,"prompt_tokens":1041,"completion_tokens":4644,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":4534}},"tokens_in":657,"tokens_out":4644,"duration_ms":30803,"temperature":1.0,"reasoning_tokens":4534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:19:15.268873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute exact and Tanimoto-similarity overlap between the MERGED training molecules and the LIT-PCBA ligands, retrain SPRINT after excluding any training pair whose drug resembles a benchmark ligand, and re-measure AUROC and enrichment factor at 0.5%; a substantial drop would indicate that the reported enrichment came from memorization rather than generalization.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the structure-aware protein language model and per-residue embeddings that SPRINT pools with attention."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the main contrastive virtual-screening baseline whose LIT-PCBA results SPRINT is designed to beat."},{"cited_title":"Chroma - the open-source embedding database","cited_arxiv_id":null,"evidence_quote":"Is the vector store that enables the reported 7 ms per-target queries and 16-minute proteome-scale screens."},{"cited_title":"A large dataset curation and benchmark for drug target interaction","cited_arxiv_id":"2401.17174","evidence_quote":"Provides the MERGED training corpus of PubChem, BindingDB, and ChEMBL interactions used for large-scale pretraining."},{"cited_title":"& Rognan, D","cited_arxiv_id":null,"evidence_quote":"Defines the LIT-PCBA benchmark that supplies the paper's zero-shot virtual screening evaluation."},{"cited_title":"& S¨ oding, J","cited_arxiv_id":null,"evidence_quote":"Is used to split the MERGED data by protein homology and to remove proteins with high homology to LIT-PCBA."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the docking software and CNN scoring used in the CACHE2 comparison."},{"cited_title":"of pittsburgh] computational methods - cache2","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepDocking screen that SPRINT is measured against and used to seed with higher-scoring molecules."}],"review_version":1}