{"id":"16bbb636-d239-43d8-9e81-e8f4bf43a66f","arxiv_id":"2507.21404","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LIT-PCBA has extensive molecular duplication and analog overlap that a parameter-free ECFP4 similarity baseline exploits to match state-of-the-art reported performance.","lead":"An audit of the widely used LIT-PCBA virtual screening benchmark finds thousands of identical and near-identical molecules across training, validation, and query sets. A trivial no-learning baseline exploits these overlaps and matches the reported enrichment scores of state-of-the-art deep learning models, calling published results into question.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central parity claim conflates different reference sets: the 'Mem' baseline scores with training actives, while CHEESE scores with query ligands; the paper's own query-only baseline does not match CHEESE medians.","rationale":"The reader identified transcription and protocol equivalence as the weakest assumption. This stress-test sharpens the protocol part: even granting the transcribed CHEESE numbers, Mem and CHEESE do not share a reference set. The manuscript itself provides the decisive control (the Qry row of Table 4), which fails to match CHEESE, so the claimed parity is reached only by injecting training actives into the baseline's reference set. That is a meaningful and informative leakage result, but it is not an apples-to-apples SOTA comparison. I also flag the metric-selection issue, because the paper criticizes selective use of one statistic while its headline relies on median raw EF1%; Table 2 shows Mem trails on mean raw and nEF1% statistics. These concerns do not overturn the paper's core leakage findings, which are concrete and independent of the CHEESE comparison, so the appropriate verdict remains conditional on revising the framing and headline claims. A single concrete check, recomputing the Qry baseline and the four summary statistics from Table 2, would settle the scope of the central claim.","tokens_in":19408,"tokens_out":14293,"duration_ms":161946,"concrete_test":"Run the released code to recompute the four baselines in Table 4 and the CHEESE comparison, then test the query-only baseline ('Qry') against the per-target CHEESE values from Fig. 7. If Qry's median raw EF1% is confirmed to be about 1.3 while the CHEESE medians are about 4.15, the parity claim depends on using training actives as the reference set and must be reframed as a leakage demonstration rather than a same-protocol SOTA match.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 claims that a zero-parameter ECFP4 baseline matches the median EF1% of the CHEESE 3D encoders 'under the same multi-ligand protocol.' But Mem is not run under the same protocol as CHEESE: CHEESE ranks the validation set by similarity to the PDB query ligands, whereas Mem scores each validation molecule by its maximum ECFP4 similarity to any training active. The paper's Table 4 supplies the directly comparable query-only baseline ('Qry'), whose median raw EF1% is 1.301, far below the 4.15 median attributed to EspSim and ShapeSim. Mem reaches 4.15 only by adding training actives as the reference set. That is exactly the information channel through which the documented train/validation leakage flows, so the exercise remains a valid leakage demonstration, but the headline 'matches SOTA under the same protocol' is not supported. The parity is also fragile because it relies solely on median raw EF1%: the paper's own Section 3.1.1 recommends reporting mean, median, and nEF1%, and on mean raw EF1% Mem (4.35) trails CE-E (4.66) and CE-S (5.21), as it does on mean nEF1% (0.063 vs 0.065 and 0.077) and on median nEF1% (0.042 vs 0.051 for CE-E). The abstract's stronger phrasing, 'match or exceed... state-of-the-art deep learning and 3D-similarity models,' goes beyond the evidence in Table 2, which contains no deep-learning baselines at all.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript audits the LIT-PCBA benchmark for data leakage and redundancy. Using deterministic RDKit-based comparisons, the authors report 2D-identical molecules shared across training/validation/query splits, thousands of shared inactives, and pervasive analog overlap (e.g., 323 ALDH1 active training–validation pairs at Tanimoto ≥ 0.6). They additionally implement a parameter-free memorization baseline that scores validation molecules by maximum ECFP4 similarity to training actives and report that its median raw EF1% (4.15) equals the median reported for CHEESE's 3D encoders. The paper concludes that LIT-PCBA rewards scaffold memorization rather than generalization and that nearly all published results on it are undermined. Code and data are released.","tokens_in":19744,"tokens_out":5773,"duration_ms":59669,"significance":"If the audit's factual findings hold, it is a valuable cautionary contribution: it adds systematic documentation of split leakage in a widely used benchmark and provides a simple, reproducible baseline demonstrating that training-set similarity alone extracts high enrichment. The deterministic nature of the counts and the release of code are strengths. The significance is reduced, however, by the mismatch between the 'same protocol' parity claim and the actual baseline design, and by the abstract's broader claim about deep-learning and 3D-similarity state-of-the-art results, which is not supported by the comparisons actually reported.","major_comments":[{"comment":"The claim that Mem matches CHEESE's 3D encoders 'under the same multi-ligand protocol' is not supported, because the reference-set protocols differ. Mem scores each validation molecule by maximum ECFP4 similarity to any training active, whereas the CHEESE encoders rank molecules by similarity to the PDB query ligands. The paper's own query-only baseline (Qry in Table 4) has median raw EF1% 1.301, far below the 4.15 median attributed to EspSim and ShapeSim; Mem attains 4.15 only by adding training actives as references. This is exactly the leakage channel documented in Section 5, so the exercise remains a valid leakage demonstration, but it is not an apples-to-apples 'same protocol' comparison. Please revise the wording and present a query-only baseline as the correct protocol-matched comparison.","section":"Section 3.1 and Tables 2–4"},{"comment":"The parity conclusion rests on a single statistic, the median raw EF1%, even though the paper itself argues that mean, median, and nEF1% should all be reported. Under the paper's own recommended statistics, Mem does not match the CHEESE encoders: mean raw EF1% is 4.35 for Mem versus 4.66 and 5.21 for CE-E and CE-S; mean nEF1% is 0.063 versus 0.065 and 0.077; median nEF1% is 0.042 versus 0.051 for CE-E. The claim should be restricted to the specific statistic for which it holds and the full metric set should be reported in the main text.","section":"Section 3.1.1 and Table 2"},{"comment":"The statement that the baseline 'match[es] or exceed[s] ... state-of-the-art deep learning and 3D-similarity models' is not supported by the evidence in the manuscript. Tables 2 and 3 contain no deep-learning baselines; the only head-to-head comparisons are with the two CHEESE 3D encoders. Similarly, the conclusion that 'nearly all published results on LIT-PCBA are undermined' goes beyond the demonstrated scope, which is a leakage demonstration for one family of methods and a set of deterministic overlap counts. Please either provide the missing comparisons or temper the abstract and conclusions to what the data show.","section":"Abstract and Section 1"},{"comment":"The CHEESE per-target values are transcribed from figures of preprint [17] rather than from a machine-readable source. Because the parity claim depends on the exact numbers (e.g., median 4.15), please provide the underlying per-target EF1% values, or the extraction script and figure source, so the comparison is independently checkable from the manuscript or repository.","section":"Section 3.1 and Tables 2–3"}],"minor_comments":[{"comment":"The text says 'over 350' active training–validation analog pairs, but the listed per-target counts (ALDH1 323, GBA 12, MAPK1 8, plus at least one each in FEN1, PKM2, and IDH1) sum to at least 346; please either report the exact total or align the wording.","section":"Section 5.1.2 and Table 6"},{"comment":"The sentence 'the first step is to extract one or more ligands are from co-crystallized PDB structures' contains a grammatical error; the word 'are' should be removed.","section":"Section 2, paragraph 2"},{"comment":"The term nEF1% is used before it is explicitly defined; please define it at first use in Section 3 (normalization by the target's theoretical maximum enrichment) so that Section 3.1.1 does not introduce it retroactively.","section":"Section 3.1.1"},{"comment":"This section uses EF0.1% while the rest of the paper discusses EF1%; please clarify the relationship between the two metrics and include EF0.1% in the metrics definitions.","section":"Section 6.2.1"},{"comment":"Reference [17] is given only as a 2024 preprint without a journal, arXiv identifier, or version; please add full bibliographic details so readers can locate the CHEESE paper.","section":"Reference [17]"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the audit's deterministic leakage counts and released code are likely to be useful to the community, but the headline claims are currently overdrawn relative to the comparisons actually run. If the authors revise the parity claim to the median raw EF1% and frame the result as a leakage-channel demonstration, while adding a query-only protocol-matched baseline, the paper could be a solid contribution within the scope of a benchmark-critique venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first systematic LIT-PCBA audit I have seen, and the hard counts are convincing: 2,491 stereo-agnostic inactives shared between training and validation, 2D-identical query ligands in PKM2 and elsewhere, 323 ALDH1 active analog pairs at Tc ≥ 0.6, and near-duplicate query sets in MTORC1, KAT2A, and GBA. The code and data are public, the methods are deterministic, and the numbers should be reproducible from the repo. That central finding is a real contribution and deserves attention. Second, the headline that a zero-parameter baseline \"matches state-of-the-art 3D encoders under the same multi-ligand protocol\" does not hold as stated. The stress-test note is right: Mem scores each validation molecule by max ECFP4 similarity to training actives, while CHEESE EspSim and ShapeSim score by similarity to the PDB query ligands. The paper's own Table 4 gives the directly comparable query-only baseline Qry, whose median raw EF1% is 1.301, far below the 4.15 attributed to CHEESE. Mem reaches 4.15 only by adding training actives as the reference set, which is exactly the leakage channel the audit documents. So the exercise remains a valid leakage demonstration, but it is not an apples-to-apples SOTA comparison. The abstract overreaches further: \"match or exceed... deep learning\" is unsupported because Table 2 contains no deep-learning baselines, and on mean raw EF1% and on mean/median nEF1% Mem actually trails CE-E and CE-S. The paper is partially self-aware: Section 3.1.1 argues that everyone, including the authors, should report mean, median, and nEF1%, and by that standard the headline \"matches SOTA\" is exactly the kind of selective reporting the paper criticizes. The broader claim that \"nearly all published results\" on LIT-PCBA are undermined also goes beyond the evidence; the paper does not re-evaluate SPRINT, DrugCLIP, or the others, it only documents leakage that plausibly affects them. I also want to credit Section 4 and Section 7: the stereochemistry-removal rationale is careful and not prescriptive, and declining to release a \"corrected\" benchmark because the problems are systemic is an honest and defensible choice. Who is this for? Anyone using LIT-PCBA or building virtual screening benchmarks. The leak counts and the query/training/validation overlap findings are the kind of thing you need to know before trusting any enrichment number from this benchmark. It deserves a serious referee. I would send it out, with a firm request to replace the parity claim with the honest version: a trivial memorization baseline exploits documented leakage, even though it does not match the 3D encoders under the query-only protocol, and to temper the scope of the conclusions accordingly.","headline":"The leakage audit is real, reproducible, and long overdue; the headline parity claim is not supported as stated.","tokens_in":20245,"tokens_out":2515,"would_cite":true,"duration_ms":30395,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A memorization baseline with no learnable parameters matches the median EF1% of state-of-the-art 3D encoders on LIT-PCBA, so the benchmark's scores measure scaffold memorization rather than generalization to novel chemotypes.","keywords":["data leakage","virtual screening benchmark","LIT-PCBA","memorization baseline","molecular redundancy","scaffold memorization","enrichment factor","benchmark audit"],"falsifier":"Run the released memorization baseline and the two CHEESE 3D encoders on the official LIT-PCBA splits under the multi-ligand max-pooling protocol and compare median raw EF1% values computed from code; if the CHEESE medians do not reproduce the transcribed values, the parity claim fails.","tokens_in":19217,"feed_emoji":"🧪","tokens_out":7735,"duration_ms":81191,"temperature":0.7,"pith_summary":"This paper tries to establish that LIT-PCBA, a widely used virtual-screening benchmark, is so riddled with duplicated and near-duplicated molecules across its training, validation, and query splits that published state-of-the-art scores do not measure generalization to new chemotypes. Its decisive experiment is a trivial memorization baseline: for each validation molecule, take the maximum ECFP4 Tanimoto similarity to any training active and rank by that score. With no learnable parameters, this baseline reaches a median raw enrichment factor of 4.15 at the 1% threshold, matching the median EF1% reported for the CHEESE 3D encoders under the same multi-ligand protocol. The audit also counts 2,491 2D-identical inactives shared between training and validation sets and more than 350 active analog pairs at Tanimoto similarity at least 0.6, arguing that the leakage is structural, not incidental. If the claim holds, published LIT-PCBA results should be reinterpreted as measuring scaffold memorization rather than screening skill.","feed_headline":"Zero-parameter baseline ties 3D encoders on LIT-PCBA","feed_subtitle":"An audit finds 2,491 duplicate inactives and 350+ analog pairs across splits; enrichment scores reward memorization.","key_machinery":"The load-bearing object is the memorization baseline: a 4096-bit ECFP4 fingerprint is computed for each molecule, stereochemistry is stripped so that stereo-agnostic duplicates count as identical, and each validation molecule is scored by its maximum Tanimoto similarity to any fingerprint in the set of training actives. The baseline has no learnable parameters and uses no physical modeling, so any enrichment it achieves must come from information the benchmark itself puts into both splits. The named comparison object is EF1%, the enrichment factor at the top 1% of the screened library, along with its normalized variant nEF1%, which divides raw EF1% by that target's theoretical maximum. The baseline's parity with the 3D encoders is the step that turns the leakage inventory into a claim about published results.","core_discovery":"The central discovery is that LIT-PCBA's validation set is not held out in any meaningful sense. Across 15 targets the paper finds stereo-agnostic duplicate molecules inside the same split and across splits, extensive analog overlap at ECFP4 Tanimoto similarity at least 0.6, and query sets dominated by near-identical ligands; in MTORC1, for example, nine of the eleven query entries reduce to three highly similar scaffolds. Because of this, a stereo-agnostic memorization baseline that simply max-pools ECFP4 similarity to training actives scores a median raw EF1% of 4.15, equal to the median reported for the CHEESE 3D encoders in the multi-ligand comparison, while selective reporting of the median rather than the mean, and failure to report normalized nEF1%, can invert which method appears superior. The paper concludes that nearly all published LIT-PCBA results, including zero-shot evaluations, are inflated by these artifacts, and that the benchmark cannot be repaired by de-duplication because analog leakage and low query diversity are systemic.","pith_inferences":["The same audit recipe could be applied prospectively to any new virtual-screening benchmark: compute a zero-parameter ECFP4 max-similarity baseline as part of the release, and flag any target where its EF1% exceeds a small fraction of the theoretical maximum as failing to test generalization.","The metric-sensitivity result suggests that published rankings across virtual-screening methods may be less robust than assumed wherever only one summary statistic is reported; reporting full per-target distributions and normalized scores would settle this for other benchmarks.","A testable extension would be to build a leakage-corrected split that removes duplicate inactives and analog pairs at Tanimoto 0.6 and then check whether the memorization baseline's EF1% drops to the level of a random ranker; the paper's claims imply that it would, but the paper does not run this experiment."],"forward_implications":["Published LIT-PCBA enrichment factors and AUROC scores should be read as measuring scaffold memorization, not recovery of novel chemotypes.","Zero-shot evaluations are also affected because query sets leak analogs of validation molecules, so their claims of generalization are not supported.","Choosing mean versus median raw EF1% can reverse the apparent ranking of methods; reporting normalized scores, both statistics, and full evaluation code is necessary for fair comparison.","Simple duplicate removal cannot salvage the benchmark, so future benchmarks need explicit controls on molecular overlap, scaffold similarity, and query diversity.","A no-learnable-parameter memorization baseline can serve as a lower-bound control that any serious virtual-screening benchmark should be checked against."],"supporting_citations":[{"why":"The LIT-PCBA benchmark itself; supplies the splits, labels, and query ligands that the audit examines.","marker":"[4]"},{"why":"The preprint whose reported EF1% values for the 3D encoders the memorization baseline is matched against.","marker":"[17]"},{"why":"The paper that introduced the asymmetric validation embedding used to create the LIT-PCBA training and validation splits.","marker":"[7]"},{"why":"The argument that most ligand-based classification benchmarks reward memorization rather than generalization; this is the framing the memorization baseline tests.","marker":"[25]"},{"why":"A prior audit showing hidden bias in an earlier docking benchmark, used as precedent for the claim that benchmark artifacts can inflate deep-learning performance.","marker":"[5]"},{"why":"The cheminformatics toolkit used to canonicalize SMILES and compute the ECFP4 fingerprints for the baseline and similarity analysis.","marker":"[30]"},{"why":"The screening study whose Tanimoto cutoff motivates the paper's conservative analog-leakage threshold.","marker":"[31]"}],"fun_headline_variants":["LIT-PCBA leaks: zero-parameter baseline ties deep models","Audit: 2,491 duplicate inactives invalidate LIT-PCBA","Memorization baseline matches 3D encoders on LIT-PCBA","LIT-PCBA: data leakage lets trivial baseline match SOTA","Zero-param baseline exploits LIT-PCBA leakage to tie 3D nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The head-to-head parity claim depends on the assumption that the published CHEESE EF1% values, which the authors transcribed from figures, were produced under the same multi-ligand max-pooling protocol used for the memorization baseline, and those values are not supplied in a machine-readable form in this paper.","fun_headline_variants_meta":{"raw":{"variants":["LIT-PCBA leaks: zero-parameter baseline ties deep models","Audit: 2,491 duplicate inactives invalidate LIT-PCBA","Memorization baseline matches 3D encoders on LIT-PCBA","LIT-PCBA: data leakage lets trivial baseline match SOTA","Zero-param baseline exploits LIT-PCBA leakage to tie 3D nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000613,"raw_usage":{"total_tokens":2896,"prompt_tokens":1037,"completion_tokens":1859,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1759}},"tokens_in":653,"tokens_out":1859,"duration_ms":13317,"temperature":1.0,"reasoning_tokens":1759,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:47:16.177852+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the released memorization baseline and the two CHEESE 3D encoders on the official LIT-PCBA splits under the multi-ligand max-pooling protocol and compare median raw EF1% values computed from code; if the CHEESE medians do not reproduce the transcribed values, the parity claim fails.","supporting_citations":[{"cited_title":"Lit-pcba: an unbiased data set for machine learning and virtual screening.Journal of chemical information and modeling, 60(9):4263–4273, 2020","cited_arxiv_id":null,"evidence_quote":"The LIT-PCBA benchmark itself; supplies the splits, labels, and query ligands that the audit examines."},{"cited_title":"Cheese: 3d shape and electrostatic virtual screening in a vector space","cited_arxiv_id":null,"evidence_quote":"The preprint whose reported EF1% values for the 3D encoders the memorization baseline is matched against."},{"cited_title":"Maximum unbiased validation (muv) data sets for virtual screening based on pubchem bioactivity data.Journal of chemical information and modeling, 49(2):169–184, 2009","cited_arxiv_id":null,"evidence_quote":"The paper that introduced the asymmetric validation embedding used to create the LIT-PCBA training and validation splits."},{"cited_title":"Most ligand-based classification benchmarks reward memorization rather than generalization","cited_arxiv_id":null,"evidence_quote":"The argument that most ligand-based classification benchmarks reward memorization rather than generalization; this is the framing the memorization baseline tests."},{"cited_title":"Hidden bias in the dud-e dataset leads to misleading performance of deep learning in structure-based virtual screening.PloS one, 14(8):e0220113, 2019","cited_arxiv_id":null,"evidence_quote":"A prior audit showing hidden bias in an earlier docking benchmark, used as precedent for the claim that benchmark artifacts can inflate deep-learning performance."},{"cited_title":"Rdkit: Open-source cheminformatics.https://www.rdkit.org","cited_arxiv_id":null,"evidence_quote":"The cheminformatics toolkit used to canonicalize SMILES and compute the ECFP4 fingerprints for the baseline and similarity analysis."},{"cited_title":"Alphafold2 structures guide prospective ligand discovery.Science, 384(6702):eadn6354, 2024","cited_arxiv_id":null,"evidence_quote":"The screening study whose Tanimoto cutoff motivates the paper's conservative analog-leakage threshold."}],"review_version":1}