{"id":"d5ebc912-4e0e-4a9a-9acb-747909a29547","arxiv_id":"2411.14157","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A DrugGPT-derived model, fine-tuned on approved drugs and optimized with reinforcement learning, produces more valid molecules and higher purported binding affinities, though the affinity metric is the training reward.","lead":"DrugGen is an AI model that generates small drug-like molecules from protein sequences, built by fine-tuning the existing DrugGPT model on approved drugs and then optimizing it with reinforcement learning to favor valid, high-binding molecules. The paper reports higher validity and predicted binding scores than DrugGPT, but its main evaluation metric is the same predictor used as the training reward, so the affinity gains are partly circular.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline affinity advantage is not independently established: PLAPT serves as both the PPO training reward and the evaluation metric, so higher DrugGen scores are an expected optimization outcome; the docking evidence is sparse, selective, and never compares DrugGPT.","rationale":"The reader's weakest_assumption correctly identifies the circular PLAPT evaluation. This is the single most load-bearing concern because the headline numeric claim (7.22 vs 5.81) comes from the same oracle used as the PPO reward. The validity improvement (95.45% to 99.90%) is credible because it is a direct chemical validity check and does not depend on PLAPT. The docking results do not salvage the affinity claim: they are one-sided, selective, and lack statistical comparison to DrugGPT. The internal contradiction on diversity/novelty further weakens the abstract's claim, though it is secondary to the circularity. I therefore do not see a reason to change the reader's rejection, but the concrete test could rescue the affinity claim if it shows a consistent independent advantage.","tokens_in":15298,"tokens_out":3556,"duration_ms":50865,"concrete_test":"Generate 100 unique molecules per target from both DrugGPT and DrugGen using the same sampling procedure, prepare them with LigPrep, and dock all of them into the corresponding crystal structures (as in Section 4.3.4) using GLIDE XP. Compare the docking score distributions with a Mann–Whitney U test. If DrugGen's scores are not significantly better than DrugGPT's across most targets, the PLAPT-based affinity improvement is an artifact of reward overfitting. Equivalent test: rescore the same molecule sets with an independent ML predictor (e.g., GNINA) and check for a consistent advantage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 4.3.3, PLAPT's neg_log10_affinity_M is used as the reward signal for PPO, and in Section 4.3.4 the same PLAPT model is the evaluation metric for binding affinity. PPO explicitly maximizes this reward, so generating molecules that score high on PLAPT is the optimization target; reporting higher PLAPT scores for DrugGen than for DrugGPT does not demonstrate improved real binding affinity, only successful optimization against the chosen oracle. The paper's own docking results are limited: they are reported only for DrugGen (Table 2), with no parallel DrugGPT docking scores, no statistical test, and a small, possibly cherry-picked subset of molecules. The FABP5 example is anecdotal, and the ACE scores are confounded by ligand binding to different sites. Additionally, the abstract's claim that diversity and novelty are 'maintained' is contradicted by Figures 3A/B: diversity falls from 84.54% to 60.32% and novelty from 66.84% to 41.88%, both highly significant. Thus the central claim of a higher-affinity, equally diverse/novel tool rests on circular evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents DrugGen, a DrugGPT-based generative model for target-specific SMILES generation. DrugGen is produced by supervised fine-tuning DrugGPT on approved drug-target pairs and then applying PPO with a reward that combines PLAPT-predicted binding affinity, an RDKit-based validity checker, and penalties for repetition and approved-drug matches. The authors evaluate DrugGen against DrugGPT on eight protein targets using validity, diversity, novelty, PLAPT affinity, and molecular docking. They report that DrugGen improves validity from 95.45% to 99.90%, produces higher predicted binding affinities (7.22 vs. 5.81), and claim that diversity and novelty are maintained; docking scores for selected DrugGen molecules are also presented. The paper includes public code and data links.","tokens_in":15532,"tokens_out":6659,"duration_ms":69541,"significance":"If the central claims were supported, DrugGen would be a useful, well-documented contribution to RL-based molecular generation. The validity result is a concrete, reproducible improvement: the RDKit-based validity checker is described, and the validity statistics are straightforward. The authors also share code, checkpoints, and data, which supports reproducibility. However, the headline affinity claim is not independently established: PLAPT is used both as the PPO reward and as the evaluation metric, so the higher DrugGen scores are an expected optimization outcome rather than evidence of higher true binding affinity. The docking analysis is selective and is not compared with DrugGPT. In addition, the paper's own data show statistically significant decreases in diversity and novelty, contradicting the abstract's claim that these properties are maintained. The contribution, properly scoped, may be a demonstration that PPO with PLAPT can optimize a generative model's PLAPT scores while improving validity, but the current manuscript overclaims biological and medicinal significance.","major_comments":[{"comment":"PLAPT's neg_log10_affinity_M is the PPO reward and also the evaluation metric for binding affinity. Since PPO explicitly maximizes this reward, the observed difference (7.22 vs. 5.81) is a check that optimization worked, not evidence that DrugGen molecules bind better in reality. To support the affinity claim, the authors need an independent evaluation: a held-out or different affinity predictor, docking of both DrugGen and DrugGPT outputs under identical protocols, or experimental data. The Discussion's caveat that the reward model has inherent accuracy and specificity limitations does not address this circularity.","section":"Sections 4.3.3 and 4.3.4, Table 1"},{"comment":"The abstract states that DrugGen maintains diversity and novelty, but the reported values show large significant declines: diversity falls from 84.54% to 60.32% and novelty from 66.84% to 41.88%, both with highly significant test statistics. The text acknowledges these decreases, calling generated molecules more similar and fewer novel, yet concludes there is a good balance. No criterion or comparison is provided for what counts as maintained or balanced, so the claim is unsupported and the abstract is misleading.","section":"Abstract, Section 2.2, Figures 3A and 3B"},{"comment":"The docking evidence is not a controlled comparison. Only DrugGen molecules are docked; DrugGPT molecules are not docked with the same protocol, and only a small subset of the 122 docked molecules is reported in Table 2 without a stated selection criterion. The ACE results are acknowledged to bind different sites from the reference, so they cannot be read as target-site validation, and the FABP5 and NAMPT examples are anecdotal. Moreover, the method text calls the docking blind, while Table 3 lists grid boxes of 30-40 Å, which is not a whole-protein search; the exact search space needs to be clarified. Thus the docking section cannot substitute for an independent affinity assessment.","section":"Section 2.3, Table 2, Methods 4.3.4"},{"comment":"The aggregate PLAPT comparison pools eight targets, of which one (FABP5) shows no significant difference after Bonferroni correction, and the Discussion concedes variability across targets. Reporting a pooled median across targets hides per-target differences and overstates consistency. The authors should present per-target effect sizes and interpret the aggregate statistic with the non-significant target explicitly excluded or modeled.","section":"Section 2.3, Table 1"}],"minor_comments":[{"comment":"There are several textual typos, including 'DrugGen/quotesingle.ts1' in Sections 2.3 and 3 and inconsistent spacing in 'F ABP5'; these should be corrected.","section":"Throughout"},{"comment":"The validity number is reported inconsistently: the abstract and novelty section say 100% valid, while Section 2.2 reports 99.90% for DrugGen; the manuscript should state which generation set each number refers to.","section":"Abstract and Section 2.2"},{"comment":"The p-values are reported as 'P = 0' for diversity and as 'P = 1' for FABP5; exact p-values should be given as inequalities (for example, P < 10^-300) rather than 0, and the FABP5 value should be reported as non-significant after correction.","section":"Section 2.2 and Table 1"},{"comment":"PLAPT is cited only as a 2024 bioRxiv preprint; since it is central to both the method and the evaluation, the authors should cite the version used and specify whether it was used with default weights or fine-tuned.","section":"Section 4.3.3 and References"},{"comment":"The caption of Figure 1 says the assessment includes binding affinity for both DrugGen and DrugGPT, but docking simulations are only reported for DrugGen; the figure and text should be aligned.","section":"Figure 1"}],"recommendation":"reject","confidential_remarks":"The reader's report already captures the main technical issues. For the editor: the manuscript's contribution is closer to a technical demonstration than a validated drug-discovery method. The circularity of using PLAPT as both reward and metric is central; even with careful reframing, the docking analysis would need to be substantially expanded and compared against DrugGPT before the affinity claims are credible. The paper may be better suited to a machine-learning methods venue after recalibration, but in its present form it does not meet the standard for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about DrugGen. Short version: the validity result is real, the code/data release is a plus, but the main affinity claim doesn't survive contact with the methods.\n\nThe validity jump (95.45% to 99.90%, chi-squared p < 1e-38) is a clean RDKit-based result. The authors also ship the dataset, checkpoints, and code, which is more than many papers in this space do. The docking protocol is standard Glide XP with redocking RMSDs under 2.1 Å, showing care. The SFT+PPO pipeline is clearly described.\n\nThe problem is Section 4.3.3/4.3.4: PLAPT is the PPO reward and the binding-affinity metric. Molecules that score higher on PLAPT after PPO are exactly what the optimization was supposed to produce; that is not independent evidence of higher real affinity. Figures 3A/B contradict the abstract's 'maintaining diversity and novelty': diversity drops from 84.5% to 60.3% and novelty from 66.8% to 41.9%, both with tiny p-values. The docking data is only for DrugGen, not DrugGPT, so it cannot support a comparative claim; the ACE numbers are confounded by different binding sites, and the FABP5 example is anecdotal. The free parameters (reward penalty 0.7, KL coefficient) are not swept, which matters because the diversity drop might be tuned back.\n\nAlso, the novelty metric compares to approved drugs; since DrugGen was fine-tuned on approved drugs, lower novelty is an expected trade-off, not a bug—but then the paper shouldn't claim it's maintained.\n\nWho is this for? Someone working on RL fine-tuning of molecular generators would find a useful case study, but the claims need to be reworked. I would not cite the affinity gain. I would accept it for peer review because the validity improvement and released assets deserve a careful referee, and the circularity deserves a public airing.\n\nMy recommendation: reject in current form, but invite a revision with a held-out affinity predictor or docking-based comparison across both models, corrected diversity/novelty claims, and a sensitivity analysis of the reward coefficients.","headline":"Useful validity improvement and clean release, but the headline affinity gain is a reward-overfitting artifact and the abstract overstates diversity/novelty.","tokens_in":16091,"tokens_out":2645,"would_cite":false,"duration_ms":27511,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By fine-tuning a GPT-style molecule generator on approved drugs and then optimizing it with reinforcement learning against a predicted binding-affinity reward, the paper's DrugGen achieves 100% valid structure generation and higher…","keywords":["drug discovery","large language model","reinforcement learning","proximal policy optimization","binding affinity prediction","SMILES generation","molecular docking","drug repurposing"],"falsifier":"Take the top-scoring DrugGen molecules for ACE, NAMPT, and FABP5 and measure their actual binding (for example, IC50 or Kd) alongside an equal number of DrugGPT molecules; if DrugGen's molecules do not bind more tightly, the central affinity claim is refuted. A faster in silico check is to re-run the PPO phase with a completely different affinity predictor and see whether the apparent gain persists.","tokens_in":15113,"feed_emoji":"💊","tokens_out":6744,"duration_ms":60783,"temperature":0.7,"pith_summary":"The paper claims that a large language model trained to write molecular formulas can be made substantially more useful for early drug discovery by first fine-tuning it on known approved drug-target pairs and then optimizing it with reinforcement learning, using a predicted binding-affinity score as the reward. The resulting model, DrugGen, produces 100% valid molecular structures versus 95.5% for its base model DrugGPT, and its molecules receive higher predicted binding affinities (median 7.22 versus 5.81) across all eight test targets. The authors argue this combination makes the generator both more reliable and more potent, and docking simulations suggest some generated molecules bind their targets more strongly than the reference drugs. If the affinity predictor's scores reflect real binding, the recipe could accelerate lead generation and drug repurposing.","feed_headline":"DrugGen generates 100% valid molecules with stronger predicted affinity.","feed_subtitle":"Fine-tuned on approved drugs and rewarded by predicted affinity, the model beats DrugGPT across eight targets.","key_machinery":"The load-bearing mechanism is the two-phase training loop. First, DrugGPT is supervised fine-tuned on 9,398 sequence-SMILES strings built from 1,660 approved small molecules and their targets, so the model learns the chemical language of drugs that have already passed regulatory review. Second, the fine-tuned model is run through proximal policy optimization (PPO), where each generated SMILES string receives a reward equal to PLAPT's predicted negative log affinity, multiplied by a validity check that assigns zero to invalid structures and by a factor of 0.7 if the molecule duplicates an approved drug. A KL-divergence penalty keeps the policy close to the fine-tuned model. This reward function is what simultaneously enforces chemical validity, affinity, and a balance between novelty and repurposing.","core_discovery":"On the paper's own terms, the central discovery is that a GPT-style molecule generator can be steered toward chemically valid, high-affinity candidates by the right reward signal: supervised fine-tuning on approved drugs alone is not enough, but adding proximal policy optimization with a reward that combines a transformer-based affinity predictor (PLAPT), a rigid validity check, and a penalty for reproducing known approved molecules pushes the model to 100% validity and higher predicted affinity than the DrugGPT baseline. Across eight targets, including six with no known approved drugs, DrugGen's median PLAPT affinity was 7.22 [6.30-8.07] versus 5.81 [4.97-6.63] for DrugGPT, and for FABP5 and NAMPT the best docked poses beat the reference ligands' docking scores (FABP5/11 at -9.537 versus palmitic acid at -6.177). The paper also frames the evaluation metrics themselves—validity, diversity, novelty, and binding affinity—as a reusable standard for comparing future generative drug-design models.","pith_inferences":["The affinity gain is measured with the same PLAPT model used as the training reward, so part of the improvement could be reward overfitting; an independent affinity predictor or wet-lab binding data would be needed to confirm the gain is real.","Because DrugGen's novelty is lower than DrugGPT's (41.88% versus 66.84%), fine-tuning on approved drugs seems to pull the generator toward known drug-like space; whether this is a cost depends on whether the goal is de novo scaffold discovery or repurposing.","Ablating the reward components (validity penalty vs affinity reward vs repetition penalty) would show which term causes the validity jump; if it is mainly the penalty, a cheaper post-hoc filter might match DrugGen without reinforcement learning.","The docking validation covers only four targets and a handful of molecules, so a broader docking screen over all generated molecules would reveal how often high PLAPT scores correspond to strong predicted binding geometry."],"forward_implications":["If the PLAPT reward transfers to real binding, DrugGen-style training should increase the fraction of generated molecules that advance to synthesis and assay, since invalid structures are almost eliminated.","The pipeline applies to proteins with no known approved drugs, so it could be used to generate early leads for novel or understudied targets.","Because the reward is pluggable, the same PPO setup could optimize other objectives—synthesis feasibility, toxicity, solubility—by swapping in the corresponding predictor.","Docking results like the NAMPT novel pharmacophore suggest the method can propose chemically different scaffolds that still occupy the intended binding site, which is useful for escaping crowded chemical space.","The evaluation metrics (validity, Tanimoto-based diversity, novelty, and binding affinity) give later models a standard way to report generation quality."],"supporting_citations":[{"why":"the GPT-based molecule generator that DrugGen starts from and is compared against","marker":"[19]"},{"why":"the PLAPT model whose predicted binding affinity serves as both the PPO reward and the evaluation metric","marker":"[16]"},{"why":"the database that supplies the approved drug-target pairs used for fine-tuning","marker":"[31]"},{"why":"the trainer modules that implement the supervised fine-tuning and PPO optimization","marker":"[36]"},{"why":"the 'summarize from feedback' reinforcement-learning recipe on which the PPO training is based","marker":"[37]"},{"why":"the cheminformatics library used to build the invalid-structure assessor","marker":"[38]"},{"why":"the Glide docking program used to validate that generated molecules bind the intended sites","marker":"[46]"},{"why":"the prior study that identified the six targets without known approved drugs","marker":"[24]"}],"fun_headline_variants":["DrugGen achieves 100% valid molecules via RL feedback","RL fine-tuning makes DrugGen generate valid high-affinity molecules","DrugGen beats DrugGPT with 100% validity and better affinity","DrugGen uses RL to ensure 100% valid drug candidates","DrugGen's RL feedback yields perfect validity and stronger binding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that PLAPT's predicted binding affinity is a reliable proxy for true drug-target binding, because that same predictor is used to define the training reward and to score the final comparison; if optimizing against PLAPT merely exploits its blind spots, the reported affinity advantage would not survive contact with real assays.","fun_headline_variants_meta":{"raw":{"variants":["DrugGen achieves 100% valid molecules via RL feedback","RL fine-tuning makes DrugGen generate valid high-affinity molecules","DrugGen beats DrugGPT with 100% validity and better affinity","DrugGen uses RL to ensure 100% valid drug candidates","DrugGen's RL feedback yields perfect validity and stronger binding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1652,"prompt_tokens":1096,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":712,"completion_tokens_details":{"reasoning_tokens":471}},"tokens_in":712,"tokens_out":556,"duration_ms":6136,"temperature":1.0,"reasoning_tokens":471,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:00:55.212523+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the top-scoring DrugGen molecules for ACE, NAMPT, and FABP5 and measure their actual binding (for example, IC50 or Kd) alongside an equal number of DrugGPT molecules; if DrugGen's molecules do not bind more tightly, the central affinity claim is refuted. A faster in silico check is to re-run the PPO phase with a completely different affinity predictor and see whether the apparent gain persists.","supporting_citations":[{"cited_title":"DrugGPT: A GPT-based Strategy for Desi gning Potential Ligands Targeting Speciﬁc Proteins,","cited_arxiv_id":null,"evidence_quote":"the GPT-based molecule generator that DrugGen starts from and is compared against"},{"cited_title":"PLAPT: Protein-Ligand Binding Aﬃnity Prediction Using Pretraine d Transformers,","cited_arxiv_id":null,"evidence_quote":"the PLAPT model whose predicted binding affinity serves as both the PPO reward and the evaluation metric"},{"cited_title":"A vailable at: https://huggingface.co/docs/trl/en/index","cited_arxiv_id":null,"evidence_quote":"the trainer modules that implement the supervised fine-tuning and PPO optimization"},{"cited_title":"A vailable at:https://github.com/openai/summarize-from-feedback","cited_arxiv_id":null,"evidence_quote":"the 'summarize from feedback' reinforcement-learning recipe on which the PPO training is based"},{"cited_title":"A vailable at: https://www.rdkit.org/","cited_arxiv_id":null,"evidence_quote":"the cheminformatics library used to build the invalid-structure assessor"},{"cited_title":"and Banks, Jay L","cited_arxiv_id":null,"evidence_quote":"the Glide docking program used to validate that generated molecules bind the intended sites"},{"cited_title":"DrugTar Improves Druggability Pred iction by Integrating Large Language Models and Gene Ontologies,","cited_arxiv_id":null,"evidence_quote":"the prior study that identified the six targets without known approved drugs"}],"review_version":1}