{"id":"e1ff434d-8e56-4919-92a1-56c600d4195d","arxiv_id":"2412.10743","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A flow-based model, NeuralPLexer3, predicts biomolecular complex structures with improved physical validity and speed over AlphaFold3 on the PoseBusters benchmark, and introduces NPBench and ConfBench for broader evaluation.","lead":"NeuralPLexer3 is a flow-based machine learning model that predicts 3D structures of protein-ligand and other biomolecular complexes from sequence and chemical graphs. On the PoseBusters benchmark it reports a higher physically-valid success rate than AlphaFold3, with much faster inference, and introduces new benchmarks for interactions and conformational changes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline PoseBusters margin over AF3 may reflect NP3's validity-aware ranking rather than the generative model; AF3 baselines are not protocol-matched.","rationale":"The reader's weakest_assumption (AF3 numbers not rerun under the same protocol) is correct and is part of the problem, but the sharper load-bearing issue is where the reported advantage comes from: the two models are tied on pure coordinate accuracy, and NP3's only edge is PB-valid, which its own ranking function directly penalizes. This makes the 'greater accuracy than AF3' claim contingent on a ranking detail that is not matched in the published AF3 baseline. The concern is concrete and testable by an internal ablation; it does not allege any misreporting, and the paper's disclosure of the ranking penalty is a point in its favor. The system-level speed and NPBench/ConfBench results provide independent evidence of utility, so I do not recommend rejection. The verdict stays CONDITIONAL: the stated condition should include a common-protocol rerun and/or a ranking-ablation demonstration.","tokens_in":20318,"tokens_out":7591,"duration_ms":69828,"concrete_test":"Rerun NP3 on the exact PoseBusters-V2 targets with sample ranking changed from the S.5 penalized score to plain pLDDT (disabling the clash/chirality penalty), keeping sampling, alignment, and PB-valid scripts fixed. If the combined success rate drops from 78.4% to at or below AF3's 73.1%, the claimed accuracy advantage is attributable to the ranking penalty rather than to the flow model; if it stays above 73.1%, the ranking concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"NP3's quantitative superiority over AF3 rests on a 78.4% vs 73.1% combined PoseBusters success rate, but Table 1 shows the coordinate-accuracy component is effectively tied (80.2% vs 80.4% RMSD < 2 Å). The entire margin is in the PB-valid component, and NP3 explicitly optimizes that component at inference: SI S.5 scores samples by pLDDT minus 1000 times (clash + chirality-violation), and Figure 2C credits these penalties for PoseBusters gains. The AF3 baseline is not produced under the same protocol: Methods states baseline numbers are taken from Abramson et al. Extended Data Table 1, with no common rerun of dataset version, pocket-aligned RMSD, PB-valid filters, or sample ranking. The headline claim therefore currently cannot distinguish 'a more accurate generative model' from 'a validity-aware ranking postprocessor applied to a tied-accuracy model.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"NeuralPLexer3 (NP3) is a flow-matching generative model for all-atom biomolecular complex structure prediction. The model uses a Langevin-relaxed globular polymer prior, optimal-transport entity and atom permutation, rigid-alignment vector-field reparameterization, an encoder-decoder architecture with Flash-TriangularAttention, and a confidence module. The authors report state-of-the-art protein-ligand accuracy on PoseBusters (78.4% combined success vs 73.1% for AlphaFold3, with coordinate-only accuracy 80.2% vs 80.4%), 98.8% ligand stereochemistry accuracy, roughly 30 seconds of inference on one L40S GPU versus about six A100-minutes reported for AF3, broad interaction coverage on a new NPBench suite, and a new ConfBench benchmark for ligand-induced conformational change, where NP3 outperforms AlphaFold2-Multimer. They also present a scaling-law analysis of encoder/decoder capacity and decoder-replica count, and they release the NPBench code.","tokens_in":20490,"tokens_out":8711,"duration_ms":75738,"significance":"The paper is a substantial engineering contribution to biomolecular structure prediction. The flow-matching formulation with informative priors, the detailed ablations in Figure 2C, the open NPBench code, and the explicit sampling and training algorithms in the SI are valuable. The speed and physical-validity improvements are plausible and practically important for drug discovery. However, the headline superiority over AF3 is not established as stated: the AF3 baseline is not protocol-matched, the entire margin is in a PB-valid component that the conformer-ranking step explicitly optimizes, and the new NPBench and ConfBench benchmarks are self-built. With those caveats, the value of the work lies more in an efficient, broadly applicable predictor with comparable coordinate accuracy and strong physical validity than in a demonstrated accuracy advantage over AF3.","major_comments":[{"comment":"The central claim that NP3 improves protein-ligand binding structure prediction with greater accuracy than AF3 is not supported by a protocol-matched comparison. The Methods state that AF3 baseline numbers are taken from Abramson et al. Extended Data Table 1, and the paper does not rerun AF3 under its own evaluation pipeline. Because the coordinate-only success rate is essentially tied (80.2% vs 80.4% RMSD < 2 Å in Table 1), the reported 5.3-point margin in the combined PoseBusters metric rests entirely on the PB-valid component. To make the headline claim credible, the authors should either rerun AF3 with the identical PoseBusters-V2 dataset definition, the same pocket-aligned RMSD calculation, the same physical-validity filters, and the same sample-ranking protocol, or explicitly and prominently reframe the comparison as literature-based. In either case, confidence intervals or bootstrap error bars for the success rates are needed.","section":"Methods – Baselines; Table 1; Figure 1B"},{"comment":"The combined PoseBusters success rate is partly an artifact of the ranking procedure. SI S.5 defines Score = pLDDT(LG) – 1000 × (is_clash + is_chirality_violation), and Figure 2C explicitly credits the clash and chirality penalties for PoseBusters gains. Since PB-valid includes clash and chirality checks, selecting the top-ranked sample with this score directly optimizes the validity component of the combined metric. The paper should report per-sample or unranked PoseBusters accuracy, such as the accuracy of a randomly chosen sample or the full distribution over samples, and apply the same ranking protocol to both NP3 and AF3 before claiming that NP3's generative model is more accurate.","section":"SI S.5; Figure 2C"},{"comment":"Equation (3) defines Compute (FLOPs) = (α · β · P) · D with α and β described as encoder and decoder FLOPs. As written, the units are not FLOPs: if α and β are per-sample FLOPs, the total should be (α + β) · P · D. This inconsistency affects the compute-optimal frontier analysis in Figure 2B and the relative capacity comparison in Figure 2A; please correct the formula and verify that the scaling conclusions are unchanged.","section":"Results – Compute-optimal scaling; Eq. (3)"},{"comment":"The main-text ConfBench score in Eq. (4) differs from the SI formulas in S.7.2. With the main-text denominator sqrt(RMSD_alt^2 + RMSD_ref^2 + RMSD_altref^2) multiplied by 1/2, a query identical to the reference does not receive score 1, contradicting the stated interpretation; the SI version uses a denominator of sqrt((1/2)(...)), which does satisfy that boundary. This discrepancy changes the meaning of a ConfBench score greater than 0 and could affect the reported win rates of 51.9% versus 29.6% and the apo/holo correctness rates. Please reconcile the equations and re-verify the conformational statistics with the corrected scoring function.","section":"Methods – Conformational Benchmarks; Eq. (4); SI S.7.2 Eqs. (5)-(6)"}],"minor_comments":[{"comment":"In the baselines paragraph, 'Abramsom' should be 'Abramson'.","section":"Methods – Baselines"},{"comment":"The main text says NPBench contains structures released after 2023, while the Methods give a deposition window of 2022-05-01 to 2023-01-12; these dates should be reconciled.","section":"NPBench description; Methods – Structure Prediction Benchmarks"},{"comment":"The Figure 2C caption uses 'C: Compute in GFLOPs' and 'P: number of decoder replicas', but the text and Eq. (3) use α and β for encoder and decoder FLOPs and P for replicas; the notation should be made consistent.","section":"Figure 2C caption; Eq. (3)"},{"comment":"The inference-time comparison of roughly 30 seconds on one L40S GPU versus about six A100-minutes for AF3 is based on timing statistics reported in the AF3 paper on different hardware; this should be described as an approximate, non-protocol-matched comparison.","section":"Results – Computational Efficiency"},{"comment":"The paper should state more prominently that NPBench and ConfBench are newly introduced benchmarks built by the authors, and that independent external validation is pending; the release of NPBench code is helpful but does not by itself establish community acceptance of the benchmark.","section":"Results – NPBench and ConfBench"},{"comment":"The Discussion acknowledges limitations of ConfBench, namely imperfect scores and apo structures that are not always fully ligand-free; these limitations should also be stated next to the headline conformational win rates in the Results section.","section":"Discussion – Limitations of ConfBench"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the main-text claims should not be taken at face value until the AF3 baseline is protocol-matched or explicitly downgraded. The authors' own Figure 2C and SI S.5 make clear that the validity-aware ranking is responsible for a large part of the PoseBusters gain, so the editorial decision should not rest on the 78.4% versus 73.1% number. The core methodological contributions and the released NPBench code are valuable and justify a major revision rather than rejection, but the paper's central accuracy claim needs substantial rework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious systems paper with real engineering contributions, but the headline claim that NP3 beats AlphaFold3 on PoseBusters is not actually supported by the numbers as presented. The RMSD<2 Å component is a tie (80.2 vs 80.4); the entire margin comes from the PB-valid component, and NP3 explicitly optimizes that component during sample ranking via the 1000× clash/chirality penalty (SI S.5). The AF3 baseline numbers come straight from the published paper, with no common rerun of dataset version, alignment, filters, or ranking. So the current evidence supports 'a validity-aware ranker can raise combined PoseBusters success rates on top of a tied-accuracy model,' not 'a more accurate generative model.'\n\nThe genuinely new pieces are the globular polymer prior, the OT-based symmetry correction, Flash-TriangularAttention, the scaling analysis, and the two benchmarks. NPBench (1,143 recent PDB chains/interfaces, with cluster-based splits) and ConfBench (apo/holo conformational scores) address real gaps, and the curation protocols in SI are detailed and sensible. The ablation in Fig 2C is useful; it shows which components moved the metric and includes the ranking penalty as a major factor, which is honest.\n\nSoft spots, in proportion: (1) the missing protocol-matched AF3 baseline is the main one; it is exactly the comparison the abstract leads with. (2) No confidence intervals on any success rate, and the benchmarks are authored and scored by the same group, so the numbers should be treated as vendor-reported until someone else reruns them. (3) The model and weights are not released, so no independent reproduction is possible; NPBench code is open, which helps. (4) The conformer ranking penalty is a post-hoc filter that directly targets the headline metric; that is not a flaw per se, but it means the reported margin conflates generation and ranking.\n\nThis is not a takedown. For readers interested in structure prediction systems and benchmarking, the paper is worth reading, and the benchmarks will probably be useful. But the central SOTA claim is provisional. I would send it to peer review — the work is detailed enough to deserve referee time — and instruct the referee to require a common-protocol AF3 rerun and a breakdown of success rates with and without the ranking penalty. Also ask for confidence intervals. The paper could become a solid contribution after that.","headline":"Serious systems paper with useful benchmarks, but the AF3 beat rests on a ranking postprocessor and unmatched baselines, so the SOTA claim is provisional.","tokens_in":21086,"tokens_out":3184,"would_cite":true,"duration_ms":29060,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"NeuralPLexer3 claims state-of-the-art biomolecular complex prediction using flow-based generative modeling, beating AlphaFold3 on combined PoseBusters accuracy while running in roughly 30 seconds per prediction.","keywords":["protein–ligand structure prediction","flow matching","continuous normalizing flows","biomolecular complex prediction","PoseBusters benchmark","conformational change prediction","structure-based drug design","physical validity"],"falsifier":"A head-to-head rerun of AF3 and NP3 on the same PoseBusters-V2 targets with identical code for pocket-aligned RMSD, PB-valid checks, and conformer ranking would settle the claim; if AF3's combined success rate under that protocol exceeds 78.4% or NP3's falls below 73.1%, the stated advantage is reversed.","tokens_in":20084,"feed_emoji":"🧬","tokens_out":8140,"duration_ms":67654,"temperature":0.7,"pith_summary":"NeuralPLexer3 (NP3) is a flow-based generative model that predicts the 3D structure of biomolecular complexes—proteins, nucleic acids, ligands, ions, and post-translational modifications—from sequence and molecular topology alone. The paper's central claim is that NP3 outperforms AlphaFold3 on the combined PoseBusters protein–ligand benchmark, with a 78.4% success rate requiring both sub-2 Å RMSD and physical validity against AF3's 73.1%, while running a prediction in roughly 30 seconds on one L40S GPU rather than several A100-minutes. The work also introduces two new evaluation resources: NPBench, for low-homology recent PDB structures covering diverse interaction types and stoichiometry, and ConfBench, for apo/holo ligand-induced conformational changes. A sympathetic reader should care because these are the capabilities drug discovery actually needs: physically plausible poses, generalization to novel targets, conformational sensitivity, and speed.","feed_headline":"Flow model beats AlphaFold3 on drug-binding poses","feed_subtitle":"78.4% physically valid poses in seconds per target, plus benchmarks for conformational change and diverse interactions.","key_machinery":"The central object is a continuous normalizing flow (CNF) trained by flow matching, which maps a simple prior distribution to the distribution of all heavy-atom coordinates of a biomolecular complex by learning a velocity field. Three mechanisms carry the argument: a physics-inspired globular polymer prior (random atom configurations relaxed by a short Langevin dynamics with harmonic connectivity and confinement terms) that starts sampling from chemically sensible structures; a simulation-free optimal-transport symmetry correction that permutes equivalent entities and atoms so the conditional flow trajectories between prior and ground truth are straightened; and a vector-field reparameterization that predicts denoised coordinates and applies optimal rigid alignment, reducing the number of integrator steps to 40. The architecture itself is an encoder–decoder transformer with anchor-level conditioning, MSA and language-model embeddings, and Flash-TriangularAttention, a kernel that avoids explicit bias broadcasting in the QK^T + bias operation and cuts peak memory by about 5x so training can use larger structure crops.","core_discovery":"On its own terms, the paper establishes that a conditional flow-matching model with a physics-informed prior can be the best current sequence-only predictor of biomolecular interactions: NP3 achieves 78.4% combined PoseBusters success versus 73.1% for AF3 and 98.8% ligand stereochemistry accuracy, matches AF3 on raw coordinate error (80.2% vs 80.4%), and outperforms AlphaFold2-Multimer on protein–peptide interfaces and on apo/holo conformational-change prediction while remaining competitive on monomers, PPIs, and CASP15 RNA. The authors attribute this to flow matching rather than diffusion: an informative globular-polymer prior relaxed by Langevin dynamics, optimal-transport-based symmetry correction that straightens conditional flows, prediction of denoised coordinates with rigid alignment, and a 40-step sampler that removes expensive diffusion rollouts. They further report compute-optimal scaling behavior with an encoder/decoder FLOP ratio near 10 and 20 decoder replicas, and describe Flash-TriangularAttention as the memory optimization that allows large-crop training.","pith_inferences":["Editorial inference: if the timing comparison is done on matched hardware and includes the full sampling and ranking pipeline, the reported seconds-per-target speed would enable enumerating hundreds of conformers per ligand, which current diffusion models cannot afford; the paper only reports a single-inference time, so the end-to-end gain is likely smaller but still substantial.","Editorial inference: the AF3 comparison rests on published benchmark tables rather than a local rerun; a protocol-matched head-to-head could shift the 5.3-point gap in either direction, so the durable claim is that NP3 is competitive and faster, not necessarily that it wins by that exact margin.","Editorial inference: the ConfBench scoring function is transferable: any structure prediction model could be scored on the same apo/holo pairs, making 'does this model see ligand-induced motion?' a standard quantitative question rather than a qualitative one.","Editorial inference: the success of RNA language-model conditioning suggests that for organisms with sparse MSA coverage, structure prediction may rely more on learned sequence embeddings, which could extend NP3-style models to non-model organisms."],"forward_implications":["If the 78.4% versus 73.1% PoseBusters result holds, sequence-only prediction has overtaken the previous best model on the combined accuracy-plus-physical-validity metric that matters for structure-based drug design.","A prediction in about 30 seconds on a single L40S GPU would make virtual screening of large compound libraries practical, since the cost per target-ligand pair drops by roughly two orders of magnitude relative to the reported AF3 timing.","The 98.8% ligand stereochemistry accuracy and PB-valid rates imply predicted poses need less post-hoc filtering before use in medicinal chemistry.","The ConfBench results imply the model can sometimes track ligand-induced conformational changes, including apo-to-holo transitions in kinases, which is a prerequisite for predicting allosteric effects and induced-fit selectivity.","NPBench's low-homology recent-PDB evaluation provides a reusable standard for testing generalization to unseen chains, ligands, and stoichiometries, and the paper's RNA results suggest MSA-free conditioning with language models can approach MSA-based accuracy."],"supporting_citations":[{"why":"Defines the strongest baseline and supplies the AF3 PoseBusters numbers and benchmark protocol that NP3 claims to beat.","marker":"[1]"},{"why":"Provides the benchmark dataset and physical-validity (PB-valid) checks that define the headline 78.4% versus 73.1% success rates.","marker":"[9]"},{"why":"Supplies the conditional-flow-matching training framework and ODE sampling that NP3's generative model is built on.","marker":"[10]"},{"why":"Predecessor method whose broad biomolecular coverage and accuracy NP3 extends and compares against.","marker":"[2]"},{"why":"Earlier protein-ligand model that contributes the rigid-alignment and multiscale generative ideas reused in NP3.","marker":"[4]"},{"why":"Provides the optimal-transport principle used by NP3's symmetry-correction module to straighten flow trajectories.","marker":"[12]"},{"why":"Source dataset and linkage definitions that ConfBench builds on for apo/holo conformational-change pairs.","marker":"[29]"},{"why":"Baseline for NPBench protein-protein and protein-nucleic-acid interface comparisons and for ConfBench conformational accuracy.","marker":"[31]"},{"why":"Supplies the DockQ metric and chain-mapping tools used to score polymer-polymer interfaces in NPBench.","marker":"[30]"},{"why":"Used to generate the AlphaFold2-Multimer baseline predictions on NPBench targets.","marker":"[8]"}],"fun_headline_variants":["Flow model beats AlphaFold3 on valid drug-binding pose rate","NeuralPLexer3: physics-informed flow model tops AF3 pose validity","78.4% valid poses: flow-based NeuralPLexer3 outperforms AF3","Flow matching yields better biomolecular complex poses than diffusion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparison assumes that AlphaFold3's published PoseBusters numbers were computed under the same benchmark protocol—same dataset version, same pocket-aligned RMSD, same physical-validity filters—as the numbers reported for NeuralPLexer3, because the paper does not rerun AlphaFold3 itself.","fun_headline_variants_meta":{"raw":{"variants":["Flow model beats AlphaFold3 on valid drug-binding pose rate","NeuralPLexer3: physics-informed flow model tops AF3 pose validity","78.4% valid poses: flow-based NeuralPLexer3 outperforms AF3","Flow matching yields better biomolecular complex poses than diffusion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000321,"raw_usage":{"total_tokens":1773,"prompt_tokens":880,"completion_tokens":893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":814}},"tokens_in":496,"tokens_out":893,"duration_ms":8288,"temperature":1.0,"reasoning_tokens":814,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:39:00.435035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A head-to-head rerun of AF3 and NP3 on the same PoseBusters-V2 targets with identical code for pocket-aligned RMSD, PB-valid checks, and conformer ranking would settle the claim; if AF3's combined success rate under that protocol exceeds 78.4% or NP3's falls below 73.1%, the stated advantage is reversed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the strongest baseline and supplies the AF3 PoseBusters numbers and benchmark protocol that NP3 claims to beat."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the benchmark dataset and physical-validity (PB-valid) checks that define the headline 78.4% versus 73.1% success rates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier protein-ligand model that contributes the rigid-alignment and multiscale generative ideas reused in NP3."},{"cited_title":"& Elofsson, A","cited_arxiv_id":null,"evidence_quote":"Used to generate the AlphaFold2-Multimer baseline predictions on NPBench targets."}],"review_version":1}