{"id":"56e3a570-dc9a-4f56-b142-725b773a123d","arxiv_id":"2509.02642","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"BioMD generates realistic all-atom protein-ligand dynamics trajectories by flow-matching over a two-stage forecast-then-interpolate schedule, reaching an unbinding success of 97.1% on the DD-13M test set.","lead":"BioMD, a new machine learning model, generates atom-level trajectories of protein-ligand complexes, including complete drug unbinding paths, in seconds rather than hours. It forecasts coarse time steps and then interpolates the intermediate frames with a single flow matching network.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 97.1% unbinding success rate rests on a permissive convex-hull criterion that may count AR drift as dissociation; Table 2's own RMSD and clash numbers are consistent with this artifact.","rationale":"The reader's weakest assumption exactly identifies the load-bearing condition: the convex-hull exit metric must correspond to physical unbinding. My review of the full text confirms this is the most fragile step in the argument. The strongest evidence BioMD provides—MISATO RMSF correlations on an external dataset, and the qualitative 6EY8 pathways matching metadynamics—does not rescue the quantitative unbinding claim, because the 97.1% figure is computed with a geometric criterion that is insensitive to whether the exit is a genuine transition or a byproduct of accumulated error. Table 2 supplies internal warning signs: BioMD-rel (AR-5)'s unbinding path RMSD is worse than Static, and its clash score is much higher, while Sec. 5.3 admits AR error accumulation. The paper's lack of energetic validation (no PMF, no kinetic comparison, no contact/solvent analysis) means the gap between 'centroid outside convex hull' and 'dissociation pathway' is unaddressed. This does not refute the method—it may still be a useful generative trajectory model—but it does mean the central headline should be treated as unverified. I therefore recommend keeping the reader's CONDITIONAL verdict (UNCHANGED), since the reader already conditioned acceptance on validating the unbinding metric. A single concrete re-evaluation of the success cases with sustained-exit and distance criteria, plus a random-walk control, would settle whether the concern lands.","tokens_in":14264,"tokens_out":2670,"duration_ms":35024,"concrete_test":"For the DD-13M test set, take the trajectories that BioMD-rel (AR-5) counts as successes under the convex-hull criterion. Re-evaluate each using a stricter physical dissociation test: (i) after first hull exit, require the ligand centroid to remain beyond the hull for at least 20 consecutive frames (or >25% of remaining frames), and (ii) require at least one post-exit frame with protein-ligand minimum heavy-atom distance > 4 Å, and (iii) require zero steric clashes (threshold 1.5 Å) at the exit frame. Additionally, run a negative control: generate ligand random-walk trajectories from the bound pose with per-step displacement calibrated to BioMD-rel (AR-5)'s observed per-frame RMSD drift, and apply the same convex-hull criterion. If the negative control achieves comparable success rates, or if the stricter criteria collapse the reported @1/@5/@10 rates, the headline success claim is a met","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that BioMD generates complete ligand unbinding pathways (Sec. 5.2, Table 2)—depends entirely on the Sec. A.3.4 definition of success: a trajectory is a successful unbinding if at least one predicted ligand centroid lies outside the convex hull of the protein's heavy atoms in the initial bound state. This criterion is necessary for dissociation but far from sufficient. A ligand can cross the hull by slow accumulated drift, by sliding along the protein surface, or by AR error accumulation, without ever overcoming the physical dissociation barrier. The paper's own numbers expose this vulnerability: the variant with the highest success, BioMD-rel (AR-5), has an unbinding path RMSD (0.7055 Å) worse than the Static baseline (0.6504 Å), indicating that its 'successful' trajectories are, on average, no closer to reference unbinding paths than a no-motion baseline. The same variant also shows a protein-ligand steric clash score (0.6375) far above Static (0), consistent with distorted geometries rather than clean egress. Section 5.3 explicitly documents AR error accumulation. No energetic validation is offered: no free-energy profiles, no residence-time comparison, no contact analysis at the supposed exit events, and no demonstration that identified exits are irreversible dissociations rather than transient geometric escapes. If many of the 70.9/92.9/97.1% successes are drift-generated crossings, the headline claim is substantially weakened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BioMD, an all-atom generative model for protein-ligand dynamics, built on a unified conditional flow-matching framework with hierarchical coarse-grained forecasting and fine-grained interpolation. The model is evaluated on two datasets: MISATO for ligand dynamics in the binding pocket and DD-13M for ligand unbinding. The authors report high physical stability and competitive RMSF correlations on MISATO, and on DD-13M report that the auto-regressive variant BioMD-rel (AR-5) achieves 70.9%/92.9%/97.1% unbinding success at @1/@5/@10 attempts, with the claim that BioMD is the first all-atom generative model to simulate long-timescale ligand-unbinding pathways. The central scientific claim is the 97.1% unbinding success and the physical meaningfulness of the generated unbinding paths.","tokens_in":14542,"tokens_out":3860,"duration_ms":46067,"significance":"If the central claim is correct, the work would be a notable advance: a generative model that produces plausible all-atom trajectories, including ligand dissociation events, at orders-of-magnitude lower cost than metadynamics. The methodological core is sensible: using a single flow-matching model with different masking schedules for forecasting and interpolation is elegant, and the architecture is described in sufficient detail to be reproducible. The qualitative 6EY8 example, in which BioMD reproduces two known metadynamics pathways and finds a third, is compelling but anecdotal. The paper also honestly acknowledges the error-accumulation problem in Section 5.3. The main weakness is that the headline unbinding metric is not physically validated, and the paper's own path-RMSD and clash numbers are consistent with the concern that many 'successful' unbinding events are artifacts of cumulative AR drift or distorted geometries rather than genuine dissociation transitions. The result is potentially important, but the current empirical support is insufficient to sustain the strong conclusions.","major_comments":[{"comment":"The unbinding success metric—defined as the ligand centroid leaving the convex hull of the initial protein heavy atoms—is necessary but not sufficient for a physical dissociation pathway. The paper's own numbers in Table 2 are consistent with the concern that convex-hull exits are produced by cumulative drift: BioMD-rel (AR-5), which has the highest success rate, has an unbinding path RMSD (0.7055 Å) worse than the Static baseline (0.6504 Å) and an protein-ligand clash score (0.6375) far above Static (0). Section 5.3 documents AR error accumulation. The authors should either replace this metric with a stricter one (e.g., sustained outward displacement of the ligand from the binding site, requiring the exit to be irreversible and not solely occurring in the last frame) or provide additional physical validation: contact maps at the exit event, ligand-pocket distance profiles, free-energy o","section":"Sec. A.3.4 / Table 2"},{"comment":"The Unbinding Path RMSD metric needs clarification and appears internally inconsistent. The DD-13M dataset consists of dissociation trajectories, so a static ligand centroid should be very far from a reference unbinding path. Yet the Static baseline reports an RMSD of only 0.6504 Å, comparable to the best BioMD variants. This suggests either that the metric's best-match/resampling procedure is not sensitive to unbinding, or that the reference trajectories used for comparison are not predominantly unbinding-like, or that an alignment step removes the signal. The authors should specify the alignment, how trajectories are resampled, whether only successful reference dissociation paths are used, and how the best-match search is performed. As reported, Table 2's path-accuracy column does not provide meaningful evidence for the physical fidelity of the generated unbinding paths.","section":"Table 2 / Sec. A.3.3"},{"comment":"The auto-regressive variant that achieves the headline unbinding success also shows substantially degraded physical stability: bond MAE increases from 0.0308 (BioMD-rel) to 0.0580 (BioMD-rel AR-5), angle MAE from 0.0606 to 0.0918, and protein-ligand clashes from 0.0004 to 0.6375 (Table 2). Section 5.3 acknowledges this but argues the errors are correctable by 'a simple local refinement step' without demonstrating that such refinement preserves the unbinding pathways. If the reported 97.1% success is computed on unrefined trajectories where the ligand may be sterically clashing with the protein, the physical interpretation of 'complete unbinding paths' is questionable. The authors should report results after the proposed refinement, or at minimum show that the geometric distortions do not drive the hull-exit events.","section":"Sec. 5.3 / Table 2"},{"comment":"The Euler solver uses dt=0.1 with 10 integration steps (Algorithm 3: τ in {0, 0.1, ..., 0.9}), and the paper does not provide a sensitivity analysis of success/path metrics to the integrator step size. Since the headline results depend on long autoregressive generation, the numerical integration scheme is a potentially important source of error. A convergence check or a comparison with a smaller dt would strengthen the claim that the generated trajectories are stable and not numerical artifacts.","section":"Sec. A.1 / Algorithms 3-4"}],"minor_comments":[{"comment":"The claim 'first all-atom generative model to simulate long-timescale protein-ligand dynamics' is strong. Given that MDGen and other trajectory models exist, the 'first' claim should be qualified with respect to protein-ligand systems specifically and justified in the related-work section.","section":"Abstract / Sec. 1"},{"comment":"Typo: 'nmasked' should be 'unmasked'.","section":"Sec. 4.2.1"},{"comment":"Typo: 'Indepentent' should be 'Independent'.","section":"Algorithm 2"},{"comment":"The 'best-match search' over reference trajectories needs a precise statement of whether the same reference path can be selected for multiple generated trajectories and how the reported mean is aggregated over complexes; otherwise the RMSD values are hard to interpret.","section":"A.3.3"},{"comment":"The RMSF correlation for BioMD-abs on protein atoms (0.685) is good, but the authors do not report confidence intervals or significance tests against NeuralMD; adding these would strengthen the comparison.","section":"Sec. 5.1 / Table 1"}],"recommendation":"major_revision","confidential_remarks":"The DD-13M dataset and its generative predecessor are from the same group (ref [18], overlapping authors), and the primary success metric is self-defined. This is not disqualifying by itself, but it increases the burden on objective, physics-based validation. The paper's own Table 2 numbers (Static path RMSD 0.6504; AR-5 path RMSD 0.7055 with high clashes) are difficult to reconcile with the claim of successful unbinding-path generation. If the authors can provide trajectory-level evidence that the hull exits correspond to genuine dissociation events (e.g., sustained outward motion, contact breaking, and no re-binding), the work could become publishable; otherwise the central claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"BioMD is a well-built generative model for all-atom protein-ligand trajectories, with a genuinely useful design: it splits long trajectory generation into coarse forecasting and fine interpolation, both handled by one flow-matching transformer with different masking schedules. That is a real advance over single-step or protein-only existing models. On the MISATO dataset the model produces stable structures and captures protein flexibility better than the baselines (RMSF correlation 0.685 for protein, which no other method achieves). The speed gain (seconds vs ~1 hour for metadynamics on 6EY8) is real, though expected for an emulator.\n\nThe soft spot is the unbinding claim. The success criterion (A.3.4) is just whether the ligand centroid leaves the convex hull of the initial protein. That's necessary but nowhere near sufficient for a dissociation path. Table 2 shows BioMD-rel (AR-5) — the most successful variant with 97.1% @10 — actually has a worse unbinding path RMSD (0.7055 Å) than the Static baseline (0.6504 Å), and a high protein-ligand clash score (0.6375 vs 0 for Static). So many of those 'successes' look like AR drift and distorted geometries, not physical egress. The paper itself acknowledges error accumulation (Sec 5.3) but never validates the exits with any energetic or kinetic measure: no free-energy profiles, no residence times, no contact analysis at the supposed exit events. Given that the headline result is the 97.1%, this is a load-bearing gap.\n\nTwo smaller issues. The stability metrics overlap with the auxiliary geometric losses (Lgeom) used in training, so those numbers are partly self-evaluation. And the novelty claim relative to ref [18] (the same group's DD-13M dataset and generative model) is not made explicit; that should be clarified.\n\nNone of this is fatal. The method is coherent, the MISATO results are credible, and the authors are transparent about the AR error problem. The unbinding evaluation needs to be redone or at least supplemented with physics-based checks, and ideally with error bars and code. If that's done, the paper could be a solid contribution.\n\nI'd send this to peer review rather than desk-reject it, but the reviewers should push hard on the unbinding metric. It's worth a reading-group discussion on how to validate generative MD. I wouldn't cite it yet, but I'd keep an eye on the revision.","headline":"A well-engineered generative model for all-atom trajectories whose headline 97.1% unbinding rate is likely inflated by a weak convex-hull criterion and AR drift.","tokens_in":15140,"tokens_out":3158,"would_cite":false,"duration_ms":34303,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BioMD claims to be the first all-atom generative model that simulates long-timescale protein-ligand dynamics, generating physically plausible all-atom trajectories and complete ligand unbinding pathways for up to 97.1% of protein-ligand sys","keywords":["all-atom generative model","molecular dynamics","ligand unbinding","trajectory generation","flow matching","protein-ligand interactions","hierarchical generation","conformational flexibility"],"falsifier":"Re-run molecular dynamics or enhanced sampling seeded from BioMD's predicted exit structures and check whether they are metastable dissociated states; if the predicted exits collapse back to the binding site or sit in high free-energy barriers, the claimed unbinding paths fail. A faster check: apply the same convex-hull success metric to random-walk trajectories and compare the resulting Success@10 rate.","tokens_in":14060,"feed_emoji":"🧬","tokens_out":5975,"duration_ms":59226,"temperature":0.7,"pith_summary":"BioMD attempts to show that a single all-atom generative model, trained on molecular dynamics trajectories, can produce long-timescale protein-ligand dynamics that currently require expensive enhanced-sampling simulations. The paper's central claim is that by splitting trajectory generation into coarse forecasting followed by fine interpolation inside one flow-matching model, the method keeps long trajectories physically plausible while staying fast. On the DD-13M ligand-unbinding set, the auto-regressive variant claims successful exit paths for 97.1% of test complexes within ten generation attempts; on MISATO it reports RMSF correlations of 0.685 for protein atoms and 0.486 for ligands, capturing flexibility that static or protein-only baselines miss. If these results hold, BioMD would turn a slow rare-event calculation into a fast sampling problem.","feed_headline":"BioMD draws complete protein-ligand unbinding paths in seconds","feed_subtitle":"Full protein-ligand trajectories, unbinding included, now take seconds instead of the hour metadynamics needs.","key_machinery":"The central object is a unified trajectory flow-matching model with a 'noising-as-masking' schedule: each frame is independently noised according to a time variable, so known conditioning frames are kept clean and frames to be generated are initialized from noise, with different masking schedules realizing coarse-grained forecasting or fine-grained interpolation within the same architecture. This reduces sequence length by decoupling long-term evolution from local dynamics and helps manage error accumulation for long trajectories. An SE(3)-equivariant graph transformer encodes the initial conformation as conditional embeddings, and the FlowTrajectoryTransformer uses AttentionPairBias for int","core_discovery":"BioMD's core discovery claim is that long-timescale biomolecular motion can be generated all-atom by one conditional flow-matching velocity network, with the distinction between forecasting and interpolation encoded only as different noise/masking schedules. The model takes the first frame as a clean condition, denoises coarse frames spaced every k steps, then refills the gaps from two clean anchors; at inference the auto-regressive mode reconditions on its own history. The paper reports that this scheme yields complete ligand unbinding trajectories for 92.9% of systems within five attempts and 97.1% within ten, while also reproducing two known metadynamics pathways on 6EY8 and discovering a","pith_inferences":["If convex-hull exits are not validated energetically, the reported success rates may partly reflect accumulated drift; a fair test would compare free-energy profiles or committor probabilities of generated exits against reference dissociation paths.","The forecasting-then-interpolation design suggests a general recipe for other rare-event trajectory tasks, such as folding or conformational transitions, where coarse milestones are known but fine dynamics are not.","BioMD-abs versus BioMD-rel trade-off implies a Pareto frontier: the same model family can be pushed toward accurate reproduction or toward exploration, and a tuned interpolation may combine both.","Because the model is trained on simulated trajectories, its 'discovered' novel pathways are only as good as the training distribution's coverage; genuinely new pathways would need experimental or enhanced-sampling confirmation."],"forward_implications":["Long-timescale ligand dissociation, normally requiring enhanced sampling, becomes a direct generative sampling problem.","The same architecture handles both coordinate accuracy (BioMD-abs) and exploratory sampling (BioMD-rel) by switching prediction target or masking schedule.","Generated trajectories retain local chemical plausibility, with bond and angle errors below thermal fluctuation thresholds, even with autoregressive error accumulation.","The method can run in seconds on a single GPU, making pathway exploration practical for drug-discovery screening.","Protein flexibility, not just ligand motion, is captured when many baselines treat the receptor as static."],"supporting_citations":[{"why":"Supplies the noising-as-masking schedule that lets one model switch between forecasting and interpolation.","marker":"[5]"},{"why":"Motivates the all-atom velocity-network design that operates on full Cartesian coordinates.","marker":"[1]"},{"why":"Defines the flow-matching training objective that the trajectory velocity model minimizes.","marker":"[19]"},{"why":"Supplies the MISATO protein-ligand trajectory dataset used to test conformational flexibility.","marker":"[27]"},{"why":"Provides the DD-13M unbinding trajectories and the metadynamics-derived reference used to compute unbinding success.","marker":"[18]"},{"why":"Baseline that models protein-ligand dynamics with a static protein, which BioMD's RMSF results are compared against.","marker":"[20]"},{"why":"Baseline trajectory model limited to peptides and proteins, setting the protein-ligand gap BioMD targets.","marker":"[13]"},{"why":"Metadynamics reference for enhanced sampling that produces the unbinding pathways and the runtime comparison.","marker":"[16]"}],"fun_headline_variants":["BioMD nets 97% success on protein-ligand unbinding paths","One generative model forecasts and interpolates long MD trajectories","All-atom flow matching speeds protein-ligand unbinding simulation","BioMD: 97.1% unbinding routes from a single starting conformation","Hierarchical generator mimics MD to produce ligand exit paths"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that a ligand counted as 'unbound' because its center moved outside the protein's starting shape has really dissociated, not just drifted or distorted as the model's own error-accumulation numbers suggest.","fun_headline_variants_meta":{"raw":{"variants":["BioMD nets 97% success on protein-ligand unbinding paths","One generative model forecasts and interpolates long MD trajectories","All-atom flow matching speeds protein-ligand unbinding simulation","BioMD: 97.1% unbinding routes from a single starting conformation","Hierarchical generator mimics MD to produce ligand exit paths"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001198,"raw_usage":{"total_tokens":4772,"prompt_tokens":735,"completion_tokens":4037,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":3947}},"tokens_in":479,"tokens_out":4037,"duration_ms":34364,"temperature":1.0,"reasoning_tokens":3947,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:58:18.621230+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run molecular dynamics or enhanced sampling seeded from BioMD's predicted exit structures and check whether they are metastable dissociated states; if the predicted exits collapse back to the binding site or sit in high free-energy barriers, the claimed unbinding paths fail. A faster check: apply the same convex-hull success metric to random-walk trajectories and compare the resulting Success@10 rate.","supporting_citations":[],"review_version":1}